Sign in

Egor Marin

@marinegor.bsky.social
301 followers 604 following 145 posts

ML Engineer in Cofolding @ Apheris (Berlin, Germany) computational biology, ML, protein design, cheminformatics, fancy dev tooling, tinge of bouldering marinegor.dev

PostsRepliesMedia
Reposted by Egor Marin
Itai Yanai @itaiyanai.bsky.social · 12h
1. CRISPR came from basic research on bacteria. 2. Optogenetics started with a discovery in algae. 3. The GFP protein comes from jellyfish. The biggest breakthroughs all trace to scientists following their curiosity, and this is one huge reason why we must keep funding basic research!
9463167
Egor Marin @marinegor.bsky.social · 15h
python 3.15 looks so good! especially curious to see what C ABI improvements will bring to projects like maturin, unfortunately I'm not that knowledgeable to figure it out myself :(
000
Egor Marin @marinegor.bsky.social · 07/10/2026
If I'd still lecture macromolecular crystallography course, I'd definitely use that for demonstrations
040
Egor Marin @marinegor.bsky.social · 28/09/2026
favourite notebook + favourite package manager = 💞
010
Egor Marin @marinegor.bsky.social · 25/09/2026
following worst dev practices and deploying on friday!
001
Egor Marin @marinegor.bsky.social · 25/09/2026
gosh I love cheese
010
Egor Marin @marinegor.bsky.social · 18/09/2026
Ngl there's something epic to this
git diff of two versions of text -- first one shows user's employment at company called ENPICOM, which is in the past, and second one at company called Apheris, the current one
000
Egor Marin @marinegor.bsky.social · 16/09/2026
cool name for a indie band though
100
Egor Marin @marinegor.bsky.social · 14/09/2026
Oh that's about us, I work there! It was genuinely surprising to see companies actually collaborate with actual data (20,000 structures, can you imagine) to improve folding models. Naively I'd imagine that it's not what any company would do, since it's their hard-earned data😰
020
Egor Marin @marinegor.bsky.social · 12/09/2026
true! I have different attitude towards specifically cheminformatics in rust since my rust interest sparkled with Rich Apodaca's projects, which are sadly never came to life as the author intended. But I see a great future for the technology nevertheless :)
110
Egor Marin @marinegor.bsky.social · 12/09/2026
that reminds me of BPE tokenizer properties a bit, ie you can't have an arbitrary dictionary for them: github.blog/ai-and-ml/ll...
020
Egor Marin @marinegor.bsky.social · 12/09/2026
how realistic do you think that rust will eventually take over c++ for cheminformatics? There have been few attempts to do "rdkit but in rust", none of them came even close to 100 github stars. I wonder what's stopping them?
101
Egor Marin @marinegor.bsky.social · 11/09/2026
Linear baselines FTW!
010
Egor Marin @marinegor.bsky.social · 05/09/2026
I have been a fan of @rs-station.bsky.social 's approach to software for a few years now, and now I start admiring your approach to structural biology! I feel like ensembles are somewhat underexplored in modern structural biology and especially crystallography. Hope there'll be more of them soon :)
000
Egor Marin @marinegor.bsky.social · 05/09/2026
very much! not sure what's your use-case, but there's a tool (shameless plug -- I am its co-developer) likely not as good as structure-based numbering, but aims to be (on at least not VHH/VNAR sequences) on par with classical ANARCI. And it has a private in-browser demo: immunum.enpicom.com/demo
immunum.enpicom.com
immunum: antibody numbering and segmentation
030
Reposted by Egor Marin
Etowah Adams @etowah0.bsky.social · 21/08/2026
OpenBind intends to collect 10,000s of protein-ligand structures & affinities. To prioritize what we collect next, we need cofolding models trained on the latest data. Today we're releasing OpenBind-0 and 717 new ligand-bound structures.
1154
Egor Marin @marinegor.bsky.social · 26/08/2026
here's the competition page: marimo.io/pages/events...
marimo.io
Bring Cheminformatics to Life: molab Notebook Competition #3
OpenADMET x marimo notebook competition, co-hosted with Pat Walters. Pick an ADMET dataset, build a marimo notebook, and win prizes. Enter by October 4, 2026.
010
Egor Marin @marinegor.bsky.social · 19/08/2026
Starting this week, I'm joining Apheris (apheris.com). The company builds big federated training networks, which is cool, but the new job title is even cooler: Forward-Deployed ML Engineer in Cofolding (longest title I've had so far). Not sure if that has any practival implications, but we'll see😁
010
Egor Marin @marinegor.bsky.social · 19/08/2026
I figured bsky is not really the place for flashy job updates, but I'm actually getting back to structural biology, and feel very excited about that🫨
apheris.com
Superior drug discovery models
Superior drug discovery models through federated data networks
110
Egor Marin @marinegor.bsky.social · 14/08/2026
cpawthon🐾
010
Egor Marin @marinegor.bsky.social · 01/07/2026
Salary gives people incentive to work faster and actually deliver on the deadlines, who would have thought.
020
Egor Marin @marinegor.bsky.social · 09/06/2026
*why god why, I'd ask
010
Egor Marin @marinegor.bsky.social · 22/04/2026
we also had a very weird example of "oh every n-th layer of membrane protein crystal is point the other way while we do XFEL crystallography", but I doubt it's the case here.
010
Egor Marin @marinegor.bsky.social · 22/04/2026
some twinning perhaps? eg perfect or near-perfect twin can cause similar behaviour (normal maps with high rfactors).
210
Egor Marin @marinegor.bsky.social · 16/04/2026
>Any suggestions on setups for such storage that has worked well for folks in the past? create 2 Tb swapfile and enjoy limitless RAM (but make sure you replace these drives often enough since their lifetime will be limited).
010
Reposted by Egor Marin
Phil Ewels @ewels.bsky.social · 02/04/2026
Super excited to be launching two things today: #RustQC 🦀🧬 and rewrites.bio 🚀 I used AI to rewrite 15 RNA-seq QC tools into a single Rust binary (I've never written any Rust). It ended up being over 60x faster. Here's the story 🧵 seqeralabs.github.io/RustQC/
seqeralabs.github.io
Welcome to RustQC
Fast quality control tools for sequencing data, written in Rust.
38536
Egor Marin @marinegor.bsky.social · 30/03/2026
shamelessly tagging @delalamo.xyz here since I know you as a person who I might be interested in that :)
010
Egor Marin @marinegor.bsky.social · 30/03/2026
open-source details: - MIT license - CI/CD + continuous benchmarks - dedication for maintenance -- we have internal roadmap for the package, as we're using it internally as well, so it won't go unmaintained after 1.0 release - plans for R and duckdb bindings (stay tuned!)
110
Egor Marin @marinegor.bsky.social · 30/03/2026
Rarely show my work stuff here, but we did something cool (and open-source!) last week: github.com/ENPICOM/immu... TL;DR: - antibody numbering and segmentation with Rust - bindings to python, polars and WASM - VERY fast numbering at scale (got up to 1,000,000 seqs per second on 48 CPUs)
github.com
GitHub - ENPICOM/immunum: A high-performance antibody and TCR sequence numbering tool for Rust, Python, Polars and JS/TS.
A high-performance antibody and TCR sequence numbering tool for Rust, Python, Polars and JS/TS. - ENPICOM/immunum
151
Egor Marin @marinegor.bsky.social · 23/03/2026
me: big tech companies probably have automated everything big tech companies in 2026:
a screenshot of an sms from KPN. The message says:

Beste klant, je bent nu in Verenigd Koninkrijk. Binnen de EU bel en sms je zoals in Nederland....
010
Egor Marin @marinegor.bsky.social · 23/03/2026
your scripts got me through my masters and phd (and endless refinement cycles), thank you so much! truly think they should be a part of coot's distribution :)
010
Egor Marin @marinegor.bsky.social · 20/03/2026
with coding it feel less productive sometimes, but with infrastructure it's actually a life-saver for me personally. Like, configuring github actions / deploys / ... is actually so much better with it, mainly because I know exactly what I want to do but don't know how😁
030
Egor Marin @marinegor.bsky.social · 19/03/2026
I have a rust joke but it's still compiling
000
Egor Marin @marinegor.bsky.social · 18/03/2026
from cat's perspective, they're resting on the right side of the cat, no?
000
Egor Marin @marinegor.bsky.social · 28/02/2026
I constantly wonder how much the crystallographic data quality actually matters -- not overall, like resolution or rfactors, but local, like per-residue modelling scores. And if cleaning the dataset better will result in better model🤔
010
Egor Marin @marinegor.bsky.social · 26/02/2026
from my discussions with PDB maintainers, it's a legacy thing. The "auth" things are there to carefully preserve information that authors put some meaning in chain names (eg H and L for antibody chains, L for lipids, S for solvent etc)
050
Egor Marin @marinegor.bsky.social · 22/02/2026
I've seen them talk online and at a PEGS conference -- no mentions of preparing a publication there.
010
Egor Marin @marinegor.bsky.social · 21/02/2026
I wonder if this person has ever seen electron density maps that yielded pdb structures for the AI training🙃
000
Egor Marin @marinegor.bsky.social · 10/02/2026
...and this is how silly hoomans will help clever agents building new things✨
010
Egor Marin @marinegor.bsky.social · 06/02/2026
to me the switch was so easy since it's declarative. And also interactivity is just so easily done with altair, definitely a killer feature.
010
Egor Marin @marinegor.bsky.social · 06/02/2026
bonus point is that you can embed them with html onto your website easily!
110
Egor Marin @marinegor.bsky.social · 06/02/2026
not sure if it's something you're interested in, but I usually plot things like that with altair, and then make them interactive and with a tooltip, with 5-10-50 different sliding window options, just to see how it behaves instead of plotting it every time with matplotlib :)
110
Egor Marin @marinegor.bsky.social · 06/02/2026
would it be more informative to plot first derivatives perhaps?
110
Egor Marin @marinegor.bsky.social · 27/01/2026
you said "I'll leave the link in the show notes" on 7:53, but you never did💔 I assume you're talking about this link, right: docs.marimo.io/guides/wasm/
docs.marimo.io
WebAssembly notebooks - marimo
The next generation of Python notebooks
110
Egor Marin @marinegor.bsky.social · 21/01/2026
yes! and it has amazing apps too, can not recommend this stack enough: github.com/navilg/media...
github.com
GitHub - navilg/media-stack: A self-hosted stack for media management and streaming, with AI-powered movie and show recommendations. Includes Sonarr, Radarr, qBitTorrent, Prowlarr, Jellyfin, Jellyseer...
A self-hosted stack for media management and streaming, with AI-powered movie and show recommendations. Includes Sonarr, Radarr, qBitTorrent, Prowlarr, Jellyfin, Jellyseerr, Recommendarr, and VPN s...
020
Egor Marin @marinegor.bsky.social · 10/01/2026
it's probably fine as it is for archival purposes, but certainly not for consumption😁 and low visibility of libraries such as gemmi/mdanalysis/biotite lead to abundance of self-written PDB/cif parcers, which imo has a lot of drawbacks.
010
Egor Marin @marinegor.bsky.social · 10/01/2026
to be honest, I don't really care for space -- iirc, whole RCSB is under 200 Gb, and significantly less if you care about only cryoEM/crystallography structures under certain size. I honestly wouldn't change anything (except for probably GraphQL API), and perhaps work on tutorials and docs more.
100
Egor Marin @marinegor.bsky.social · 09/01/2026
😁 I'm just trying to understand whether your problem is format itself or the underlying data model
110
Egor Marin @marinegor.bsky.social · 09/01/2026
why though, may I ask?
110
Egor Marin @marinegor.bsky.social · 09/01/2026
cif and pdb are indeed the worst formats, the only problem is that all others are even worse :) also, I'd highly recommend using gemmi that allows you to parse cif into json. Although arguably, the data model itself is very messy, which imo is expected for half-aa-century old legacy.
110