Sign in

Jeremie Kalfon 👨‍💻🧬🤖🚀

@jkobject.com
731 followers 3.3K following 171 posts

Doing a Ph.D. AI in Bio. | Ex @WhiteLabGx @BroadInstitute @MIT | Built @PiPleteam | ML, Cancer, Genomics, Data Sci, Entrepreneur, FullStack Dev | All views are mine

PostsRepliesMedia
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 22/05/2026
This morning I had the chance to present my research at the Cancéropôle IDF "AI & Cancer" day. Talking single-cell foundation models to cancer researchers and clinicians was a great exercise! 😅 1/2
111
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 20/04/2026
🚨 2026 Lilly x Nucleate Grand Challenge: Aging Reimagined 🔬 $100K non-dilutive 🏛️ Pitch at Lilly HQ 🤝 Lilly's science + venture teams Focus: mobility, cognition, immune resilience, regenerative medicine — the frontiers of healthspan. 📅 May 15 👉 linktr.ee/lillygrand...
000
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 13/04/2026
Instead of attending to every pair of positions, you make two lightweight passes — one along rows, one along columns. I Just wrote up a small blogpost about it: 📖 jkobject.com/criss-c... Would love to hear from anyone exploring efficient attention mechanisms 🙂 #Transformers #Attention 2/2
000
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 05/12/2025
It was a blast hosting our Nucleate Inside AI roundtable at the France Techbio 2025 event.
110
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 25/09/2025
The first 1 million prime numbers vizualized in 2D according to their prime factors (Umap) Source: johnhw.github.io/uma...
031
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 20/09/2025
what they sell you, what you get...
010
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 20/06/2025
By using common fine-tuning mechanism we show how one can train from one scale to the next by back-propagating signal to the compressed tokens and lower scale model.
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 20/06/2025
Each group of biologist are on their own niche and so too are the models. But These models talk about different steps of the same stair. We present ideas on how we might end up training models from atoms to organs by using transformers to compress 🔺 🔻 data into tokens used by larger scale models
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 20/06/2025
Very happy to share that my new paper got accepted at the ICML workshop for Foundation model for Life Sciences!! www.biorxiv.org/cont... Foundation Models are being trained from atoms to molecules ⚛️, molecule chains 🧬, entire cells 🦠, and even groups of cell across tissue slices 🫁
111
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 23/05/2025
As part of my research, I believe that scientific outreach is essential. Last week, I had the pleasure of presenting how AI can help us understand biology and the cell at Pint of Science 2025! I also put together a short video recap (in French) for those curious: youtu.be/fc8L8Dn_7tw... 1/2
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 23/05/2025
As part of my research, I believe that scientific outreach is essential. Last week, I had the pleasure of presenting how AI can help us understand biology and the cell at Pint of Science 2025! I also put together a short video recap (in French) for those curious: youtu.be/fc8L8Dn_7tw... 1/2
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 09/12/2024
Thanks again team and congrats on the First Place!! 6/6
120
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 09/12/2024
Nothing is publication worthy of course and many important problems were not solved like batch effects, adaptive patch size etc. But still seeing what can be done with drive and elbow grease makes me optimistic about the future! 5/6
110
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 09/12/2024
Secondly we have ran scPRINT on gene panel ST datasets like Xenium to predict cell level cell type, disease, and impute the remaining gene's expression. With surprising ability to predict some cancerous cells in BRCA slides. Many other ideas were unfinished. 4/6
110
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 09/12/2024
Finding similar patients through their slides or slide subsets, To retrieve associated molecular profiles, disease subtypes, and treatment and health journey. Our tool automatically downloads the HEST1K database and generates anndata with patch embeddings. 3/6
110
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 09/12/2024
We manage to introduce a pipeline mixing spatialdata, #scPRINT #HEST1K #CONCH #SPATIALDATA and #NAPARI. Our first POC was STsimilarity: finding similar image patches across a large database of histopatologic slides. 2/6
110
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
020
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
I thus decided to formalize it by creating GRnnData: First it is a tool to import and store many different network format to an AnnData. 💁 But it also contains more bells and whistles to work with gene networks! (like subsetting some genes, extracting targets, plotting 💹. 5/6
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
Interestingly, there is a possible standard for it! 🎉AnnData contains the .varp field which is made to store var to var (e.g. genes to gene) relationships. However not many people use it… 4/6
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
However, working with gene networks, I have seen various ways to store them throughout the different papers and benchmarks. Often as some kind of tsv/csv/… file with some kind of a gene-gene list. This lack of standard makes it quite hard to work with gene networks🙉 3/6
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
Multi-modal AI in Biology will certainly be some kind of multi-scale approach where each modalitity is feeding into the next. Transformer use embeddings of element they look at. each can be produced by the previous scale transformer model. going from molecules to whole tissues!
000
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
💯 🙏 Also, I would like to acknowledge the important pioneering work from Geneformer, UCE, scFoundation and scGPT. Thanks to FlashAttention, pytorch, lightning, and scanpy for their toolkits. Thanks to Omnipath, Scenic+, Openproblems, Replogle et al. and Mc Calla et al.
000
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
We propose to use the specificity of scRNAseq data and define a multi task pre-training composed of expression denoising, bottleneck learning and classification. We also propose a new hierarchical classification method to work with the rich hierarchical ontologies used to label cells in cellxgene.
100
Jeremie Kalfon 👨‍💻🧬🤖🚀 @jkobject.com · 19/11/2024
-> it is for now a pre-print and more is to come but here are some of our results: scPRINT is a transformer model trained on 50M cells 🦠 from the cellxgene database, it has novel expression encoding and decoding schemes and new pre-training methodologies 🤖. www.biorxiv.org/content/10.1...
110