Sign in

Jacob Schreiber

@jmschreiber91.bsky.social
6.7K followers 1.4K following 824 posts

Studying genomics, machine learning, and fruit. My code is like our genomes -- most of it is junk. Assistant Professor UMass Chan Previously IMP Vienna, Stanford Genetics, UW CSE.

PostsRepliesMedia
Reposted by Jacob Schreiber
Shenzhi Chen @shenzhichen1999.bsky.social · 25/08/2026
Excited to share our paper in @natgenet.nature.com on the de novo design of tissue-specific mammalian enhancers that function in vivo in mouse embryos. 🧬🐭 www.nature.com/articles/s41...
35826
Jacob Schreiber @jmschreiber91.bsky.social · 14/03/2026
There was concern it would not be helpful for identification purposes for the few people who do not already know me
010
Jacob Schreiber @jmschreiber91.bsky.social · 09/03/2026
The Programmable Genomics Lab at @umasschan.bsky.social just officially launched our site! We have a "simple" goal: to develop synthetic regulatory elements that target every cell type in every tissue, and control payload dosage/duration Check it out! programmable-genomics.github.io
programmable-genomics.github.io
Programmable Genomics Laboratory | Home
1218
Reposted by Jacob Schreiber
Alex Stark @alex-stark.bsky.social · 24/12/2025
Our preprint "Predictive design of tissue-specific mammalian enhancers that function in vivo in the mouse embryo" is on bioRxiv: www.biorxiv.org/content/10.6... . Amazing collaboration by @shenzhichen1999.bsky.social, Vincent Loubiere (@impvienna.bsky.social,@viennabiocenter.bsky.social),... (1/2)
biorxiv.org
Predictive design of tissue-specific mammalian enhancers that function in vivo in the mouse embryo
Enhancers control tissue-specific gene expression across metazoans. Although deep learning has enabled enhancer prediction and design in mammalian cell lines and invertebrate systems, it remains uncle...
210448
Jacob Schreiber @jmschreiber91.bsky.social · 10/12/2025
There's a ton of really cool stuff in the paper. Take a skim if you have time!
010
Jacob Schreiber @jmschreiber91.bsky.social · 10/12/2025
In collaboration with @alex-stark.bsky.social , we designed cell type-specific enhancers and experimentally verified them using STARR-seq. Strikingly, we found that Ledidi can even design enhancers with stronger regulatory activity than the strongest endogenous enhancers.
130
Jacob Schreiber @jmschreiber91.bsky.social · 10/12/2025
We designed regulatory DNA in dozens of settings and found that we could achieve our target objective with surprisingly few edits (biology alert). Because we focus on edits, our designs are less likely to overfit to the model or inadvertently alter properties that are not well captured by the model.
100
Jacob Schreiber @jmschreiber91.bsky.social · 10/12/2025
Ledidi is a DNA design method that designs *edits* to a template while explicitly minimizing the number of needed edits. This allows you to build upon informative starting material, which our genomes are full of, and simply edit in the final touches, similar to an Instagram filter or video touch-up.
101
Jacob Schreiber @jmschreiber91.bsky.social · 10/12/2025
After a huge amount of work w/ @alex-stark.bsky.social's group, a new version of our Ledidi preprint is now out! In an era of AI-designed proteins, the next leap will be controlling when, where, and how much of these proteins are expressed in living cells. www.biorxiv.org/content/10.1...
biorxiv.org
Programmatic design and editing of cis-regulatory elements
The development of modern genome editing and DNA synthesis has enabled researchers to edit DNA sequences with high precision but has left unsolved the problem of designing these edits. We introduce Le...
26026
Jacob Schreiber @jmschreiber91.bsky.social · 25/11/2025
The @impvienna.bsky.social is a unique place where talented researchers are doing amazing science. I'm glad I had an opportunity to spend some time there, and am looking forward to the next excuse to visit Vienna!
1120
Jacob Schreiber @jmschreiber91.bsky.social · 20/11/2025
That sounds too biologically important for me to be involved.
000
Jacob Schreiber @jmschreiber91.bsky.social · 20/11/2025
My first @umasschan.bsky.social/@impvienna.bsky.social affiliated paper is up! tomtom-lite is a re-implementation of tomtom targeting the ML age of genomics. Fast annotations ("what is this motif?") and simple large-scale discovery of motifs. Check it out! academic.oup.com/bioinformati...
academic.oup.com
Tomtom-lite: accelerating Tomtom enables large-scale and real-time motif similarity scoring
AbstractSummary. Pairwise sequence similarity is a core operation in genomic analysis, yet most attention has been given to sequences made up of discrete c
13514
Jacob Schreiber @jmschreiber91.bsky.social · 31/10/2025
Why is it a problem to translate C++ to changng standards when you can just use Agenic AI?
000
Reposted by Jacob Schreiber
Pinar Demetci @pinard.bsky.social · 29/10/2025
Tune in for a great MIA talk on ML for regulatory genomics by @jmschreiber91.bsky.social and Gregory Andrews now! 🧪 broad.io/mia (Will also be available online on our YouTube playlist later: www.youtube.com/watch?v=sMTO...)
052
Reposted by Jacob Schreiber
Avantika Lal @avantikalal.bsky.social · 15/10/2025
I'm happy to share that our gReLU package is now published in Nature Methods! www.nature.com/articles/s41...
nature.com
gReLU: a comprehensive framework for DNA sequence modeling and design - Nature Methods
gReLU advances deep-learning-based modeling and analysis of DNA sequences with comprehensive toolsets and versatile applications.
0206
Jacob Schreiber @jmschreiber91.bsky.social · 13/10/2025
they told the flight attendants to sit down for the second half of the flight (an hour) because it was entirely turbulence
200
Jacob Schreiber @jmschreiber91.bsky.social · 13/10/2025
had to land in a nor'eastern storm in boston and i'm ready to move back to the west coast forever
270
Reposted by Jacob Schreiber
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
Now that I'm settled in at @umasschan.bsky.social, I'm hiring at all levels: grad students, post-docs, and software engineers/bioinformaticians! The goal of my lab is to understand the regulatory role of every nucleotide in our genomes and how this changes across every cell in our bodies.
64326
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
I figure I can spend the time until I get tenure on answering the question, and then the time after tenure arguing about what "regulatory" means
170
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
If you're interested, please reach out with your CV and which topics you'd be interested in working on!
010
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
- Genomics Software Ecosystem: A major obstacle to our goal is the lack of simple+scalable software that everyone can use. Come build this with me. Training a lightweight deep learning model and using it for design/interpretability/VE prediction should be no more challenging than mapping reads.
160
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
- Foundation Models: As someone involved in ML, I am legally required to be working on this topic.
120
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
We have an array of ML-based projects for going after this, focusing on the following topics: - DNA Design ( 🧬 ) We have shown that Ledidi (www.biorxiv.org/content/10.1...) can precisely design DNA, and now it's time to push the boundaries in several directions w/ some very cool collaborations.
biorxiv.org
Programmatic design and editing of cis-regulatory elements
The development of modern genome editing tools has enabled researchers to make such edits with high precision but has left unsolved the problem of designing these edits. As a solution, we propose Ledi...
120
Jacob Schreiber @jmschreiber91.bsky.social · 07/10/2025
Now that I'm settled in at @umasschan.bsky.social, I'm hiring at all levels: grad students, post-docs, and software engineers/bioinformaticians! The goal of my lab is to understand the regulatory role of every nucleotide in our genomes and how this changes across every cell in our bodies.
64326
Jacob Schreiber @jmschreiber91.bsky.social · 03/10/2025
It was suggested that the audience may not appreciate/understand :(
200
Jacob Schreiber @jmschreiber91.bsky.social · 30/09/2025
is it a good idea to wear a "join, or die!" hat to a big talk in europe? please say yes
130
Jacob Schreiber @jmschreiber91.bsky.social · 22/09/2025
the greatest productivity hack is having a grant deadline. there's so much other stuff you can do when you're supposed to be working on a grant.
0150
Jacob Schreiber @jmschreiber91.bsky.social · 19/09/2025
I was delighted to have the unexpected opportunity to give a keynote at MLCB 2025 in NYC last week. I used it to explain how I view deep learning models in genomics not as "uninterpretable black boxes" but as indispensable tools for understanding genomics + designing the next gen of synthetic DNA.
0121
Jacob Schreiber @jmschreiber91.bsky.social · 09/09/2025
for some reason i thought being a professor would involve more mentoring and research and less filling out disclosures concerning whether plants and seeds were used in my computational study
0101
Jacob Schreiber @jmschreiber91.bsky.social · 03/09/2025
stocking up the new apartment with essentials
040
Reposted by Jacob Schreiber
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
In the genomics community, we have focused pretty heavily on achieving state-of-the-art predictive performance. While undoubtedly important, how we *use* these models after training is potentially even more important. tangermeme v1.0.0 is out now. Hope you find it useful!
14414
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
For some reason, hitting "comment" on GitHub is significantly more responsive than a month ago and it freaks me out. Surely there are some important calculations that need to be done before letting my thoughts into the wild?
010
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Thanks! Let me know if you want me to stop in virtually, we can try to figure out a time.
010
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Hope you find tangermeme helpful in your work! Please reach out if you have any comments + questions.
000
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Because everything is automatic, we can probe models. What motifs are driving model predictions? Calculate attributions, call + annotate seqlets, and count the annotations! BPNet is relying on MYC, whereas Beluga is relying on many more TFs. Easy comparison now.
110
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Frequently, people manually annotate seqlets and draw bars or boxes around these high-attribution characters themselves. This is not really a problem, but it's just slow and does not scale genome-wide. In the above picture, everything is automatically done.
111
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
People *talk* about seqlets a lot but tangermeme is the first package for complete functionality. Here is a complete example of using tangermeme for attributions, seqlet calling + annotation, and plotting, to visualize what five models think of the same locus
130
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Expanding past these implementations, tangermeme has a large focus on automatic seqlet calling and usage. Seqlets are short contiguous spans of high-attribution characters that usually correspond to the binding of a TF.
100
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
By considering attributions you can see how variants disrupt or change usage of motifs. Maybe you'll even find that a variant causes alternative binding by inducing a new motif or slightly changing competition! That would be challenging to see from the predictions alone.
110
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Past simply re-implementing algorithms people use (in a convenient repo), tangermeme offers flexibility not usually offers in other implementations. As an example, instead of calculating variant effect as predictions before/after a substitution, why not look at attributions?
100
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
This care extends to each of our operations. For example, one-hot encoding the entirety of chr1 takes <2s on a single thread. This is significantly faster than other one-hot encoding methods out there, and is fast enough to enable real-time batch generation from FASTAs.
130
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Here is a (twitter) thread on the issue: x.com/jmschreiber9...
x.com
Jacob Schreiber on X: "Using Captum to interpret your @PyTorch models using DeepLift/DeepLiftShap? If you specify your activations incorrectly, you will silently get incorrect attributions. In this genomics example, the TTTGCAT.ACAAT motif is the important thing and is entirely missed. https://t.co/MuVsAO5isz" / X
Using Captum to interpret your @PyTorch models using DeepLift/DeepLiftShap? If you specify your activations incorrectly, you will silently get incorrect attributions. In this genomics example, the TTTGCAT.ACAAT motif is the important thing and is entirely missed. https://t.co/MuVsAO5isz
100
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
By focusing in this manner, we can "delve" deeply into these downstream algorithms. For instance, we found a bug in many DeepLIFT/SHAP implementations that will cause them to silently fail when you don't register your operations. Didn't know you needed to do that? Same!
120
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
This design choice is intentional. Because model components and training strategies are hugely variable and evolving quickly, I did not even want to touch those aspects. You define and train your model however you want, and then use tangermeme to do genomic discovery with it.
220
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
tangermeme is a toolkit that implements "everything-but-the-model" for genomic machine learning. This includes sequence manipulations, batched predictions, attributions, ablations, marginalizations, variant effect prediction, design, etc...
120
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
Preprint: biorxiv.org/content/10.1... Installation: `pip install tangermeme`
biorxiv.org
tangermeme: A toolkit for understanding cis-regulatory logic using deep learning models
Deep learning models have achieved state-of-the-art performance at predicting diverse genomic modalities, yet their promise for biological discovery lies in how they are used after demonstrating their...
131
Jacob Schreiber @jmschreiber91.bsky.social · 27/08/2025
In the genomics community, we have focused pretty heavily on achieving state-of-the-art predictive performance. While undoubtedly important, how we *use* these models after training is potentially even more important. tangermeme v1.0.0 is out now. Hope you find it useful!
14414
Jacob Schreiber @jmschreiber91.bsky.social · 26/08/2025
It's about both. Conceptually, they are intertwined.
010
Jacob Schreiber @jmschreiber91.bsky.social · 26/08/2025
An excellent post about the receptive range of convolution models. "You might reasonably ask: "If I have 100 layers with W=1000W=1000, that's a theoretical receptive field of 100,000 tokens. Doesn't that matter?" The answer is no, and here's why:" guangxuanx.com/blog/stackin...
guangxuanx.com
Why Stacking Sliding Windows Can't See Very Far
Modern LLMs use sliding window attention for efficiency, but why can't stacking sliding windows see as far as theory suggests? A mathematical exploration of information dilution and the exponential ba...
150
Jacob Schreiber @jmschreiber91.bsky.social · 25/08/2025
Multiple, in fact
000