Sign in

Omer Moussa

@omermosa.bsky.social
23 followers 99 following 14 posts

Neurosci x AI PhD Student @the Max Planck Institute for Software Systems, supervised by @mtoneva.bsky.social; CS@MaxPlanck -- ML and CogNeurosci Enthusiast.

PostsRepliesMedia
Omer Moussa @omermosa.bsky.social · 01/10/2026
13/ We're making everything public. 📄 Paper: arxiv.org/abs/2607.05171 🌐 Interactive website + demos: bridge-ai-neuro.github.io/rabbit/ 💻 Code: github.com/bridge-ai-ne... We would love to hear from you after trying it! Huge thanks to @mtoneva.bsky.social for making this work possible.
arxiv.org
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This ...
040
Omer Moussa @omermosa.bsky.social · 01/10/2026
12/ Third, Shared–Idiosyncratic Decomposition (SID) captures what responses share and how individuals differ. It gives us a population starting point, then lets us adapt only the small idiosyncratic heads. During few-shot adaptation, these idiosyncratic heads are the only part tuned in the model.
110
Omer Moussa @omermosa.bsky.social · 01/10/2026
11/ Second, our Temporal Brain Transformer lets each of the cortical regions learn what to attend to in the speech output. These regional representations feed a readout predicting ~41K cortical surface vertices. Speech goes in; a detailed brain prediction comes out
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
10/ We brought together three ideas to make this work. First, brain-tuning: fine-tune a pretrained speech model directly on paired audio and fMRI from CNeuroMod Friends. Brain data helps shape the speech representations from which we predict cortical responses.
110
Omer Moussa @omermosa.bsky.social · 01/10/2026
9/ Second, each brain region in RABBiT learns its own representation. If we start in the primary auditory cortex and move to the most similar region, the model reconstructs the cortical progression: primary auditory → belt → STG/STS → temporal → frontal language, with no anatomical supervision.
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
8/ Beyond accuracy, we also asked: does RABBiT organize speech and language information in a brain-like way? The answer is “sounds like it does”. First, it reproduces the classic left-lateralized language network, without being trained on any non-naturalistic experiment.
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
7/ The few-shot gains are concentrated in higher-order language regions, regions that we found to be most idiosyncratic: IFG, angular gyrus, supramarginal gyrus, mPFC, MFG/DLPFC. These are regions that no strong zero-shot predictor will predict well because they are not shared across the population.
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
6/ That tiny update beats voxel-wise ridge regression while updating roughly 1000× fewer parameters. Few-shot brain prediction becomes lightweight, fast, personalized, and realistic for settings where collecting hours of fMRI per person simply isn't an option.
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
5/ The few-shot result is where things get exciting for us. With only 5-10 minutes of fMRI, RABBiT personalizes its predictions to totally new subjects and stimuli. Not by retraining the whole model, or fitting a huge voxel-wise model, but by updating a compact pathway (only ~115K parameters).
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
4/ The zero-shot result is the first big gain. Across 324 unseen participants from 15 held-out studies, RABBiT reaches the inter-subject consistency estimate. RABBiT also outperforms the current state-of-the-art, TRIBEv2, across auditory and language regions (despite RABBiT being much smaller).
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
3/ That gives one model two modes natively. 🟦 Zero-shot: no fMRI. Just speech in, and RABBiT predicts the reliable group-level brain response shared across people. 🟥 Few-shot: give it only ~10 minutes of the new person's fMRI, and RABBiT learns a tiny personalized correction.
101
Omer Moussa @omermosa.bsky.social · 01/10/2026
2/ The key problem is that not all brain regions’ responses are shared across people. Early auditory cortex is highly consistent, while higher-order language regions are more idiosyncratic. So RABBiT doesn't force one solution everywhere. It separates what's shared from what's idiosyncratic.
100
Omer Moussa @omermosa.bsky.social · 01/10/2026
1/ Foundation models changed AI because they learn strong priors and can adapt quickly to different tasks. Brain encoding models never had this; existing models either produce shared predictions or need to be fit fully per participant. We wanted the equivalent of a this for language in the brain.
110
Omer Moussa @omermosa.bsky.social · 01/10/2026
🚨 Very excited to share our latest work: RABBiT: a foundation model for speech- and language-evoked brain activity. 🧠🎧 W/ @mtoneva.bsky.social One model enables: ✅ accurate zero-shot prediction ✅ efficient few-shot adaptation with only 5-10 mins of data ✅ convenient in-silico neuroscience. 🧵👇
1112