Sign in

Arna Ghosh

@arnaghosh.bsky.social
324 followers 229 following 59 posts

Research Scientist at Google Research, working on Bio-inspired AI • PhD at Mila & McGill University, Vanier scholar • Ex-RealityLabs, Meta AI • Comedy+Cricket enthusiast

PostsRepliesMedia
Reposted by Arna Ghosh
Carsen Stringer @computingnature.bsky.social · 22/05/2026
🚨🧠 Our paper is out! We introduce a simple computational model that generates macroscopic, long-timescale dynamics as seen in large-scale neural recordings. #neuroscience #dynamics @marius10p.bsky.social @zhong-lin.bsky.social @hhmijanelia.bsky.social Link: go.nature.com/4tRNjIu
414053
Reposted by Arna Ghosh
Blake Richards @tyrellturing.bsky.social · 04/05/2026
Here's what I think everyone is missing about this issue: Prediction is how we identify understanding! If we had a model that could do OOD prediction, then we would rightly say it understands. Of course, that doesn't mean understanding *necessarily* emerges from training to predict. #NeuroAI
1123
Reposted by Arna Ghosh
(((Dr Hannah Wirtshafter))) 🔬 @aheadofthenerve.bsky.social · 19/02/2026
New preprint out 🎉 What happens to the hippocampal “place code” when an animal is actively engaged in a task? The answer surprised us (and might surprise you too!). Let's dive in ⬇️ Link: "Hippocampal trace coding dominates and disrupts place coding" www.biorxiv.org/content/10.6...
biorxiv.org
36122
Reposted by Arna Ghosh
Jesper Sjöström @pjsjostrom.bsky.social · 16/02/2026
🧪🧠 New preprint: helping resolve a decades-long debate in synaptic plasticity NMDA receptors are central to Hebbian learning. Yet for >30 years, the existence and function of presynaptic NMDA receptors have remained controversial. 📄 doi.org/10.64898/202... 1/6
doi.org
15519
Arna Ghosh @arnaghosh.bsky.social · 12/02/2026
Someone added a `.stop_gradient()` and left it running. 😜
070
Reposted by Arna Ghosh
deanpospisil.bsky.social @deanpospisil.bsky.social · 27/01/2026
New paper out at PNAS: www.pnas.org/doi/10.1073/... Revisiting the high-dimensional geometry of population responses in the visual cortex with @jpillowtime.bsky.social. The review took forever because a reviewer was doubtful our new estimator can infer eigenvalues beyond the rank of the data! (1/6)
pnas.org
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
26819
Reposted by Arna Ghosh
Blake Richards @tyrellturing.bsky.social · 13/01/2026
Are you thinking about doing neuroscience outreach but want to make it more exciting or hands on? Check out RetINaBox! (A collab led by the Trenholm lab) We tried to bring the experience of experimental neuroscience to a classroom setting: www.eneuro.org/content/13/1... #neuroscience 🧪
eneuro.org
RetINaBox: A Hands-On Learning Tool for Experimental Neuroscience
An exciting aspect of neuroscience is developing and testing hypotheses via experimentation. However, due to logistical and financial hurdles, the experiment and discovery component of neuroscience is...
03914
Arna Ghosh @arnaghosh.bsky.social · 05/12/2025
Whoaaa!! This is a fantastic effort, and an amazing resource. Huge congratulations to the authors! 🎉
010
Reposted by Arna Ghosh
Mila - Institut québécois d'IA @mila-quebec.bsky.social · 05/12/2025
Last day of poster sessions and presentations at @neuripsconf.bsky.social. Full schedule featuring Mila-affiliated researchers presenting their work at #NeurIPS2025 here mila.quebec/en/news/foll...
011
Arna Ghosh @arnaghosh.bsky.social · 05/12/2025
In San Diego attending #NeurIPS2025? Come to our poster to talk more about representation geometry in LLMs. 😃 🗓️ Friday 4:30-7:30 pm session 📍 Exhibit Hall C, D, E 🏁 Poster # 2502
041
Reposted by Arna Ghosh
Portugues Lab @portugueslab.bsky.social · 24/11/2025
(1/n) We are excited to share our new paper in Nature Communications, by Hagar Lavian (@hlavian.bsky.social) and team, revealing how the zebrafish brain integrates visual navigation signals! www.nature.com/articles/s41...
nature.com
Visual motion and landmark position align with heading direction in the zebrafish interpeduncular nucleus - Nature Communications
How are various visual signals integrated in the vertebrate brain for navigation? Here authors show that different spatial signals are topographically organized and align to one another in the zebrafi...
35421
Reposted by Arna Ghosh
Blake Richards @tyrellturing.bsky.social · 03/12/2025
1/ Why does RL struggle with social dilemmas? How can we ensure that AI learns to cooperate rather than compete? Introducing our new framework: MUPI (Embedded Universal Predictive Intelligence) which provides a theoretical basis for new cooperative solutions in RL. Preprint🧵👇 (Paper link below.)
Image of robots struggling with a social dilemma.
56728
Arna Ghosh @arnaghosh.bsky.social · 02/12/2025
Population coding 🙌
030
Reposted by Arna Ghosh
Konrad Kording @kordinglab.bsky.social · 02/12/2025
How I contributed to rejecting one of my favorite papers of all times, Yes, I teach it to students daily, and refer to it in lots of papers. Sorry. open.substack.com/pub/kording/...
open.substack.com
How I contributed to rejecting one of my favorite papers of all time
I believe we should talk about the mistakes we make.
111928
Arna Ghosh @arnaghosh.bsky.social · 18/11/2025
Thanks Ken! ☺️ Here's the (more updated) NeurIPS version: proceedings.neurips.cc/paper_files/... Also, more recently we extended the use of powerlaws for characterizing how representations change over (pre/post) training in LLMs. 🙂 🧵 here: bsky.app/profile/arna...
proceedings.neurips.cc
$\alpha$-ReQ : Assessing Representation Quality in Self-Supervised Learning by measuring eigenspectrum decay
040
Arna Ghosh @arnaghosh.bsky.social · 16/11/2025
This is an excellent blueprint on a very fascinating use of AI scientist! And the results and super cool and interesting! 🤩 I have been asked this when talking about our work on using powerlaws to study representation quality in deep neural networks, glad to have a more concrete answer now! 😃
1143
Reposted by Arna Ghosh
Greg Priest @gregpriest.bsky.social · 08/11/2025
Conrad Hal Waddington was born OTD in 1905. His “epigenetic landscape” is a diagrammatic representation of the constraints influencing embryonic development. On his 50th birthday, his colleagues gave him a pinball machine on the model of the epigenetic landscape. 🧪 🦫🦋 🌱🐋 #HistSTM #philsci #evobio
512335
Arna Ghosh @arnaghosh.bsky.social · 08/11/2025
You mean the algorithms "generate" some auxilliary targets and then do supervised learning?
000
Arna Ghosh @arnaghosh.bsky.social · 08/11/2025
I got you 😉
120
Reposted by Arna Ghosh
Shahab Bakhtiari @shahabbakht.bsky.social · 07/11/2025
I’m looking for interns to join our lab for a project on foundation models in neuroscience. Funded by @ivado.bsky.social and in collaboration with the IVADO regroupement 1 (AI and Neuroscience: ivado.ca/en/regroupem...). Interested? See the details in the comments. (1/3) 🧠🤖
ivado.ca
AI and Neuroscience | IVADO
14624
Reposted by Arna Ghosh
Olivier Codol @oliviercodol.bsky.social · 06/11/2025
A tad late (announcements coming) but very happy to share the latest developments in my previous preprint! Previously, we show that neural representations for control of movement are largely distinct following supervised or reinforcement learning. The latter most closely matches NHP recordings.
1468
Arna Ghosh @arnaghosh.bsky.social · 03/11/2025
Thank you! 😁
000
Arna Ghosh @arnaghosh.bsky.social · 03/11/2025
Indeed! We show in the paper that the DPO objective is analogous to contrastive learning objectives used for self-supervised vision pretraining, which is indeed entropy-seeking in nature (shown in prev works). I feel spectral metrics can go a long way in unlocking LLM understanding+design. 🚀
150
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
A big shoutout to @koustuvsinha.com for insightful discussions that shaped this work, and @natolambert.bsky.social + the OLMo team! Paper 📝: arxiv.org/abs/2509.23024 👩‍💻 Code : Coming soon! 👨‍💻
arxiv.org
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learned representations a...
060
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
This work was done with dream team 🤩 @melodylizx.bsky.social @kumarkagrawal.bsky.social Komal Teru @glajoie.bsky.social @adamsantoro.bsky.social @tyrellturing.bsky.social at @mila-quebec.bsky.social @berkeleyair.bsky.social @cohere.com & @googleresearch.bsky.social! 🧵9/9
140
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
Takeaway: LLM training exhibits multi-phasic information geometry changes! ✨ - Pretraining: Compress → Expand (Memorize) → Compress (Generalize). - Post-training: SFT/DPO → Expand; RLVR → Consolidate. Representation geometry offers insights into when models memorize vs. generalize! 🤓 🧵8/9
The multi-phasic information geometry changes in LLM pretraining and post-training. Pretraining undergoes an initial warmup phase, which corresponds to echolalia behavior, followed by entropy-seeking where the model learns high-frequency n-gram statistics, and finally a compression-seeking phase, where the model learns long-range dependencies. The post-training stages of SFT and DPO exhibit entropy-seeking behavior, where the model memorizes instruction following behavior, whereas RLVR exhibits compress-seeking behavior, where the model learns generalized reasoning at the cost of exploration.
130
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
BONUS: Is task-relevant info contained in the top eigendirections? On SciQ: - Removing top 10/50 directions barely hurts accuracy.✅ - Retaining only top 10/50 directions CRUSHES accuracy.📉 As supported by our theoretical results, eigenspectrum tail encodes critical task information! 🤯 🧵7/9
120
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
Why do these geometric phases arise?🤔 We show, both through theory and with simulations in a toy model, that these non-monotonic spectral changes occur due to gradient descent dynamics with cross-entropy loss under 2 conditions: 1. skewed token frequencies 2. representation bottlenecks 🧵6/9
140
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
Post-training also yields distinct geometric signatures: - SFT & DPO exhibit entropy-seeking expansion, favoring instruction memorization but reducing OOD robustness.📈 - RLVR exhibits compression-seeking consolidation, learning reward-aligned behaviors at the cost of reduced exploration.📉 🧵5/9
Supervised Finetuning (SFT) exhibits entropy-seeking expansion, coupled with decreased OOD robustness, whereas Reinforcement Learning from Verifiable Rewards (RLVR) exhibits compression-seeking consolidation, coupled with reduced exploration.
120
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
How do these phases relate to LLM behavior? - Entropy-seeking: Correlates with short-sequence memorization (♾️-gram alignment). - Compression-seeking: Correlates with dramatic gains in long-context factual reasoning, e.g. TriviaQA. Curious about ♾️-grams? See: bsky.app/profile/liuj... 🧵4/9
Different geometric phases correspond to acquiring different behaviors: entropy-seeking phase correlates with increased short-sequence memorization, compression-seeking phase correlates with better TriviaQA performance.
150
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
LLMs have 3 pretraining phases: Warmup: Rapid compression, collapsing representation to dominant directions. Entropy-seeking: Manifold expansion, adding info in non-dominant directions.📈 Compression-seeking: Anisotropic consolidation, selectively packing more info in dominant directions.📉 🧵3/9
OLMo-2 and Pythia models undergo multiple distinct geometric phases during pretraining, indicating a non-monotonic change in representation complexity underlying monotonic decrease in training loss.
160
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
When investigating OLMo (@ai2.bsky.social) & Pythia (@eleutherai.bsky.social) model checkpoints, as expected, pretraining loss ⬇️monotonically. BUT 🎢The spectral metrics (RankMe, αReQ) change non-monotonically (with more pretraining)! Takeaway: We discover geometric phases of LLM learning! 🧵2/9
150
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
📐We measured representation complexity using the #eigenspectrum of the final layer representations. We used 2 spectral metrics: - Spectral Decay Rate, αReQ: Fraction of variance in non-dominant directions. - RankMe: Effective Rank; #dims truly active. ⬇️αReQ ⇒ ⬆️RankMe ⇒ More complex! 🧵1/9
Spectral decomposition methods and metrics used to quantify representation space complexity in LLMs.
170
Arna Ghosh @arnaghosh.bsky.social · 31/10/2025
LLMs are trained to compress data by mapping sequences to high-dim representations! How does the complexity of this mapping change across LLM training? How does it relate to the model’s capabilities? 🤔 Announcing our #NeurIPS2025 📄 that dives into this. 🧵below #AIResearch #MachineLearning #LLM
New paper titled "Tracing the Representation Geometry of Language Models from Pretraining to Post-training" by Melody Z Li, Kumar K Agrawal, Arna Ghosh, Komal K Teru, Adam Santoro, Guillaume Lajoie, Blake A Richards.
16112
Arna Ghosh @arnaghosh.bsky.social · 19/09/2025
Very cool study, with interesting insights about theta sequences and learning!
110
Reposted by Arna Ghosh
Blake Richards @tyrellturing.bsky.social · 09/07/2025
Together with @repromancer.bsky.social, I have been musing for a while that the exponentiated gradient algorithm we've advocated for comp neuro would work well with low-precision ANNs. This group got it working! arxiv.org/abs/2506.17768 May be a great way to reduce AI energy use!!! #MLSky 🧪
arxiv.org
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
Studies in neuroscience have shown that biological synapses follow a log-normal distribution whose transitioning can be explained by noisy multiplicative dynamics. Biological networks can function sta...
33913
Arna Ghosh @arnaghosh.bsky.social · 24/06/2025
Congratulations, Dan!! 😁
010
Arna Ghosh @arnaghosh.bsky.social · 16/06/2025
This looks like a very cool result! 😀 Can't wait to read in detail.
030
Arna Ghosh @arnaghosh.bsky.social · 09/06/2025
Fantastic work on Multi-agent RL from @dvnxmvlhdf5.bsky.social & @tyrellturing.bsky.social! 🤩
082
Reposted by Arna Ghosh
Yiğit Demirağ @yigit.ai · 08/05/2025
Our team is hiring in Zurich! www.google.com/about/career...
google.com
Research Scientist, Paradigms of Intelligence — Google Careers
0153
Arna Ghosh @arnaghosh.bsky.social · 02/04/2025
Re diff implicit biases of architecture: the metrics implemented here (roughly) characterize the eigenspectrum (eigenval distribution) of the representation space. They don't really incorporate the eigenvector information --> hence, "what" features don't matter, only "how" matters.
000
Arna Ghosh @arnaghosh.bsky.social · 02/04/2025
Comparing models of different architectures often is tricky because of the different implicit biases of each architecture. RSA can be helpful to in some cases. But if you are looking for a metric (number) that tells you which model is better, I have some ideas but they are not implemented here. 😜
100
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
Indeed! The metrics work best when comparing networks of comparable architectures, though. So, if you are looking to select the best model checkpoint from a pretraining routine or diff hyperparam configurations, and the loss function is not as insightful, these metrics are incredibly helpful. :)
110
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
Also, big shoutout to @quentin-garrido.bsky.social+gang and @aggieinca.bsky.social+gang for developing Rankme and Lidar, respectively. Reptrix incorporates these representation quality metrics. 🚀 Let's make it easier to select good SSL/foundation models. 💪
020
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
PS: It was fun to put Reptrix together with @daniebenes.bsky.social, our open-source expert and new addition to the α-Squad comprising of Arnab Mondal, @kumarkagrawal.bsky.social @tyrellturing.bsky.social and I. [6/6]
030
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
If you do end up giving Reptrix a try, we would love to hear from you! We also encourage you to contribute to our repo and add example notebooks trying out these metrics for your networks, beyond natural language and vision domains. github.com/BARL-SSL/rep... [5/6]
github.com
reptrix/CONTRIBUTING.md at main · BARL-SSL/reptrix
Library that provides metrics to assess representation quality - BARL-SSL/reptrix
110
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
Metrics for evaluating representation quality: - α-ReQ: Measures discriminativeness. Lower alpha = better! - RankMe: Assesses representation capacity. Higher rank = higher capacity! - LiDAR: Evaluates separability among object manifolds. Higher rank = better separability! [4/6]
120
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
✨ Key Features of Reptrix: - 📈 Suite of metrics to assess representation quality: α-ReQ, RankMe, LiDAR, and more! - 🤝 Seamless PyTorch integration for minimal setup, maximum insights - 💻 Open Source: Contribute and enhance! [3/6]
110
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
Inspired by conversations after our α-ReQ paper (NeurIPS 2022) and subsequent work, we created Reptrix as an open-source library for assessing representation quality across models of vision, language… and more. Check out our @mila-quebec.bsky.social blogpost: mila.quebec/en/article/a... [2/6]
mila.quebec
α-ReQ: Assessing Representation Quality in SSL | Mila
The success of self-supervised learning algorithms has drastically changed the landscape of training deep neural networks. With well-engineered architectures and training objectives, SSL models learn ...
120
Arna Ghosh @arnaghosh.bsky.social · 01/04/2025
Are you training self-supervised/foundation models, and worried if they are learning good representations? We got you covered! 💪 🦖Introducing Reptrix, a #Python library to evaluate representation quality metrics for neural nets: github.com/BARL-SSL/rep... 🧵👇[1/6] #DeepLearning
3279