Sign in

Giovanni Monea

@giomonea.bsky.social
31 followers 25 following 7 posts

PhDing at Cornell University / Cornell Tech

PostsRepliesMedia
Giovanni Monea @giomonea.bsky.social · 16h
Very excited to finally share my IBM work. It was an intense but rewarding summer, and I had a blast working with my talented collaborators at IBM & Cornell ( @keshavramji.bsky.social @yelkurdi.bsky.social @lastrasl.bsky.social @yoavartzi.com @nthngdy.bsky.social @ramon-astudillo.bsky.social )
091
Giovanni Monea @giomonea.bsky.social · 16h
So, what does memory sharing bring to the table to *improve* quality? Prediction recursions attend heavily to the shared memory. Each read sends a gradient signal back to the state recursion. We find that this "gradient highway" makes the state recursion learn much more.
180
Giovanni Monea @giomonea.bsky.social · 16h
We train LPT and Hybrid LPT at 150M-1B parameters, on 20-50B tokens at 2, 3, and 5 recursions. The more recursions, the more our memory-sharing models outperform their respective baselines.
170
Giovanni Monea @giomonea.bsky.social · 16h
Prior work laid the foundation for looped models and reduced their extra memory at inference. The result? Surprisingly strong despite the train-test mismatch. That made us ask: could this be a good inductive bias? One recursion memorizes, the others reuse it to predict.
170
Giovanni Monea @giomonea.bsky.social · 16h
What is memory sharing? Our simple recipe: at later recursions, just concatenate the attention keys of the first recursion and those of the current recursion before performing standard attention. Then, restrict the later recursions to a short window for memory savings.
190
Giovanni Monea @giomonea.bsky.social · 16h
Sharing memory makes looping cheaper and better. At 1B with 5 recursions, H-LPT beats a same-size Transformer by 1.1 perplexity with 76% less memory. Locally, it's easier to fit on your GPU. For serving, it fits 3.6x bigger batches and keeps 81% of the Transformer's throughput.
180
Giovanni Monea @giomonea.bsky.social · 16h
⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *memory sharing* not only saves memory but is also a net-positive inductive bias! Less memory, same flops, higher quality. 🔗 arxiv.org/abs/2610.02383 🧵
16211
Reposted by Giovanni Monea
Conference on Language Modeling @colmweb.org · 20/03/2025
A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏
33631
Reposted by Giovanni Monea
Sophie Greenwood @sjgreenwood.bsky.social · 10/03/2025
Please repost to get the word out! @nkgarg.bsky.social and I are excited to present a personalized feed for academics! It shows posts about papers from accounts you’re following bsky.app/profile/pape...
8171118
Reposted by Giovanni Monea
Yoav Artzi @yoavartzi.com · 18/02/2025
We recently pushed an update to our in-context RL paper. Usually, updates don't justify a post, but this one is exceptionally contentful -> 🧵 tl;dr: all the findings are stronger, and the behaviors are super cool! arxiv.org/abs/2410.05362
1184