Sign in

Anej Svete

@anejsvete.bsky.social
265 followers 129 following 9 posts

PhD student in NLP at ETH Zurich. anejsvete.github.io

PostsRepliesMedia
Anej Svete @anejsvete.bsky.social · 21/04/2026
Also presenting at the poster session on Saturday at 3:15 PM :)
000
Anej Svete @anejsvete.bsky.social · 21/04/2026
Very excited to give an oral talk at ICLR this Saturday at 10:30 AM! We show that masked diffusion language models are computationally equivalent to looped transformers, and can solve anything chain of thought can; sometimes even faster, thanks to parallelism. Paper: arxiv.org/abs/2510.13117
110
Reposted by Anej Svete
Ai2 @ai2.bsky.social · 05/03/2026
Introducing Olmo Hybrid, a 7B fully open model combining transformer and linear RNN layers. It decisively outperforms Olmo 3 7B across evals, w/ new theory & scaling experiments explaining why. 🧵
1315
Reposted by Anej Svete
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 02/10/2025
Andy Yang, Christopher Watson, Anton Xue, Satwik Bhattamishra, Jose Llarena, William Merrill, Emile Dos Santos Ferreira, Anej Svete, David Chiang: The Transformer Cookbook arxiv.org/abs/2510.00368 arxiv.org/pdf/2510.00368 arxiv.org/html/2510.00368
002
Reposted by Anej Svete
pentagonalize.bsky.social @pentagonalize.bsky.social · 03/10/2025
We present The Transformer Cookbook: a collection of recipes for programming algorithms directly into transformers! Hungry for an induction head? Craving a Dyck language recognizer? We show you step-by-step how to cook up transformers for these algorithms and many more!
arxiv.org
The Transformer Cookbook
We present the transformer cookbook: a collection of techniques for directly encoding algorithms into a transformer's parameters. This work addresses the steep learning curve of such endeavors, a prob...
155
Reposted by Anej Svete
Ai2 @ai2.bsky.social · 26/08/2025
Introducing Asta—our bold initiative to accelerate science with trustworthy, capable agents, benchmarks, & developer resources that bring clarity to the landscape of scientific AI + agents. 🧵
3214
Reposted by Anej Svete
Ai2 @ai2.bsky.social · 26/08/2025
As part of Asta, our initiative to accelerate science with trustworthy AI agents, we built AstaBench—the first comprehensive benchmark to compare them. ⚖️
163
Reposted by Anej Svete
Ai2 @ai2.bsky.social · 28/07/2025
Ai2 is excited to be at #ACL2025 in Vienna, Austria this week. Come say hello, meet the team, and chat about the future of NLP. See you there! 🤝📚
093
Reposted by Anej Svete
Ai2 @ai2.bsky.social · 12/06/2025
We are #1 on the @huggingface heatmap - this is what true openness looks like!🥇🎉 750+ models 230+ datasets And counting... Come build with us huggingface.co/spaces/cfahl...
huggingface.co
Model Release Heatmap - a Hugging Face Space by cfahlgren1
Search this app to see model release activity for any Hugging Face organization or user over time. Just enter the org name to view their heatmap.
083
Anej Svete @anejsvete.bsky.social · 17/05/2025
6/ The work refines the landscape of transformer expressivity and demonstrates that seemingly minor implementation details can have major theoretical consequences for what neural architectures can represent.
000
Anej Svete @anejsvete.bsky.social · 17/05/2025
5/ This might help explain why positional encodings that skew attention toward recent (rightmost) tokens—like ALiBi—work so well in practice. They're compensating for an inherent limitation in conventional attention mechanisms.
100
Anej Svete @anejsvete.bsky.social · 17/05/2025
4/ Here's why this matters: leftmost-tiebreaking transformers are actually equivalent to soft-attention transformers in terms of expressivity! This suggests they might better approximate real-world transformers than right-attention models.
100
Anej Svete @anejsvete.bsky.social · 17/05/2025
3/ Specifically, we show that leftmost tiebreaking models correspond to a strictly weaker fragment of Linear Temporal Logic (LTL). While rightmost tiebreaking enables the full power of LTL, leftmost models are limited to the "past" fragment.
100
Anej Svete @anejsvete.bsky.social · 17/05/2025
2/ We analyzed future-masked unique hard attention transformers and found that those with leftmost tiebreaking are strictly less expressive than those with rightmost tiebreaking. The "Tale of Two Sides" nicely describes about how these two models differ.
100
Anej Svete @anejsvete.bsky.social · 17/05/2025
1/ When multiple positions achieve the maximum attention score in a transformer, we need a tiebreaking mechanism. Should we pick the leftmost or rightmost position? Turns out, this trivial implementation detail dramatically affects what transformers can express!
100
Anej Svete @anejsvete.bsky.social · 17/05/2025
🧵 Excited to share our paper "Unique Hard Attention: A Tale of Two Sides" with Selim, Jiaoda, and Ryan, where we show that the way transformers break ties in attention scores has profound implications on their expressivity! And it got accepted to ACL! :) The paper: arxiv.org/abs/2503.14615
arxiv.org
Unique Hard Attention: A Tale of Two Sides
Understanding the expressive power of transformers has recently attracted attention, as it offers insights into their abilities and limitations. Many studies analyze unique hard attention transformers...
121
Reposted by Anej Svete
Afra Amini @afraamn.bsky.social · 06/05/2025
Current KL estimation practices in RLHF can generate high variance and even negative values! We propose a provably better estimator that only takes a few lines of code to implement.🧵👇 w/ @xtimv.bsky.social and Ryan Cotterell code: arxiv.org/pdf/2504.10637 paper: github.com/rycolab/kl-rb
173
Reposted by Anej Svete
Javier Rando @javirandor.com · 09/12/2024
I will be at #NeurIPS2024 in Vancouver. I am excited to meet people working on AI Safety and Security. Drop a DM if you want to meet. I will be presenting two (spotlight!) works. Come say hi to our posters.
141
Reposted by Anej Svete
Marco @mcognetta.bsky.social · 26/11/2024
No joke, FLaNN is one of the most interesting servers around. Check out the website for talk information! flann.super.site
flann.super.site
FLaNN Seminars
We organize a series of weekly online seminars on Formal Language Theory, Natural Language Processing, Machine Learning and Computational Linguistics in an informal setting.
081
Reposted by Anej Svete
Shauli Ravfogel @shauli.bsky.social · 12/11/2024
Happy to share our work "Counterfactual Generation from Language Models" with @AnejSvete, @vesteinns, and Ryan Cotterell! We tackle generating true counterfactual strings from LMs after interventions and introduce a simple algorithm for it. (1/7) arxiv.org/pdf/2411.07180
2143