Sign in

Daniel Scalena

@danielsc4.it
436 followers 185 following 24 posts

Intern of TS @Cohere | PhDing @unimib 🇮🇹 & @GroNlp 🇳🇱, interpretability et similia danielsc4.it

PostsRepliesMedia
Daniel Scalena @danielsc4.it · 01/07/2026
I'd never have guessed models commit to their final answer this early, often within the first 20% of reasoning, across math/logic tasks and model families. The rest is mostly hedging that doesn't change their mind. And turns out they encode this internally, we can decode it! 🧵👇
093
Daniel Scalena @danielsc4.it · 05/01/2026
Indeed, but we also show the other side of the coin: personalized generation and its evaluation remain extremely challenging, and IMO professional human translators are still essential to produce a truly original, publication-ready final work as of today.
000
Daniel Scalena @danielsc4.it · 04/01/2026
Want models to translate in the style you actually like? Our paper is here! See you in Marocco! 🇲🇦
051
Daniel Scalena @danielsc4.it · 16/10/2025
Takeaway: EAGer shows we can be MORE efficient & MORE effective by letting models focus compute where it matters most. 📄Paper: arxiv.org/abs/2510.11170 💻Code: github.com/DanielSc4/EA... ✨Huge thanks to my mentors and collaborators @leozotos.bsky.social E. Fersini @malvinanissim.bsky.social A. Üstün
arxiv.org
EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computation is often required to generate multiple candidate sequenc...
020
Daniel Scalena @danielsc4.it · 16/10/2025
Results: Across 3B-20B models, EAGer cuts budget by up to 80%, boosts perf 13% w/o labels & 37% w/ labels on AIME. As M scales, EAGer consistently: 🚀 Achieves HIGHER Pass@k, ✂️ Uses FEWER tokens than baseline, 🕺 Shifts the Pareto frontier favorably across all tasks. 🧵5/
100
Daniel Scalena @danielsc4.it · 16/10/2025
The fun part: EAGer-adapt reallocates saved budget to "saturating" prompts hitting the M cap, no labels needed! – Training & Verification-Free 🚀 Full EAGer uses labels to catch failing prompts, lowering threshold to branch or add sequences. Great for verifiable pipelines! 🧵4/
100
Daniel Scalena @danielsc4.it · 16/10/2025
EAGer works by monitoring token entropy during generation. High entropy token → It branches to explore new paths (reusing prefixes). Token with low entropy → It continues a single path. We cap at M sequences/prompt, saving budget on easy ones without regen. Training-free! 🧵3/
110
Daniel Scalena @danielsc4.it · 16/10/2025
Why? Reasoning LLMs shine with CoTs, but full parallel sampling—generating multiple paths per prompt—is inefficient 😤. It wastes compute on redundant, predictable tokens, esp. for easy prompts. Hard prompts need more exploration but get the same budget. Enter EAGER🧠! 🧵2/
110
Daniel Scalena @danielsc4.it · 16/10/2025
You can easily save up to 65% of compute while improving performance on reasoning tasks 🤯 👀 Meet EAGer: We show that monitoring token-level uncertainty lets LLMs allocate compute dynamically - spending MORE on hard problems, LESS on easy ones. 🧵👇
111
Daniel Scalena @danielsc4.it · 20/08/2025
I’ll be attending the NEMI 2025 workshop this Friday and presenting a poster👇. Happy to chat about cool interpretability stuff there!
010
Daniel Scalena @danielsc4.it · 23/05/2025
📝 Paper: arxiv.org/abs/2505.16612 🔗 Code: github.com/DanielSc4/st... Thanks to my amazing co-authors: @gsarti.com , @arianna-bis.bsky.social , Elisabetta Fersini, @malvinanissim.bsky.social 7/7
arxiv.org
Steering Large Language Models for Machine Translation Personalization
High-quality machine translation systems based on large language models (LLMs) have simplified the production of personalized translations reflecting specific stylistic constraints. However, these sys...
030
Daniel Scalena @danielsc4.it · 23/05/2025
🔍 What’s happening in the model? We find that SAE steering and multi-shot prompting impact internal representations similarly, suggesting insight from user examples are summarized with extra interpretability potential (look at latents) and better efficiency (no long context) 6/
110
Daniel Scalena @danielsc4.it · 23/05/2025
🌍 Across 7 languages, our SAE-based method matches or outperforms traditional prompting methods! Our method obtains better human-like translations (H) personalization accuracy (P), and maintains translation quality (Comet ☄️ @nunonmg.bsky.social) especially for smaller LLMs. 5/
110
Daniel Scalena @danielsc4.it · 23/05/2025
💡 We compare prompting (zero and multi-shot + explanations) and inference-time interventions (ActAdd, REFT and SAEs). Following SpARE (@yuzhaouoe.bsky.social @alessiodevoto.bsky.social), we propose ✨ contrastive SAE steering ✨ with mutual info to personalize literary MT by tuning latent features 4/
142
Daniel Scalena @danielsc4.it · 23/05/2025
📈 But can models recognize and replicate individual translator styles?: ✓ Classifiers can find styles with high acc. (humans kinda don’t) ✓ Multi-shot prompting boosts style a lot ✓ We can detect strong style traces in activations (esp. mid layers) 3/
120
Daniel Scalena @danielsc4.it · 23/05/2025
📘 Literary translation isn't just about accuracy, but also creatively conveying meaning across languages. But LLMs prompted for MT are very literal. Prompting & steering to the rescue! Can we personalize LLM’s MT when few examples are available, without further tuning? 🔍 2/
120
Daniel Scalena @danielsc4.it · 23/05/2025
📢 New paper: Applied interpretability 🤝 MT personalization! We steer LLM generations to mimic human translator styles on literary novels in 7 languages. 📚 SAE steering can beat few-shot prompting, leading to better personalization while maintaining quality. 🧵1/
2205
Daniel Scalena @danielsc4.it · 04/12/2024
Hellooo 👀
110
Daniel Scalena @danielsc4.it · 28/11/2024
Hey hello! 👋
010
Daniel Scalena @danielsc4.it · 21/11/2024
Now on 🦋!
020
Daniel Scalena @danielsc4.it · 19/11/2024
Hello!
010
Daniel Scalena @danielsc4.it · 19/11/2024
👋
100
Daniel Scalena @danielsc4.it · 17/11/2024
It was great, I'm starting to get tickets for next year!
010
Daniel Scalena @danielsc4.it · 16/11/2024
👀🙋‍♂️
010