Sign in

Martina G. Vilas

@martinagvilas.bsky.social
2.3K followers 456 following 18 posts

AI Evaluation & Interpretability @NVIDIA

PostsRepliesMedia
Reposted by Martina G. Vilas
Besmira Nushi @besmiranushi.bsky.social · 22/10/2025
When to call it quits in LLM reasoning? 🛑 ‪Martina's internship project suggests trace monitoring metrics and classifiers that can detect when an LLM reasoning trace is going to fail in mid way. The approach saves up to 70% of token usage, and it even helps with increasing accuracy by 2%-3%.
031
Martina G. Vilas @martinagvilas.bsky.social · 22/10/2025
Can we predict which reasoning paths will succeed before seeing the answer? 🤔 Our new paper (arxiv.org/abs/2510.10494) proposes latent-trajectory signals from LLMs' hidden states to identify high-quality reasoning, cutting inference costs by up to 70% while maintaining accuracy
arxiv.org
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
Reasoning models improve their problem-solving ability through inference-time scaling, allocating more compute via longer token budgets. Identifying which reasoning traces are likely to succeed remain...
181
Reposted by Martina G. Vilas
Besmira Nushi @besmiranushi.bsky.social · 29/04/2025
All Eureka inference-time scaling insights are now available here: www.microsoft.com/en-us/resear... It was fun sharing these and more together with Vidhisha Balachandran @vidhishab.bsky.social and Vibhav Vineet at #ICLR2025.
microsoft.com
Eureka Inference-Time Scaling Insights: Where We Stand and What Lies Ahead - Microsoft Research
Understanding and measuring the potential of inference-time scaling for reasoning. The new Eureka study tests nine state-of-the-art models on eight diverse reasoning tasks.
032
Martina G. Vilas @martinagvilas.bsky.social · 18/04/2025
Looking forward to presenting this work next week at #ICLR2025! DM me if you are attending and want to grab a coffee to discuss these topics 💫
0214
Martina G. Vilas @martinagvilas.bsky.social · 02/12/2024
December 5th our ML theory group at Cohere For AI is hosting @mathildepapillon.bsky.social to discuss their recent review arxiv.org/abs/2407.09468 on geometric/topological/algebraic ML. Join us online 💫
0141
Reposted by Martina G. Vilas
Yu Lu Liu @liuyulu.bsky.social · 21/11/2024
I’m putting together a starter pack for researchers working on human-centered AI evaluation. Reply or DM me if you’d like to be added, or if you have suggestions! Thank you! (It looks NLP-centric at the moment, but that’s due to the current limits of my own knowledge 🙈) go.bsky.app/G3w9LpE
153510
Reposted by Martina G. Vilas
Julian Skirzynski @jskirzynski.bsky.social · 23/11/2024
I tried to find everyone who works in the area but I certainly missed some folks so please lmk... go.bsky.app/BYkRryU
325318
Reposted by Martina G. Vilas
Serge Belongie @serge.belongie.com · 22/11/2024
Does anyone know of any feeds (or similar) for student internship opportunities in ML/CV/NLP?
24511
Reposted by Martina G. Vilas
Adhiraj Ghosh @adhirajghosh.bsky.social · 19/11/2024
I've found starter packs on NLP, vision, graphics, etc. But personally, I would love to know and hear from researchers working on vision-language. So, let me know if you'd like to join this starter pack, would be happy to add! go.bsky.app/TENRRBb
425613
Reposted by Martina G. Vilas
Laura @lauraruis.bsky.social · 20/11/2024
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
36850139
Reposted by Martina G. Vilas
Sung Kim @sungkim.bsky.social · 18/11/2024
LLMs tend to match problem-solving strategies based on textual similarity rather than truly understanding the underlying principles of mathematical problems. Paper: Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
0477
Reposted by Martina G. Vilas
Federico Adolfi @fedeadolfi.bsky.social · 14/11/2024
A starter pack of people working on interpretability / explainability of all kinds, using theoretical and/or empirical approaches. Reply or DM if you want to be added, and help me reach others! go.bsky.app/DZv6TSS
348026
Reposted by Martina G. Vilas
Sweta Karlekar @swetakar.bsky.social · 19/11/2024
If you’re interested in mechanistic interpretability, I just found this starter pack and wanted to boost it (thanks for creating it @butanium.bsky.social !). Excited to have a mech interp community on bluesky 🎉 go.bsky.app/LisK3CP
3368
Reposted by Martina G. Vilas
Christian Wolf @chriswolfvision.bsky.social · 18/11/2024
I forgot from whom in my feed I got this from, but anyway, this network analyzer is crazy efficient. It gives you ideas for accounts to follow based on your own followees. I just added 50 accounts or so. bsky-follow-finder.theo.io
bsky-follow-finder.theo.io
Bluesky Network Analyzer
Find accounts that you don't follow (yet) but are followed by lots of accounts that you do follow.
98224
Reposted by Martina G. Vilas
Yoav Goldberg @yoavgo.bsky.social · 16/11/2024
there are many smart speakers and thinkers around AI/ML and/or NLP. but i find almost everything to be kinda predictable by now, minor stylistic variations on the same story. who are some *interesting* speakers i should listen/read? i want things that may surprise or inspire me.
129512
Reposted by Martina G. Vilas
Federico Adolfi @fedeadolfi.bsky.social · 17/11/2024
Any Latin Americans here working in Cognitive Science, very broadly construed? (Neuroscience, Psychology, Artificial Intelligence, Anthropology, Linguistics, Economics, Ethics, Philosophy, and more…) I thought I’d create a starter pack but I could only find a handful of us. Say hi?
215
Reposted by Martina G. Vilas
Joaquín Herrero @joakinen.filosofias.es · 17/11/2024
It is intuitive to observe some complex-looking model behavior (e.g., the classification of images of different animals using an abstract category) and infer an interesting capacity of the model (e.g., the ability to build rich representations that abstract away from particular animals).
101
Martina G. Vilas @martinagvilas.bsky.social · 17/11/2024
Hi BlueSky! 🦋 I’m a computer science PhD student with a background in cognitive neuroscience. Working at the intersection of these topics, my research focuses on reverse engineer the cognitive capacities of AI models 🧠💻 Some recent examples 👇
2243
Reposted by Martina G. Vilas
Emile van Krieken @emilevankrieken.com · 11/11/2024
I made a starter pack with the people doing something related to Neurosymbolic AI that I could find. Let me know if I missed you! go.bsky.app/RMJ8q3i
179236
Reposted by Martina G. Vilas
Maike Osborne @maosbot.bsky.social · 09/11/2024
New here? Interested in AI/ML? Check out these great starter packs! AI: go.bsky.app/SipA7it RL: go.bsky.app/3WPHcHg Women in AI: go.bsky.app/LaGDpqg NLP: go.bsky.app/SngwGeS AI and news: go.bsky.app/5sFqVNS You can also search all starter packs here: blueskydirectory.com/starter-pack...
66558212