Reposted by Tomás Vergara BrowneGaurav Kamath @grvkamath.bsky.social · 04/03/2026 🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10) 35820
Reposted by Tomás Vergara BrowneGaurav Kamath @grvkamath.bsky.social · 29/07/2025Our new paper in #PNAS (bit.ly/4fcWfma) presents a surprising finding—when words change meaning, older speakers rapidly adopt the new usage; inter-generational differences are often minor. w/ Michelle Yang, @sivareddyg.bsky.social , @msonderegger.bsky.social and @dallascard.bsky.social👇(1/12) 33317
Reposted by Tomás Vergara BrowneBenno Krojer @bennokrojer.bsky.social · 25/06/2025Started a new podcast with @tomvergara.bsky.social ! Behind the Research of AI: We look behind the scenes, beyond the polished papers 🧐🧪 If this sounds fun, check out our first "official" episode with the awesome Gauthier Gidel from @mila-quebec.bsky.social : open.spotify.com/episode/7oTc...open.spotify.com02 | Gauthier Gidel: Bridging Theory and Deep Learning, Vibes at Mila, and the Effects of AI on ArtBehind the Research of AI · Episode 1176
Reposted by Tomás Vergara BrowneBenno Krojer @bennokrojer.bsky.social · 15/04/2025Overall I loved the paper, got lots of inspiration from it and would love to be part of a similar project in the future: for example an empirical investigation of many AI papers to answer "To what extent is AI is a science?" 011
Reposted by Tomás Vergara BrowneSara Vera Marjanovic @saravera.bsky.social · 01/04/2025Models like DeepSeek-R1 🐋 mark a fundamental shift in how LLMs approach complex problems. In our preprint on R1 Thoughtology, we study R1’s reasoning chains across a variety of tasks; investigating its capabilities, limitations, and behaviour. 🔗: mcgill-nlp.github.io/thoughtology/ 15116
Reposted by Tomás Vergara BrowneParishad BehnamGhader @parishadbehnam.bsky.social · 12/03/2025Instruction-following retrievers can efficiently and accurately search for harmful and sensitive information on the internet! 🌐💣 Retrievers need to be aligned too! 🚨🚨🚨 Work done with the wonderful Nick and @sivareddyg.bsky.social 🔗 mcgill-nlp.github.io/malicious-ir/ Thread: 🧵👇mcgill-nlp.github.ioExploiting Instruction-Following Retrievers for Malicious Information RetrievalParishad BehnamGhader, Nicholas Meade, Siva Reddy 1118
Reposted by Tomás Vergara BrowneXing Han Lu @xhluca.bsky.social · 10/03/2025Agents like OpenAI Operator can solve complex computer tasks, but what happens when users use them to cause harm, e.g. spread misinformation? To find out, we introduce SafeArena (safearena.github.io), a benchmark to assess the capabilities of web agents to complete harmful web tasks. A thread 👇 1167
Reposted by Tomás Vergara BrowneVagrant Gautam @dippedrusk.com · 20/11/2024After a fun and long #EMNLP2024 I'm now travelling AGAIN to Uppsala 🇸🇪, to speak at the Transdisciplinary Queer Futures of AI Conference! Any Sweden/Uppsala recs? 1161