Sign in

Marius Mosbach

@mariusmosbach.bsky.social
952 followers 275 following 28 posts

#NLP Postdoc at Mila - Quebec AI Institute & McGill University mariusmosbach.com

PostsRepliesMedia
Reposted by Marius Mosbach
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025
Our new paper in #PNAS (bit.ly/4fcWfma) presents a surprising finding—when words change meaning, older speakers rapidly adopt the new usage; inter-generational differences are often minor. w/ Michelle Yang, ‪@sivareddyg.bsky.social‬ , @msonderegger.bsky.social‬ and @dallascard.bsky.social‬👇(1/12)
33317
Reposted by Marius Mosbach
Ingmar Weber @ingmarweber.de · 18/07/2025
🚨Job Alert W2 (TT W3) Professorship in Computer Science "AI for People & Society" @saarland-informatics-campus.de/@uni-saarland.de is looking to appoint an outstanding individual in the field of AI for people and society who has made significant contributions in one or more of the following areas:
11418
Reposted by Marius Mosbach
Abhilasha Ravichander @lasha.bsky.social · 22/07/2025
📣 Life update: Thrilled to announce that I’ll be starting as faculty at the Max Planck Institute for Software Systems this Fall! I’ll be recruiting PhD students in the upcoming cycle, as well as research interns throughout the year: lasharavichander.github.io/contact.html
Kaiserslautern, Germany
139212
Reposted by Marius Mosbach
Sebastian Bordt @sbordt.bsky.social · 14/07/2025
I'm at #ICML in Vancouver this week, hit me up if you want to chat about pre-training experiments or explainable machine learning. You can find me at these posters: Tuesday: How Much Can We Forget about Data Contamination? icml.cc/virtual/2025...
111
Reposted by Marius Mosbach
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Reposted by Marius Mosbach
Sebastian Bordt @sbordt.bsky.social · 08/07/2025
Have you ever wondered whether a few times of data contamination really lead to benchmark overfitting?🤔 Then our latest #ICML paper about the effect of data contamination on LLM evals might be for you!🚀 Paper: arxiv.org/abs/2410.03249 👇🧵
1121
Reposted by Marius Mosbach
Valentina Pyatkin @valentinapy.bsky.social · 03/07/2025
💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR. But the set of constraints and verifier functions is limited and most models overfit on IFEval. We introduce IFBench to measure model generalization to unseen constraints.
1295
Reposted by Marius Mosbach
Cesare @cesare-spinoso.bsky.social · 26/06/2025
A blizzard is raging through Montreal when your friend says “Looks like Florida out there!” Humans easily interpret irony, while LLMs struggle with it. We propose a 𝘳𝘩𝘦𝘵𝘰𝘳𝘪𝘤𝘢𝘭-𝘴𝘵𝘳𝘢𝘵𝘦𝘨𝘺-𝘢𝘸𝘢𝘳𝘦 probabilistic framework as a solution. Paper: arxiv.org/abs/2506.09301 to appear @ #ACL2025 (Main)
1157
Reposted by Marius Mosbach
Benno Krojer @bennokrojer.bsky.social · 25/06/2025
Started a new podcast with @tomvergara.bsky.social ! Behind the Research of AI: We look behind the scenes, beyond the polished papers 🧐🧪 If this sounds fun, check out our first "official" episode with the awesome Gauthier Gidel from @mila-quebec.bsky.social : open.spotify.com/episode/7oTc...
open.spotify.com
02 | Gauthier Gidel: Bridging Theory and Deep Learning, Vibes at Mila, and the Effects of AI on Art
Behind the Research of AI · Episode
1176
Reposted by Marius Mosbach
Valentina Pyatkin @valentinapy.bsky.social · 17/06/2025
Interested in shaping the progress of responsible AI and meeting leading researchers in the field? SoLaR@COLM 2025 is looking for paper submissions and reviewers! 🤖 ML track: algorithms, math, computation 📚 Socio-technical track: policy, ethics, human participant research
181
Reposted by Marius Mosbach
Xing Han Lu @xhluca.bsky.social · 14/06/2025
"Build the web for agents, not agents for the web" This position paper argues that rather than forcing web agents to adapt to UIs designed for humans, we should develop a new interface optimized for web agents, which we call Agentic Web Interface (AWI). arxiv.org/abs/2506.10953
064
Reposted by Marius Mosbach
Benno Krojer @bennokrojer.bsky.social · 13/06/2025
Excited to share the results of my recent internship! We ask 🤔 What subtle shortcuts are VideoLLMs taking on spatio-temporal questions? And how can we instead curate shortcut-robust examples at a large-scale? We release: MVPBench Details 👇🔬
1165
Reposted by Marius Mosbach
Badr M. Abdullah, PhD @badralabsi.bsky.social · 10/06/2025
New paper in Interspeech 2025 🚨 @interspeech.bsky.social A Robust Model for Arabic Dialect Identification using Voice Conversion Paper 📝 arxiv.org/pdf/2505.24713 Demo 🎙️https://shorturl.at/rrMm6 #Arabic #SpeechTech #NLProc #AI #Speech #ArabicDialects #Interspeech2025 #ArabicNLP
112
Reposted by Marius Mosbach
Ziling Cheng @ziling-cheng.bsky.social · 06/06/2025
Do LLMs hallucinate randomly? Not quite. Our #ACL2025 (Main) paper shows that hallucinations under irrelevant contexts follow a systematic failure mode — revealing how LLMs generalize using abstract classes + context cues, albeit unreliably. 📎 Paper: arxiv.org/abs/2505.22630 1/n
14517
Reposted by Marius Mosbach
Michael Hahn @m-hahn.bsky.social · 05/05/2025
Chain-of-Thought (CoT) reasoning lets LLMs solve complex tasks, but long CoTs are expensive. How short can they be while still working? Our new ICML paper tackles this foundational question.
2122
Reposted by Marius Mosbach
Vagrant Gautam @dippedrusk.com · 03/05/2025
Come to my keynote tomorrow at the first official @queerinai.com workshop at #NAACL2025 to hear about how trans languaging is complex and cool, and how this makes it extra difficult to process computationally. I will have SO many juicy examples!
Title slide: Processing Trans Languaging - Vagrant Gautam (they/xe), Saarland University, with a very brightly patterned background featuring colourful people and math symbols.
34414
Reposted by Marius Mosbach
Hadas Orgad @hadasorgad.bsky.social · 03/05/2025
Deadline extended! ⏳ The Actionable Interpretability Workshop at #ICML2025 has moved its submission deadline to May 19th. More time to submit your work 🔍🧠✨ Don’t miss out!
043
Marius Mosbach @mariusmosbach.bsky.social · 02/05/2025
Check out Gaurav's video on their #NAACL paper and find @adadtur.bsky.social at the conference 👇
0111
Reposted by Marius Mosbach
Valentina Pyatkin @valentinapy.bsky.social · 27/04/2025
I'll be at #NAACL2025: 🖇️To present my paper "Superlatives in Context", showing how the interpretation of superlatives is very context dependent and often implicit, and how LLMs handle such semantic underspecification 🖇️And we will present RewardBench on Friday Reach out if you want to chat!
1285
Reposted by Marius Mosbach
Sonia @soniajoseph.bsky.social · 25/04/2025
I’m really excited about Diffusion Steering Lens, an intuitive and elegant new “logit lens” technique for decoding the attention and MLP blocks of vision transformers! Vision is much more expressive than language, so some new mech interp rules apply:
0113
Reposted by Marius Mosbach
Yanai Elazar @yanai.bsky.social · 25/04/2025
💡 New ICLR paper! 💡 "On Linear Representations and Pretraining Data Frequency in Language Models": We provide an explanation for when & why linear representations form in large (or small) language models. Led by @jackmerullo.bsky.social, w/ @nlpnoah.bsky.social & @sarah-nlp.bsky.social
34212
Reposted by Marius Mosbach
Aaron Mueller @amuuueller.bsky.social · 23/04/2025
Lots of progress in mech interp (MI) lately! But how can we measure when new mech interp methods yield real improvements over prior work? We propose 😎 𝗠𝗜𝗕: a 𝗠echanistic 𝗜nterpretability 𝗕enchmark!
Logo for MIB: A Mechanistic Interpretability Benchmark
15115
Marius Mosbach @mariusmosbach.bsky.social · 22/04/2025
Paper title of the year so far. I will be back ... have to read the paper now. Great work @saxon.me !
150
Marius Mosbach @mariusmosbach.bsky.social · 16/04/2025
Checkout Benno's notes about our impact of interpretability paper 👇. Also, we are organizing a workshop at #ICML2025 which is inspired by some of the questions discussed in the paper: actionable-interpretability.github.io
actionable-interpretability.github.io
General Information
ICML 2025 - Vancouver
0113
Marius Mosbach @mariusmosbach.bsky.social · 11/04/2025
Our Thoughtology 💭 paper finally made it to arXiv (after being on hold for more than a week 😵‍💫). Make sure to check it out if you are interested in analyzing reasoning chains of LLMs. 🔗: arxiv.org/abs/2504.07128
arxiv.org
DeepSeek-R1 Thoughtology: Let's <think> about LLM Reasoning
Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-st...
0150
Marius Mosbach @mariusmosbach.bsky.social · 09/04/2025
Check out our new paper on unlearning for LLMs 🤖. We show that *not all data are unlearned equally* and argue that future work on LLM unlearning should take properties of the data to be unlearned into account. This work was lead by my intern @a-krishnan.bsky.social 🔗: arxiv.org/abs/2504.05058
Diagram illustrating a hypothesis about knowledge unlearning in language models. The left side shows a training corpus with varying frequencies of facts, such as 'Montreal is a city in Quebec' (high frequency) and 'Atlantis is a city in the ocean' (lower frequency). The center shows a language model being trained on this data, then undergoing unlearning. The right side demonstrates the 'Forget Quality' results, where the model more effectively unlearns the less frequent fact ('Atlantis is in Greece') while retaining the more frequent knowledge. Labels A, B, and C mark key points in the hypothesis: A (frequency variations in training data), B (influence of frequency), and C (unlearning effectiveness).
1325
Reposted by Marius Mosbach
Amirhossein Kazemnejad @a-kazemnejad.bsky.social · 04/04/2025
Introducing nanoAhaMoment: Karpathy-style, single file RL for LLM library (<700 lines) - super hackable - no TRL / Verl, no abstraction💆‍♂️ - Single GPU, full param tuning, 3B LLM - Efficient (R1-zero countdown < 10h) comes with a from-scratch, fully spelled out YT video [1/n]
181
Marius Mosbach @mariusmosbach.bsky.social · 01/04/2025
Having access to the reasoning chains of models like DeepSeek-R1 allows us to systematically study the reasoning behavior of LLMs, an endeavour which we term Thoughtology. Check out our paper below for what we have found 👇
080
Marius Mosbach @mariusmosbach.bsky.social · 31/03/2025
Check out our new workshop on Actionable Interpretability @ ICML 2025. We are also looking forward to submissions that take a position on the future of interpretability research more broadly. 👇
091
Reposted by Marius Mosbach
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Marius Mosbach
Parishad BehnamGhader @parishadbehnam.bsky.social · 12/03/2025
Instruction-following retrievers can efficiently and accurately search for harmful and sensitive information on the internet! 🌐💣 Retrievers need to be aligned too! 🚨🚨🚨 Work done with the wonderful Nick and @sivareddyg.bsky.social 🔗 mcgill-nlp.github.io/malicious-ir/ Thread: 🧵👇
mcgill-nlp.github.io
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
Parishad BehnamGhader, Nicholas Meade, Siva Reddy
1118
Reposted by Marius Mosbach
Xing Han Lu @xhluca.bsky.social · 10/03/2025
Agents like OpenAI Operator can solve complex computer tasks, but what happens when users use them to cause harm, e.g. spread misinformation? To find out, we introduce SafeArena (safearena.github.io), a benchmark to assess the capabilities of web agents to complete harmful web tasks. A thread 👇
1167
Reposted by Marius Mosbach
Karolina Stańczak @karstanczak.bsky.social · 04/03/2025
📢New Paper Alert!🚀 Human alignment balances social expectations, economic incentives, and legal frameworks. What if LLM alignment worked the same way?🤔 Our latest work explores how social, economic, and contractual alignment can address incomplete contracts in LLM alignment🧵
12713
Reposted by Marius Mosbach
Arkil Patel @arkil.bsky.social · 21/02/2025
Presenting ✨ 𝐂𝐇𝐀𝐒𝐄: 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐧𝐠 𝐜𝐡𝐚𝐥𝐥𝐞𝐧𝐠𝐢𝐧𝐠 𝐬𝐲𝐧𝐭𝐡𝐞𝐭𝐢𝐜 𝐝𝐚𝐭𝐚 𝐟𝐨𝐫 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧 ✨ Work w/ fantastic advisors Dima Bahdanau and @sivareddyg.bsky.social Thread 🧵:
1178
Reposted by Marius Mosbach
Fabian David Schmidt @fdschmidt.bsky.social · 21/02/2025
Introducing MVL-SIB, a massively multilingual vision-language benchmark for cross-modal topic matching in 205 languages! 🤔Tasks: Given images (sentences), select topically matching sentence (image). Arxiv: arxiv.org/abs/2502.12852 HF: huggingface.co/datasets/Wue... Details👇
145
Reposted by Marius Mosbach
Shauli Ravfogel @shauli.bsky.social · 12/02/2025
Our paper "A Practical Method for Generating String Counterfactuals" has been accepted to the findings of NAACL 2025! a joint work with @matan-avitan.bsky.social , @yoavgo.bsky.social and Ryan Cotterell. We propose "Intervention Lens", a technique to explain intervention in natural language. (1/6)
1374
Reposted by Marius Mosbach
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 11/02/2025
Looking for a PhD student to come work with me on the ethical implications of NLP from September! Please share widely and point any interesting students my way! 😊
02615
Reposted by Marius Mosbach
Karen Hao @karenhao.bsky.social · 27/01/2025
As someone who has reported on AI for 7 years and covered China tech as well, I think the biggest lesson to be drawn from DeepSeek is the huge cracks it illustrates with the current dominant paradigm of AI development. A long thread. 1/
21161252340
Reposted by Marius Mosbach
Conference on Language Modeling @colmweb.org · 17/12/2024
Announcement #1: our call for papers is up! 🎉 colmweb.org/cfp.html And excited to announce the COLM 2025 program chairs @yoavartzi.com @eunsol.bsky.social @ranjaykrishna.bsky.social and @adtraghunathan.bsky.social
06624
Marius Mosbach @mariusmosbach.bsky.social · 16/12/2024
You always get the worst reviews, don't you? Time to make a difference! Become a reviewer 🤗
060
Reposted by Marius Mosbach
sophie-xhonneux.bsky.social @sophie-xhonneux.bsky.social · 12/12/2024
Come to our Spotlight Poster #4702! East Exhibition Hall A-C
0175
Reposted by Marius Mosbach
Benno Krojer @bennokrojer.bsky.social · 10/12/2024
Come by tomorrow (Wed) 11am-2pm at @neuripsconf.bsky.social Poster #1606 to chat more about AURORA 🌌, text-guided editing, and why it is arguably more interesting than image generation Or anything related to world models, evals/analysis/interp, vision+language reasoning, cogsci, academic life!
0154
Marius Mosbach @mariusmosbach.bsky.social · 26/11/2024
Sparsity is all you need?
1300
Reposted by Marius Mosbach
Marc Marone @marcmarone.com · 23/11/2024
I noticed a lot of starter packs skewed towards faculty/industry, so I made one of just NLP & ML students: go.bsky.app/vju2ux Students do different research, go on the job market, and recruit other students. Ping me and I'll add you!
10117654
Reposted by Marius Mosbach
McGill NLP @mcgill-nlp.bsky.social · 23/11/2024
Our lab members recently presented 3 papers at @emnlpmeeting.bsky.social in Miami ☀️ 📜 From interpretability to bias/fairness and cultural understanding -> 🧵
1196
Reposted by Marius Mosbach
Nathan Lambert @natolambert.bsky.social · 22/11/2024
Wtf guys, we released a SOTA 70B instruct model and a broken 70B model (which was our best) and is missing the value head. More people downloaded the model that's literally broken 🤣🤣🤣🤣
4492
Marius Mosbach @mariusmosbach.bsky.social · 21/11/2024
Free research idea 👇 train a LM head on top of this „broken“ model that matches or improves the performance of their second best model. See below for weights and more context.
1161
Marius Mosbach @mariusmosbach.bsky.social · 21/11/2024
What's the one advise you wish people would have told you before applying to a faculty position?
2192