Sign in

Ada

@adadtur.bsky.social
575 followers 814 following 15 posts

she/her incoming @ Blender Lab, UIUC prev @ McGillNLP & Mila occasionally live on ckut 90.3 fm :-) adadtur.github.io

PostsRepliesMedia
Reposted by Ada
Gaurav Kamath @grvkamath.bsky.social · 05/06/2026
Super cool project that I really enjoyed being part of! tl;dr - when a human or model encounters new visual stimuli, how closely is it mapped to other, previously encountered concepts? (Come for weird dog-monster, stay for the science 🙂 )
061
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 05/06/2026
My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision
1124
Ada @adadtur.bsky.social · 05/06/2026
Super excited to finally announce my latest research “Would you still call this Dax? Novel Visual References in VLMs and Humans”! We studied how vision-language models (VLMs) adopt new visual concepts and map them to language compared to humans, and found that…
350
Reposted by Ada
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026
Takeaway: reasoning LLMs are getting better and better on math and code—deterministic reasoning tasks. But we should also evaluate them on open-ended, inherently uncertain everyday reasoning! (9/10)
182
Reposted by Ada
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026
🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10)
35820
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Beyond our controlled setup, we also show how LatentLens works much better than baselines on off-the-shelf Qwen2-VL-7B-Instruct
131
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Building a VLM can be surprisingly simple: You keep both the LLM and vision encoder frozen, you just train a small MLP that projects into the LLM embedding space as prefixes. That’s it 😮 But how and why does that work? How do visual tokens relate to language, i.e. do they have interpretable NNs?
141
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
Paper: arxiv.org/abs/2602.00462 Code: github.com/McGill-NLP/... Demo: tinyurl.com/ce57mn4v Couldn't have imagined better collaborators to wrap up the phd: Shravan Nayak @oscmansan.bsky.social @vaibhavadlakha.bsky.social @delliott.bsky.social @sivareddyg.bsky.social @mariusmosbach.bsky.social
github.com
GitHub - McGill-NLP/latentlens: Code and data for the paper "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs"
Code and data for the paper "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs" - McGill-NLP/latentlens
131
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 11/02/2026
🚨New paper Are visual tokens going into an LLM interpretable 🤔 Existing methods (e.g. logit lens) and assumptions would lead you to think “not much”... We propose LatentLens and show that most visual tokens are interpretable across *all* layers 💡 Details 🧵
1337
Reposted by Ada
Emily M. Bender @emilymbender.bsky.social · 16/09/2025
"Not only is the ratio of AI’s resource rapacity to its productive utility indefensibly and irremediably skewed, AI-made material is itself a waste product: flimsy, shoddy, disposable, a single-use plastic of the mind." >>
25011
Reposted by Ada
Merriam-Webster @merriam-webster.com · 03/09/2025
enshittification | noun | when a digital platform is made worse for users, in order to increase profits
499290238528
Reposted by Ada
Iain @iainnd.bsky.social · 27/08/2025
Look what they did to Notepad. Shut the fuck up. This is Notepad. You are not welcome here. Oh yeah "Let me use Copilot for Notepad". "I'm going to sign into my account for Notepad". What the fuck are you talking about. It's Notepad.
Windows Notepad, the native simple text editor, now has formatting options and a Copilot button.
442172754523
Reposted by Ada
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025
Our new paper in #PNAS (bit.ly/4fcWfma) presents a surprising finding—when words change meaning, older speakers rapidly adopt the new usage; inter-generational differences are often minor. w/ Michelle Yang, ‪@sivareddyg.bsky.social‬ , @msonderegger.bsky.social‬ and @dallascard.bsky.social‬👇(1/12)
33317
Reposted by Ada
Beyza Bozdag @beyzabozdag.bsky.social · 13/05/2025
Thrilled to announce our new survey that explores the exciting possibilities and troubling risks of computational persuasion in the era of LLMs 🤖💬 📄Arxiv: arxiv.org/pdf/2505.07775 💻 GitHub: github.com/beyzabozdag/...
195
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 25/06/2025
Started a new podcast with @tomvergara.bsky.social ! Behind the Research of AI: We look behind the scenes, beyond the polished papers 🧐🧪 If this sounds fun, check out our first "official" episode with the awesome Gauthier Gidel from @mila-quebec.bsky.social : open.spotify.com/episode/7oTc...
open.spotify.com
02 | Gauthier Gidel: Bridging Theory and Deep Learning, Vibes at Mila, and the Effects of AI on Art
Behind the Research of AI · Episode
1176
Reposted by Ada
The Washington Post @washingtonpost.com · 25/06/2025
Zohran Mamdani, a 33-year-old state assemblyman, declared victory in New York City’s Democratic mayoral primary after Andrew Cuomo conceded the race. “Tonight we made history,” Mamdani said, addressing his supporters. wapo.st/44yMVoI
1394386515
Reposted by Ada
Alexandria Ocasio-Cortez @aoc.bsky.social · 21/06/2025
Mahmoud Khalil is finally home with his beautiful wife and newborn son. Each one of the 104 days he spent detained was a grave injustice. From the moment of his detention, @ccrjustice.org + @aclu.org engaged my office as we worked closely to help secure his release. They did remarkable work here.
269217553005
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 13/06/2025
The facts: We release (MVPBench) with around 55K videos (grouped as *minimal video pairs*) from diverse physical understanding sources Arxiv: arxiv.org/abs/2506.09987 Huggingface: huggingface.co/datasets/fac... GitHub: github.com/facebookrese... Leaderboard: huggingface.co/spaces/faceb...
arxiv.org
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based ...
131
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 13/06/2025
Excited to share the results of my recent internship! We ask 🤔 What subtle shortcuts are VideoLLMs taking on spatio-temporal questions? And how can we instead curate shortcut-robust examples at a large-scale? We release: MVPBench Details 👇🔬
1165
Reposted by Ada
Congressman Glenn Ivey (MD-04) @ivey.house.gov · 26/05/2025
Today, I was denied access to seeing my constituent, Mr. Kilmar Abrego Garcia. If there is nothing to hide, cut the crap. Let his lawyer and I check on him.
7263872410682
Reposted by Ada
The Washington Post @washingtonpost.com · 22/05/2025
Breaking news: The Trump administration revoked Harvard’s ability to enroll foreign students, saying it allowed anti-American agitators. Existing foreign students must transfer or risk losing their legal status, DHS said.
wapo.st
Live updates: Trump administration revokes Harvard’s ability to enroll foreign students
Get the latest news on President Donald Trump’s return to the White House and the Republican-led Congress.
81213128
Ada @adadtur.bsky.social · 07/05/2025
when in albuquerque…
040
Ada @adadtur.bsky.social · 03/05/2025
We won a Senior Area Chair Award at NAACL!! Many thanks again to my amazing coauthors Gaurav Kamath and @sivareddyg.bsky.social :-)
2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Linguistic Theories, Cognitive Modeling, and Psycholinguistics Senior Area Chair Award presented to Ada Tur, Gaurav Kamath, and Siva Reddy. May 2, 2025. Signed by Colin Cherry (General Chair), Luis Chiruzzo, Alan Ritter, and Lu Wang (Program Chairs)
0132
Reposted by Ada
Marius Mosbach @mariusmosbach.bsky.social · 02/05/2025
Check out Gaurav's video on their #NAACL paper and find @adadtur.bsky.social at the conference 👇
0111
Reposted by Ada
Benno Krojer @bennokrojer.bsky.social · 01/05/2025
Great work from labmates on LLMs vs humans regarding linguistic preferences: You know when a sentence kind of feels off e.g. "I met at the park the man". So in what ways do LLMs follow these human intuitions?
073
Reposted by Ada
Siva Reddy @sivareddyg.bsky.social · 01/05/2025
Ada is an undergrad and will soon be looking for PhDs. Gaurav is a PhD student looking for intellectually stimulating internships/visiting positions. They did most of the work without much of my help. Highly recommend them. Please reach out to them if you have any positions.
arxiv.org
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
Though English sentences are typically inflexible vis-à-vis word order, constituents often show far more variability in ordering. One prominent theory presents the notion that constituent ordering is ...
162
Reposted by Ada
Siva Reddy @sivareddyg.bsky.social · 01/05/2025
Incredibly proud of my students @adadtur.bsky.social and Gaurav Kamath for winning a SAC award at #NAACL2025 for their work on assessing how LLMs model constituent shifts.
arxiv.org
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
Though English sentences are typically inflexible vis-à-vis word order, constituents often show far more variability in ordering. One prominent theory presents the notion that constituent ordering is ...
1175
Reposted by Ada
Mila - Institut québécois d'IA @mila-quebec.bsky.social · 01/05/2025
Congratulations to Mila members @adadtur.bsky.social , Gaurav Kamath and @sivareddyg.bsky.social for their SAC award at NAACL! Check out Ada's talk in Session I: Oral/Poster 6. Paper: arxiv.org/abs/2502.05670
0137
Reposted by Ada
Senator Ed Markey @markey.senate.gov · 22/04/2025
I filmed this yesterday on my way to Lousiana where my constituent Rümeysa Öztürk is being wrongfully held by ICE. I’m there now demanding her release. More to come.
921300716154
Reposted by Ada
Sara Vera Marjanovic @saravera.bsky.social · 01/04/2025
Models like DeepSeek-R1 🐋 mark a fundamental shift in how LLMs approach complex problems. In our preprint on R1 Thoughtology, we study R1’s reasoning chains across a variety of tasks; investigating its capabilities, limitations, and behaviour. 🔗: mcgill-nlp.github.io/thoughtology/
A circular diagram with a blue whale icon at the center. The diagram shows 8 interconnected research areas around LLM reasoning represented as colored rectangular boxes arranged in a circular pattern. The areas include: §3 Analysis of Reasoning Chains (central cloud), §4 Scaling of Thoughts (discussing thought length and performance metrics), §5 Long Context Evaluation (focusing on information recall), §6 Faithfulness to Context (examining question answering accuracy), §7 Safety Evaluation (assessing harmful content generation and jailbreak resistance), §8 Language & Culture (exploring moral reasoning and language effects), §9 Relation to Human Processing (comparing cognitive processes), §10 Visual Reasoning (covering ASCII generation capabilities), and §11 Following Token Budget (investigating direct prompting techniques). Arrows connect the sections in a clockwise flow, suggesting an iterative research methodology.
15116
Reposted by Ada
Gravel Influencer @gravelinfluencer.bsky.social · 26/03/2025
Not sure if this has been shared here yet, but this is video of Rumeysa Ozturk's arrest posted by WCVB. It's terrifying.
4945612279
Reposted by Ada
Governor Tim Walz @governorwalz.mn.gov · 25/03/2025
I would like to nominate Maxwell Smart for national security advisor.
media.tenor.com
Pm Private GIF
ALT: Pm Private GIF
30087651016837
Reposted by Ada
Dr Abeba Birhane @abeba.blacksky.app · 15/03/2025
this is outrageously terrible but if there's one thing the Trump administration’s policy is doing so far it is expediting public outrage and turning the masses against everything AI www.wired.com/story/ai-saf...
wired.com
Under Trump, AI Scientists Are Told to Remove ‘Ideological Bias’ From Powerful Models
A directive from the National Institute of Standards and Technology eliminates mention of “AI safety” and “AI fairness.”
813130
Reposted by Ada
Parishad BehnamGhader @parishadbehnam.bsky.social · 12/03/2025
Instruction-following retrievers can efficiently and accurately search for harmful and sensitive information on the internet! 🌐💣 Retrievers need to be aligned too! 🚨🚨🚨 Work done with the wonderful Nick and @sivareddyg.bsky.social 🔗 mcgill-nlp.github.io/malicious-ir/ Thread: 🧵👇
mcgill-nlp.github.io
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
Parishad BehnamGhader, Nicholas Meade, Siva Reddy
1118
Ada @adadtur.bsky.social · 11/03/2025
Super excited that this is finally out! We evaluated leading LLM-based web agents from OpenAI, Anthropic, and more, on our new benchmark SafeArena and found that many are surprisingly compliant with malicious requests. Check out the leaderboard here: huggingface.co/spaces/McGil...
huggingface.co
Safearena Leaderboard - a Hugging Face Space by McGill-NLP
SafeArena Leaderboard
081
Reposted by Ada
Spandana Gella @spandanagella.bsky.social · 10/03/2025
Web agents powered by LLMs can solve complex tasks, but our analysis shows that they can also be easily misused to automate harmful tasks. See the thread below for more details on our new web agent safety benchmark: SafeArena and Agent Risk Assessment framework (ARIA).
052
Reposted by Ada
Karolina Stańczak @karstanczak.bsky.social · 10/03/2025
The potential for malicious misuse of LLM agents is a serious threat. That's why we created SafeArena, a safety benchmark for web agents. See the thread and our paper for details: arxiv.org/abs/2503.04957 👇
arxiv.org
SafeArena: Evaluating the Safety of Autonomous Web Agents
LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an onlin...
092
Reposted by Ada
Arkil Patel @arkil.bsky.social · 10/03/2025
Llamas browsing the web look cute, but they are capable of causing a lot of harm! Check out our new Web Agents ∩ Safety benchmark: SafeArena! Paper: arxiv.org/abs/2503.04957
093
Reposted by Ada
Xing Han Lu @xhluca.bsky.social · 10/03/2025
With ARIA, we find that Claude is substantially safer than Qwen, which very rarely refuses user requests, indicating limited safeguards for web-oriented tasks.
141
Reposted by Ada
Xing Han Lu @xhluca.bsky.social · 10/03/2025
The harmfulness of LLMs varies: whereas Claude-3.5 Sonnet refuses a majority of harmful tasks, Qwen-2-VL completes over a quarter of the 250 harmful tasks we designed for this benchmark. Moreover, a GPT-4o agent completes an alarming number of unsafe requests, despite extensive safety training.
131
Reposted by Ada
Xing Han Lu @xhluca.bsky.social · 10/03/2025
Agents like OpenAI Operator can solve complex computer tasks, but what happens when users use them to cause harm, e.g. spread misinformation? To find out, we introduce SafeArena (safearena.github.io), a benchmark to assess the capabilities of web agents to complete harmful web tasks. A thread 👇
1167
Reposted by Ada
Marisa Kabas @marisakabas.bsky.social · 09/03/2025
New — I wrote about how the Congressional Dems who voted to censure Rep. Al Green don’t understand their enemy and aren’t made for this moment:
thehandbasket.co
The time for a spine
Democrats who censured Rep. Al Green are as clueless as they are feckless.
742492491
Reposted by Ada
Seth Karten @sethkarten.ai · 07/03/2025
Can a Large Language Model (LLM) with zero Pokémon-specific training achieve expert-level performance in competitive Pokémon battles? Introducing PokéChamp, our minimax LLM agent that reaches top 30%-10% human-level Elo on Pokémon Showdown! New paper on arXiv and code on github!
1335
Reposted by Ada
Andrew Gordon Wilson @andrewgwils.bsky.social · 05/03/2025
My new paper "Deep Learning is Not So Mysterious or Different": arxiv.org/abs/2503.02113. Generalization behaviours in deep learning can be intuitively understood through a notion of soft inductive biases, and formally characterized with countable hypothesis bounds! 1/12
620851
Reposted by Ada
Carissa Véliz @carissaveliz.bsky.social · 05/03/2025
What if we didn't design AI to behave as impersonators? #AIEthics. The bots encouraged users' dangerous beliefs. "If given by a human therapist, he added, those answers could have resulted in the loss of a license to practice, or civil or criminal liability." www.nytimes.com/2025/02/24/h...
nytimes.com
Human Therapists Prepare for Battle Against A.I. Pretenders
Chatbots posing as therapists may encourage users to commit harmful acts, the nation’s largest psychological organization warned federal regulators.
1146
Reposted by Ada
Karolina Stańczak @karstanczak.bsky.social · 04/03/2025
📢New Paper Alert!🚀 Human alignment balances social expectations, economic incentives, and legal frameworks. What if LLM alignment worked the same way?🤔 Our latest work explores how social, economic, and contractual alignment can address incomplete contracts in LLM alignment🧵
12713
Reposted by Ada
George Takei @georgetakei.bsky.social · 01/03/2025
A reminder for these times.
6458210317563
Reposted by Ada
John Skiles Skinner @skiles.blue · 01/03/2025
I got laid off today, with the rest of 18F. 18F was an elite federal software shop. We made gov't websites work better, more efficiently for the American people. We saved taxpayers from getting screwed over by contractors. And were fired for it. We made this website to tell our story: 18f.org
18f.org
We're not done yet | 18F
31689271430856
Reposted by Ada
Arkil Patel @arkil.bsky.social · 21/02/2025
Paper: arxiv.org/pdf/2502.14678 Data: tinyurl.com/chase-data Code: github.com/McGill-NLP/C...
arxiv.org
021