Sign in

Byron Wallace

@byron.bsky.social
2.5K followers 342 following 11 posts

Assoc. Prof in CS @ Northeastern, NLP/ML & health & etc. He/him.

PostsRepliesMedia
Reposted by Byron Wallace
Sheridan Feucht @sfeucht.bsky.social · 28/09/2026
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
1107
Reposted by Byron Wallace
kaijie-mo.bsky.social @kaijie-mo.bsky.social · 11/06/2026
“Dimicillin” isn’t real. We made it up. Yet many LLMs still call it an antibiotic. Across 9 models and 653 drugs, we find that drug-name affixes alone can drive pharmacological reasoning. Models often rely on morphology over facts. We trace this shortcut from behavior to mechanism. 🧵
1176
Byron Wallace @byron.bsky.social · 12/05/2026
Surgically editing prompts to vary a factor of interest (like gender) is an intuitive way of analyzing model behavior and sensitivity. But @zihaogavinyang.bsky.social shows that we should really compare the results from such perturbations to those observed when, e.g., we simply paraphrase inputs 👇
080
Reposted by Byron Wallace
Hye Sun Yun @hyesunyun.bsky.social · 08/04/2026
Patients ask LLMs medical questions — but how they phrase it matters more than it should. Our new preprint explores how different phrasings of patient health questions can lead to inconsistent conclusions, even with the same evidence. [1/6] Full Paper: arxiv.org/abs/2604.05051
2255
Reposted by Byron Wallace
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Reposted by Byron Wallace
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️
24811
Reposted by Byron Wallace
kaijie-mo.bsky.social @kaijie-mo.bsky.social · 21/01/2026
Hello world 👋 My first paper at UT Austin! We ask: what happens when medical “evidence” fed into an LLM is wrong? Should your AI stay faithful, or should it play it safe when the evidence is harmful? We show that frontier LLMs accept counterfactual medical evidence at face value.🧵
3155
Byron Wallace @byron.bsky.social · 05/11/2025
Check out @hibaahsan.bsky.social's paper on spotting (problematic) racial biases in LLMs for healthcare applications 👇
020
Reposted by Byron Wallace
Ai2 @ai2.bsky.social · 24/10/2025
3/ 🏥 A separate team at Northeastern located where certain signals live inside Olmo and made targeted edits that reduced biased clinical predictions. This kind of audit is only possible because Olmo exposes all its components. → buff.ly/HkChr4Q
101
Byron Wallace @byron.bsky.social · 24/10/2025
Chantal (and Vinith) find that you can jailbreak LLMs with syntax! Some examples: cshaib.github.io/syntax_domai...
020
Byron Wallace @byron.bsky.social · 22/10/2025
Now to appear at #EMNLP2025 (Findings). We've added more models and experiments: arxiv.org/abs/2502.13319
020
Byron Wallace @byron.bsky.social · 01/10/2025
Can we distill *circuits* from teacher models into smaller students? 👇
010
Reposted by Byron Wallace
David Bau @davidbau.bsky.social · 27/09/2025
Who is going to be at #COLM2025? I want to draw your attention to a COLM paper by my student @sfeucht.bsky.social that has totally changed the way I think and teach about LLM representations. The work is worth knowing. And you can meet Sheridan at COLM, Oct 7! bsky.app/profile/sfe...
1398
Byron Wallace @byron.bsky.social · 24/09/2025
Can we quantify what makes some text read like AI "slop"? We tried 👇
081
Reposted by Byron Wallace
Naomi Saphra @nsaphra.bsky.social · 17/09/2025
Our new paper asks: what is the goal of “natural language verbalization” interpretability approaches? If a verbalizer is supposed to tell us something about what’s in the target LM and NOT just what’s in the verbalizer LM, how do we actually evaluate that?
0133
Reposted by Byron Wallace
Millicent Li @millicentli.bsky.social · 17/09/2025
Wouldn’t it be great to have questions about LM internals answered in plain English? That’s the promise of verbalization interpretability. Unfortunately, our new paper shows that evaluating these methods is nuanced—and verbalizers might not tell us what we hope they do. 🧵👇1/8
1268
Reposted by Byron Wallace
Hye Sun Yun @hyesunyun.bsky.social · 25/08/2025
Thrilled to share our research showing how LLM models can be influenced by bias from "spun" medical literature is now featured in Northeastern's Khoury news! This shows critical insights as AI enters healthcare. The full paper can be found at arxiv.org/abs/2502.07963
khoury.northeastern.edu
As AI expands into medicine, Northeastern study finds AI models influenced by medical bias  - Khoury College of Computer Sciences
Humans can be easily influenced by language that is one-sided, especially in complex fields like medicine. But a new Khoury-led study shows that large language models, too, can be tricked […]
031
Reposted by Byron Wallace
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
Reposted by Byron Wallace
Yuval Pinter @uvp.bsky.social · 19/07/2025
Chatted with @byron.bsky.social at icml about my recent work, so look out for his upcoming "Tokenization is More Than More Than Compression".
1131
Reposted by Byron Wallace
Lily Chen @lilywchen.bsky.social · 01/07/2025
Are we fact-checking medical claims the right way? 🩺🤔 Probably not. In our study, even experts struggled to verify Reddit health claims using end-to-end systems. We show why—and argue fact-checking should be a dialogue, with patients in the loop arxiv.org/abs/2506.20876 🧵1/
An overview of our AI-in-the-loop expert study pipeline: given a claim from a subreddit, we extract the PIO elements and retrieve the evidence automatically. The evidence, its context, and the evidence are then presented to a medical expert to provide a judgment and a rationale for the factuality of the claim.
162
Reposted by Byron Wallace
Sheridan Feucht @sfeucht.bsky.social · 07/04/2025
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
17518
Reposted by Byron Wallace
Chantal @chantalsh.bsky.social · 10/03/2025
I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏
corp.oup.com
Oxford Word of the Year 2024 - Oxford University Press
The Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world.
0108
Reposted by Byron Wallace
Jessy Li @jessyjli.bsky.social · 25/02/2025
🌟Job ad🌟 We (@gregdnlp.bsky.social, @mattlease.bsky.social and I) are hiring a postdoc fellow within the CosmicAI Institute, to do galactic work with LLMs and generative AI! If you would like to push the frontiers of foundation models to help solve myths of the universe, please apply!
0136
Reposted by Byron Wallace
Hiba Ahsan @hibaahsan.bsky.social · 22/02/2025
LLMs are known to perpetuate social biases in clinical tasks. Can we locate and intervene upon LLM activations that encode patient demographics like gender and race? 🧵 Work w/ @arnabsensharma.bsky.social, @silvioamir.bsky.social, @davidbau.bsky.social, @byron.bsky.social arxiv.org/abs/2502.13319
3177
Reposted by Byron Wallace
Hye Sun Yun @hyesunyun.bsky.social · 15/02/2025
🚨 Do LLMs fall for spin in medical literature? 🤔 In our new preprint, we find that LLMs are susceptible to biased reporting of clinical treatment benefits in abstracts—more so than human experts. 📄🔍 [1/7] Full Paper: arxiv.org/abs/2502.07963 🧵👇
36325
Reposted by Byron Wallace
Somin W @sominw.bsky.social · 11/02/2025
📢 Can we trace a small distilled model back to its teacher? 🤔New work (w/ @chantalsh.bsky.social, @silvioamir.bsky.social & @byron.bsky.social) finds some footprints left by LLMs in distillation! [1/6] 🔗 Full paper: arxiv.org/abs/2502.06659
arxiv.org
Who Taught You That? Tracing Teachers in Model Distillation
Model distillation -- using outputs from a large teacher model to teach a small student model -- is a practical means of creating efficient models for a particular task. We ask: Can we identify a stud...
182
Reposted by Byron Wallace
David Bau @davidbau.bsky.social · 31/01/2025
DeepSeek R1 shows how important it is to be studying the internals of reasoning models. Try our code: Here @canrager.bsky.social shows a method for auditing AI bias by probing the internal monologue. dsthoughts.baulab.info I'd be interested in your thoughts.
dsthoughts.baulab
1289
Reposted by Byron Wallace
ijmarshall.bsky.social @ijmarshall.bsky.social · 27/01/2025
📣 🌍 We're hiring for 2 Machine Learning researchers to join SOLACE-AI @kingscollegelondon.bsky.social , funded by @wellcometrust.bsky.social . This is your chance to develop cutting-edge AI to directly impact global health responses to climate emergencies. jobs.ac.uk/job/DLM377
023
Reposted by Byron Wallace
Luca Soldaini 🎀 @soldaini.net · 26/11/2024
OLMo 2 is out 🥳 7B and 13B trained on 5T tokens, and meticulousy instruction tuned using Tulu 3 recipe. Simply the best fully open models yet. Really proud of the work & the amazing team at @ai2.bsky.social
926044
Byron Wallace @byron.bsky.social · 09/11/2024
I'll be @ #EMNLP2024 if anyone wants to find snobby coffee / despair about election / or I guess talk research. Some work to be presented👇
1130
Reposted by Byron Wallace
Jered McInerney @dmcinerney.bsky.social · 28/02/2024
Our work on reducing diagnostic errors with interpretable risk prediction is now on arXiv! We retrieve evidence from a patient’s record, visualize how it informs a prediction, and test it in a realistic setting. 👇 (1/6) arxiv.org/abs/2402.10109 w/ @byron.bsky.social and @jwvdm.bsky.social
121
Reposted by Byron Wallace
Jessy Li @jessyjli.bsky.social · 14/11/2023
To appear #EMNLP2023! Can LMs simplify medical texts in non-English languages? We introduce⚕️MultiCochrane: the *first* multilingual, aligned dataset for this. arxiv.org/abs/2305.12532. Led by Sebastian Joseph, also w/ @byron.bsky.social Wei Xu
1135
Reposted by Byron Wallace
David Bau @davidbau.bsky.social · 11/11/2023
Work with Jiuding Sun, Andrew Yuan, and Byron Wallace @byron.bsky.social If you're going to be at EMNLP/CoNLL/BlackboxNLP in Singapore, look for her poster at CoNLL! Koyena's FutureLens preprint, code, and demo are on the project website at future.baulab.info
021
Reposted by Byron Wallace
Jessy Li @jessyjli.bsky.social · 30/10/2023
Can we use LLMs to help disseminate medical information more broadly? @byron.bsky.social, Mike Mackert, Wei Xu and I are hosting an online panel today at the HARC conference at 4:30 EST/3:30 CST on Simplifying Medical Texts with Large Language Models! harcconf.org/agenda-monda...
072
Reposted by Byron Wallace
David Bau @davidbau.bsky.social · 25/10/2023
LLMs contain Function Vectors! Eric Todd has a really interesting new preprint on arxiv functions.baulab.info showing LLMs contain vector representations of functions that compose and apply in diverse contexts. Could be a powerful analysis tool. More in his twitter thread: x.com/ericwtodd/st...
functions.baulab.info
Function Vectors in Large Language Models
Understanding the internal computations of huge autoregressive transformer neural network language models during in-context learning.
1123
Byron Wallace @byron.bsky.social · 25/10/2023
Check out Jered's work combining LLM (zero-shot) extracted features with simple linear models for healthcare data👇 (This *is* where we're doing shameless research self-promo threads now, right?)
041