Sign in

Will Held

@williamheld.com
2.2K followers 457 following 116 posts

Modeling Linguistic Variation to expand ownership of NLP tools Views my own, but affiliations that might influence them: ML PhD Student under Prof. Diyi Yang 2x RS Intern🦙 Pretraining Alum NYU Abu Dhabi Burqueño he/him

PostsRepliesMedia
Reposted by Will Held
Open Athena @openathena.ai · 25/06/2026
In a new blog, Russell Power explains how the Marin team nearly doubled its sustained TPU usage by creating a custom global scheduler: Iris. Iris searches every region where Marin has compute, places each job wherever capacity appears, and moves data along as needed. 🔗 openathena.ai/blog/cluster...
1102
Will Held @williamheld.com · 11/05/2026
To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error. Getting there took some work 🧵
2389
Will Held @williamheld.com · 29/10/2025
Super interested to what degree this interaction can be fine-tuned into models in a non-reversible fashion! Voice cloning is unfortunately a capability which inherently shows up in pretrained audio models. It would be great to be able to largely limit the capability at the level of model weights!
110
Reposted by Will Held
Dan Jurafsky @jurafsky.bsky.social · 24/08/2025
Now that school is starting for lots of folks, it's time for a new release of Speech and Language Processing! Jim and I added all sorts of material for the August 2025 release! With slides to match! Check it out here: web.stanford.edu/~jurafsky/sl...
web.stanford.edu
Speech and Language Processing
Speech and Language Processing
315358
Will Held @williamheld.com · 11/08/2025
"GPT-5 shows scaling laws are coming to an end"
060
Reposted by Will Held
George Pearkes @peark.es · 06/08/2025
We’ve discovered a literal miracle with almost unlimited potential and it’s being scrapped for *no reason whatsoever*. This isn’t even nihilism, it’s outright worship of death and human suffering.
47103173294
Will Held @williamheld.com · 06/08/2025
Really great pointer from Hao Zhang on the other site in relation to GPT OSS use of attention sinks. If I were to guess, the attention sink is what allows them to omit QK-Norm which has become otherwise standard. www.evanmiller.org/attention-is...
evanmiller.org
Attention Is Off By One
Let’s fix these pesky Transformer outliers using Softmax One and QuietAttention.
010
Will Held @williamheld.com · 28/07/2025
The SALT Lab is at #ACL2025 with our genius leader @diyiyang.bsky.social. Come see work from @yanzhe.bsky.social, @dorazhao.bsky.social @oshaikh.bsky.social, @michaelryan207.bsky.social, and myself at any of the talks and posters below!
Alt Text:

Conference schedule for July 28th (Monday) and July 29th (Tuesday), listing talk titles, locations, times, and authors:

July 28th, Monday:

1. Attacking Vision-Language Computer Agents via Pop-ups
Location: Hall 4/5, Time: 11:00–12:30
Authors: Yanzhe Zhang, Tao Yu, Diyi Yang


2. SPHERE: An Evaluation Card for Human-AI Systems
Location: Hall 4/5, Time: 18:00–19:30
Authors: Dora Zhao*, Qianou Ma*, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang*, Tongshuang Wu*
(asterisk denotes equal contribution)



July 29th, Tuesday:

1. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
Location: Hall 4/5, Time: 10:30–12:00
Authors: Michael J Ryan, Omar Shaikh, Aditri Bhagirath, Daniel Frees, William Barr Held, Diyi Yang


2. Distilling an End-to-End Voice Assistant Without Instruction Training Data
Location: Room 1.61, Time: 14:12 (Second Talk)
Authors: William Barr Held, Yanzhe Zhang, Weiyan Shi, Minzhi Li, Michael J Ryan, Diyi Yang


3. Mind the Gap: Static and Interactive Evaluations of Large Audio Models
Location: Room 1.61 (implied), follows previous talk
Authors: Minzhi Li*, William Barr Held*, Michael J Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, Diyi Yang
(asterisk denotes equal contribution)


4. EgoNormia: Benchmarking Physical Social Norm Understanding
Location: Hall 4/5, Time: 16:00–17:30
Authors: MohammadHossein Rezaei*, Yicheng Fu*, Phil Cuvin*, Caleb Ziems, Yanzhe Zhang, Hao Zhu, Diyi Yang
(asterisk denotes equal contribution)
030
Will Held @williamheld.com · 28/07/2025
I'm in Vienna for #ACL2025! My work is all presented tomorrow, but today you'll find me today at the poster session from 11-12:30 evangelizing my labmate Yanzhe Zhang's work on his behalf. If you're interested in the risks traditional pop-up attacks present for AI agents, come chat!
140
Will Held @williamheld.com · 03/07/2025
A while ago I mentioned that for marin.community project, this gradient increase led to problematic loss ascent which we patched with Z-loss. I was curious, does AdamC just work? So over the weekend, I ran 4 experiments—130M to 1.4B params—all at ~compute-optimal token counts...🧵
marin.community
Marin
141
Will Held @williamheld.com · 03/07/2025
kyutai.org/next/unmute has built in turn-detection on the ASR and full I/O streaming for the TTS. Solves the latency issues that I think are 90% of why people use end-to-end speech models in the first place! From the details, you can @kyutai-labs.bsky.social is focused on real-world utility.
unmute.sh
Unmute by Kyutai
Make LLMs listen and speak.
010
Reposted by Will Held
Haley L. @haleyhaala.bsky.social · 21/06/2025
Flattered and shocked for our paper to receive the #facct2025 best paper award.
1103
Will Held @williamheld.com · 17/06/2025
I've only seen Veo 3 (or any other video generation model) used to produce viral videos. The fake videos seem to successfully trick the majority of commenters and have no visible watermark or disclosure of AI use.
110
Reposted by Will Held
Brendan Nyhan @brendannyhan.bsky.social · 12/06/2025
What would you say if you saw it in another country? A senator from a coequal branch of government dragged away by security from asking a question of a Cabinet official
26479142
Reposted by Will Held
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
🚨 70 million US workers are about to face their biggest workplace transmission due to AI agents. But nobody’s asking them what they want. While AI R&D races to automate everything, we took a different approach: auditing what workers want vs. what AI can deliver across the US workforce.🧵
1227
Will Held @williamheld.com · 06/06/2025
Really cool to see theory connect to practice! We observed this phenomenon when trying to do deeper WSD cooldowns of our 8B model in the marin.community project! We Z-Lossed our way through the pain, but cool to see some stronger theory: marin.readthedocs.io/en/latest/re...
marin.community
Marin
0101
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 05/06/2025
What foreign power could do as much damage to the United States as Trump is doing to it right now? www.whitehouse.gov/presidential...
whitehouse.gov
Enhancing National Security by Addressing Risks at Harvard University
BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION Admission into the United States to attend, conduct research, or teach at our
712633
Will Held @williamheld.com · 02/06/2025
Based on current administration policies, China is about to have an influx of returning talent and a accelerated advantage in research investments. You need to be both sinophobic and irrational to expect the US to continue as the global scientific powerhouse with these policy own-goals.
https://www.nature.com/articles/d41586-020-00084-7
130
Reposted by Will Held
Kate Starbird @katestarbird.bsky.social · 25/05/2025
"“From time-to-time instances will arise in which the society, or segments of it, threaten the very mission of the university & its values... In such a crisis, it becomes the obligation of the university as an institution to oppose such measures & actively to defend its interests and its values.”
427872
Reposted by Will Held
David Hall @dlwh.bsky.social · 19/05/2025
Super excited Marin is finally out! Come see what we've been building! Code/platform for training fully reproducible models end-to-end, from data to evals. Plus a new high quality 8B base model. Percy did a good job explaining it on the other place. marin.community x.com/percyliang/s...
x.com
Percy Liang on X: "What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3" / X
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3
1196
Will Held @williamheld.com · 19/05/2025
How much faster would the science of large-scale AI advance if we could open-source the *process* of building a frontier model? Not just the final models/code/data, but also negative results, toy experiments, and even spontaneous discussions. That's what we're trying @ marin.community
194
Will Held @williamheld.com · 15/05/2025
It feels worth conference organizers running a study to see if this significantly impacts reviewer scores. I hope things like this are placebos, but if not we need to seriously consider whether existing peer-review processes for big ML conferences are providing value.
040
Will Held @williamheld.com · 07/05/2025
Introducing CAVA: The Comprehensive Assessment for Voice Assistants A new benchmark for evaluating the capabilities required for speech-in-speech-out voice assistants! - Latency - Instruction following - Function calling - Tone awareness - Turn taking - Audio Safety TalkArena.org/cava
talkarena.org
Comprehensive Assessment for Voice Assistants
CAVA is a new benchmark for assessing how well Large Audio Models support voice assistant capabilities.
101
Reposted by Will Held
Myra Cheng @myra.bsky.social · 02/05/2025
How does the public conceptualize AI? Rather than self-reported measures, we use metaphors to understand the nuance and complexity of people’s mental models. In our #FAccT2025 paper, we analyzed 12,000 metaphors collected over 12 months to track shifts in public perceptions.
34914
Reposted by Will Held
Naomi Saphra @nsaphra.bsky.social · 26/04/2025
I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.
nsaphra.net
The AI Researcher's Guide to a Non-Boring Bluesky Feed | Naomi Saphra
How to migrate to bsky without a boring feed.
2335994
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 21/04/2025
Worth noting that a number of universities have now sued over withheld and canceled grants, but no university has yet sued over the arrest, detention, and threatened deportation of its foreign students. www.nytimes.com/2025/04/19/o...
nytimes.com
Opinion | Our Foreign Students Are Terrified, and They’re Right to Be
The immigration crackdown has come to America’s campuses.
182002627
Reposted by Will Held
Ethan Zuckerman @ethanz.bsky.social · 20/04/2025
Mahmoud Khalil writes movingly about what his detention by ICE means for America: www.washingtonpost.com/opinions/202...
washingtonpost.com
Opinion | Mahmoud Khalil: What does my detention by ICE say about America?
A democracy for some is no democracy at all.
04416
Reposted by Will Held
Prasad Jallepalli, MD, PhD @prasad.bsky.social · 14/04/2025
aside: a stunning comment from David Baker, UW professor who won the Nobel Prize in 2024. Now 15 lab members are looking for positions overseas. “There’s so many amazing people who want to come in, & we can’t take them. The Nobel Prize was just a little blip. But things have gotten quite bleak.”
372253938
Reposted by Will Held
Sean Brodrick @seanbrodrick.bsky.social · 14/04/2025
Financial Times: "Since 1990, America has lost over 5 million manufacturing jobs. In that time, it has gained 11.8 million roles in professional and business services, and 3.3 million in transportation and logistical activities, linked to multinational supply chains." #EconSky
620682
Will Held @williamheld.com · 10/04/2025
The Model Context Protocol is cool because it gives external developers a way to add meaningful functionality on top of LLM platforms. To limit test this, I made a "Realtime Voice" MCP using free STT, VAD, and TTS systems. The result is a janky, but makes me me excited about the ecosystem to come!
141
Reposted by Will Held
ashley fairbanks @ziibiing.com · 05/04/2025
people being in the streets means something. never let your cynicism convince you otherwise.
89164332825
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 31/03/2025
The Trump administration’s roundup of students who protested Israel’s bombardment of Gaza marks an astonishing, radical break with what one might justifiably think of as the central American idea. I wrote about it for @theguardian.com. www.theguardian.com/commentisfre...
7347141
Reposted by Will Held
Joshua Weitz @joshuasweitz.bsky.social · 28/03/2025
Working with an interdisciplinary team, we have developed a website to communicate how the White House's proposed cuts to health research would cause losses of $16B and 68,500 jobs. Find out how your community may be impacted. Explore more at SCIMaP: scienceimpacts.org a 🧵
US map via scienceimpacts.org visualization of economic loss due to IDC cuts to 15% as part of Feb 7, 2025 executive order, with shading denoting intensity of cuts.
19664613511
Reposted by Will Held
Michael Hobbes @michaelhobbes.bsky.social · 29/03/2025
We have had nearly two decades of panic about Free Speech on Campus and not a single case, not even the ones they made up, were as bad as what's happening now
168176376349
Reposted by Will Held
Michael Bernstein @mbernst.bsky.social · 18/03/2025
Step 1) Install the #chi2025 module to your Claude/ChatGPT: knollapp.com/add/ZlRKvCmB... Step 2) Ask the LLM, "Given my interests, what are some CHI 2025 papers I should check out?" (If the model doesn't already know your interests, you might need to state them.)
172
Reposted by Will Held
Axios @axios.com · 17/03/2025
Exclusive: Navajo Code Talkers disappear from military websites after Trump DEI order
axios.com
Navajo Code Talkers get "DEI" label as military info disappears under Trump order
Also vanishing: The Civil War regiment from "Glory," women pilots from WWII and a Medal of Honor winner.
39640822472
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 09/03/2025
Arresting and threatening to deport students because of their participation in political protest is the kind of action one ordinarily associates with the world’s most repressive regimes. It’s genuinely shocking that this appears to be what’s going on right here. 1/
363058737
Reposted by Will Held
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
As always, we open source everything. Even our nicely made website: egonormia.org Please check out the leaderboard, the blog (w/Bibtex support), the code, data, as well as a data viewer.
egonormia.org
EgoNormia: A Benchmark for Visual Frontier Models' Normative Reasoning
A large scale video dataset and a benchmark for evaluating frontier models' understanding of physical social norms through videos.
131
Reposted by Will Held
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
We are getting closer to have agents operating in the real physical world. However, can we trust frontier models to make embodied decisions 🎮 aligned with human norms 👩‍⚖️ ? With EgoNormia, a 1.8k ego-centric video 🥽 QA benchmark, we show that this is surprisingly challenging!
1239
Reposted by Will Held
Valentin Hofmann @valentinhofmann.bsky.social · 31/01/2025
Great to see the International AI Safety Report highlight research on dialect prejudice, including our work on covert racism in LLMs! www.nature.com/articles/s41...
161
Reposted by Will Held
Kyle Mahowald @kmahowald.bsky.social · 29/01/2025
Now most *urgently*: we review the history of these models. A straight line can be traced to modern AI from basic science. Not in engineering but in the cognitive science of language. Much of it funded by NSF, whose funding has now been paused. www.goldengooseaward.org/01awardees/pdp
goldengooseaward.org
How We Think: Brain-Inspired Models of Human Cognition Contribute to the Foundations of Today’s Artificial Intelligence — The Golden Goose Award
AWARDEES : Geoffrey Hinton, James L. McClelland, David E. Rumelhart FEDERAL FUNDING AGENCIES: Department of Defense, National Institutes of Health, National Science Foundation Decades before ar...
074
Will Held @williamheld.com · 22/01/2025
Balancing data across domains is key to training the best generalist LLMs! In my summer work on the Meta Llama team, we introduce UtiliMax and MEDU, new methods to estimate data utility and optimize data mixes efficiently. HF Blog: huggingface.co/blog/WillHel... ArXiv: arxiv.org/abs/2501.11747
160
Reposted by Will Held
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)
12210
Reposted by Will Held
Dan Jurafsky @jurafsky.bsky.social · 12/01/2025
Happy New Year everyone! Jim and I just put up our January 2025 release of Speech and Language Processing! Check it out here: web.stanford.edu/~jurafsky/sl...
web.stanford.edu
Speech and Language Processing
Speech and Language Processing
115250
Reposted by Will Held
Nizar Habash @nyhabash.bsky.social · 09/01/2025
NYU Abu Dhabi Opening: Computer Science Professor in Artificial Intelligence, Machine Learning, T/TT - Open Rank! Deadline Feb 28, 2025! apply.interfolio.com/161449 #AI #ML #MachineLearning #NLP #NLProc #AIjobs #MLjobs #BigData #DataScience #AIcareers #DataJobs #JobSearch
apply.interfolio.com
Apply - Interfolio {{$ctrl.$state.data.pageTitle}} - Apply - Interfolio
074
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 19/12/2024
Doctors Without Borders: What our medical teams have witnessed is "consistent with the descriptions provided by an increasing number of legal experts and organizations concluding that genocide is taking place in Gaza." www.doctorswithoutborders.org/latest/gaza-...
doctorswithoutborders.org
Gaza death trap: MSF report exposes Israel’s campaign of total destruction
Amid continued attacks, siege, and blockade, Israel is destroying conditions of life in Gaza.
25927
Will Held @williamheld.com · 17/12/2024
Update: Gemini 2.0 Flash now supported in Talk Arena! Come try the new Gemini and determine how strong it is at Speech & Audio compared to DiVA Llama 3, Qwen 2 Audio, and GPT 4o Advanced Voice at talkarena.org
Artificial Intelligence/Tech/Google
Google launched Gemini 2.0, its new AI model for practically everything
012
Reposted by Will Held
Michael Ryan @michaelryan207.bsky.social · 10/12/2024
Introducing Talk Arena🔊: the platform for interactively evaluating Large Audio Models! With the release of speech AI models like GPT4o, Gemini, Qwen-Audio, etc. which is the best? Cast your votes and help us decide🔥
132
Will Held @williamheld.com · 10/12/2024
With an increasing number of Large *Audio* Models 🔊, which one do users like the most? Introducing talkarena.org — an open platform where users speak to LAMs and receive text responses. Through open interaction, we focus on rankings based on user preferences rather than static benchmarks. 🧵 (1/5)
Talk Arena: Interactive Evaluation of Large Audio Models
3318
Will Held @williamheld.com · 08/12/2024
New Moving Sofa Problem proof dropped!! arxiv.org/abs/2411.19826
120