Sign in

Will Held

@williamheld.com
2.2K followers 457 following 116 posts

Modeling Linguistic Variation to expand ownership of NLP tools Views my own, but affiliations that might influence them: ML PhD Student under Prof. Diyi Yang 2x RS Intern🦙 Pretraining Alum NYU Abu Dhabi Burqueño he/him

PostsRepliesMedia
Reposted by Will Held
Open Athena @openathena.ai · 25/06/2026
In a new blog, Russell Power explains how the Marin team nearly doubled its sustained TPU usage by creating a custom global scheduler: Iris. Iris searches every region where Marin has compute, places each job wherever capacity appears, and moves data along as needed. 🔗 openathena.ai/blog/cluster...
1102
Will Held @williamheld.com · 11/05/2026
Ah, sorry! Models: huggingface.co/collections/... Data Scripts: github.com/marin-commun... Recipe: github.com/marin-commun...
huggingface.co
Delphi - a marin-community Collection
Marin's first open scaling suite. 88 base models, 3e18 → 1e23 FLOPs. https://openathena.ai/blog/delphi
000
Will Held @williamheld.com · 11/05/2026
The development process and findings above are described in more detail in a blog I wrote: openathena.ai/blog/delphi/ There's interactive figures for all of the above, plus some additional results on estimating the cost of overtraining and on random seed variance at this scale!
060
Will Held @williamheld.com · 11/05/2026
To help make Delphi useful for open scaling science, we are releasing: Model Checkpoints: huggingface.co/collections/... Code to reproduce the data mixture: github.com/marin-commun... Recipe definition: github.com/marin-commun... If there are other things you want, lmk!
huggingface.co
Hugging Face – The AI community building the future.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
230
Will Held @williamheld.com · 11/05/2026
Beyond pretraining loss, we find that downstream tasks also scale predictably if you do things carefully! The important trick is that you need to use soft metrics, e.g. NLL or BPB, to predict performance, following Rylan Schaeffer's "Emergence is a Mirage" work.
130
Will Held @williamheld.com · 11/05/2026
At this point, we pre-registered the new launch on Github and later on Twitter (sorry Bsky). Stressful to pre-register but important to avoid survivorship bias on scaling law extrapolation findings!! Thankfully, the new recipe scaled predictably for held-out PPL loss.
120
Will Held @williamheld.com · 11/05/2026
Before risking another run blowing up, I sanity checked that this new recipe seemed reasonably close to the empirical optimal across two hidden dimensions, four batch sizes, and three token counts. Complete(d)P is pretty impressive and generalized well to our setting!
120
Will Held @williamheld.com · 11/05/2026
There were two red flags in the runs from that initial sweep: suspiciously large weight norms and learning rates! For the first, I switched to Kaiyue Wen's Hyperball optimizers. For the second, I found the answer in Apple's Complete(d)P paper: arxiv.org/abs/2512.22382
arxiv.org
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
Hyperparameter tuning can dramatically impact training stability and final performance of large-scale models. Recent works on neural network parameterisations, such as $μ$P, have enabled transfer of o...
140
Will Held @williamheld.com · 11/05/2026
Last fall, I naively thought that I had a decent scaling recipe built on top of muP and a few other works on hyperparameter transfer! But when we scaled to larger (8B+) scales with compute support from the Google TPU Research Cloud, things quickly went wrong!
130
Will Held @williamheld.com · 11/05/2026
To do that, you need a scaling recipe. The idea is simple: start with a small reference model and tune it heavily. Then generalize that recipe across scales by reparameterizing the config as a function of compute.
150
Will Held @williamheld.com · 11/05/2026
Predictable scaling laws are a foundational tool even for compute-rich organizations. They let you iterate faster and make decisions with less compute. What scaling laws don't tell you, though, is how to train the model at each compute scale!
140
Will Held @williamheld.com · 11/05/2026
Inspired by open scaling suites such as Pythia from @eleutherai.bsky.social, Delphi releases 3 things: - a recipe for choosing what to train at each compute budget - a suite of models trained from that recipe - a scaling law fit on smaller runs to predict larger ones
140
Will Held @williamheld.com · 11/05/2026
To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error. Getting there took some work 🧵
2389
Will Held @williamheld.com · 29/10/2025
OpenAI addresses this with a backend classifier "to detect if the GPT‑4o output is using a voice that’s different from our approved list". But that isn't possible for open-source models, so would be great to at least partially mitigate this by baking in this gating mechanism.
000
Will Held @williamheld.com · 29/10/2025
While many systems exist explicitly designed for voice cloning, some systems can do it unintentionally due to ICL. For example, in the original 4o card """ During testing, we also observed rare instances where the model would unintentionally generate an output emulating the user’s voice """
100
Will Held @williamheld.com · 29/10/2025
Super interested to what degree this interaction can be fine-tuned into models in a non-reversible fashion! Voice cloning is unfortunately a capability which inherently shows up in pretrained audio models. It would be great to be able to largely limit the capability at the level of model weights!
110
Reposted by Will Held
Dan Jurafsky @jurafsky.bsky.social · 24/08/2025
Now that school is starting for lots of folks, it's time for a new release of Speech and Language Processing! Jim and I added all sorts of material for the August 2025 release! With slides to match! Check it out here: web.stanford.edu/~jurafsky/sl...
web.stanford.edu
Speech and Language Processing
Speech and Language Processing
315358
Will Held @williamheld.com · 11/08/2025
"GPT-5 shows scaling laws are coming to an end"
060
Reposted by Will Held
George Pearkes @peark.es · 06/08/2025
We’ve discovered a literal miracle with almost unlimited potential and it’s being scrapped for *no reason whatsoever*. This isn’t even nihilism, it’s outright worship of death and human suffering.
47103173294
Will Held @williamheld.com · 06/08/2025
Really great pointer from Hao Zhang on the other site in relation to GPT OSS use of attention sinks. If I were to guess, the attention sink is what allows them to omit QK-Norm which has become otherwise standard. www.evanmiller.org/attention-is...
evanmiller.org
Attention Is Off By One
Let’s fix these pesky Transformer outliers using Softmax One and QuietAttention.
010
Will Held @williamheld.com · 28/07/2025
The SALT Lab is at #ACL2025 with our genius leader @diyiyang.bsky.social. Come see work from @yanzhe.bsky.social, @dorazhao.bsky.social @oshaikh.bsky.social, @michaelryan207.bsky.social, and myself at any of the talks and posters below!
Alt Text:

Conference schedule for July 28th (Monday) and July 29th (Tuesday), listing talk titles, locations, times, and authors:

July 28th, Monday:

1. Attacking Vision-Language Computer Agents via Pop-ups
Location: Hall 4/5, Time: 11:00–12:30
Authors: Yanzhe Zhang, Tao Yu, Diyi Yang


2. SPHERE: An Evaluation Card for Human-AI Systems
Location: Hall 4/5, Time: 18:00–19:30
Authors: Dora Zhao*, Qianou Ma*, Xinran Zhao, Chenglei Si, Chenyang Yang, Ryan Louie, Ehud Reiter, Diyi Yang*, Tongshuang Wu*
(asterisk denotes equal contribution)



July 29th, Tuesday:

1. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
Location: Hall 4/5, Time: 10:30–12:00
Authors: Michael J Ryan, Omar Shaikh, Aditri Bhagirath, Daniel Frees, William Barr Held, Diyi Yang


2. Distilling an End-to-End Voice Assistant Without Instruction Training Data
Location: Room 1.61, Time: 14:12 (Second Talk)
Authors: William Barr Held, Yanzhe Zhang, Weiyan Shi, Minzhi Li, Michael J Ryan, Diyi Yang


3. Mind the Gap: Static and Interactive Evaluations of Large Audio Models
Location: Room 1.61 (implied), follows previous talk
Authors: Minzhi Li*, William Barr Held*, Michael J Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, Diyi Yang
(asterisk denotes equal contribution)


4. EgoNormia: Benchmarking Physical Social Norm Understanding
Location: Hall 4/5, Time: 16:00–17:30
Authors: MohammadHossein Rezaei*, Yicheng Fu*, Phil Cuvin*, Caleb Ziems, Yanzhe Zhang, Hao Zhu, Diyi Yang
(asterisk denotes equal contribution)
030
Will Held @williamheld.com · 28/07/2025
Paper: aclanthology.org/2025.acl-lon...
000
Will Held @williamheld.com · 28/07/2025
I'm in Vienna for #ACL2025! My work is all presented tomorrow, but today you'll find me today at the poster session from 11-12:30 evangelizing my labmate Yanzhe Zhang's work on his behalf. If you're interested in the risks traditional pop-up attacks present for AI agents, come chat!
140
Will Held @williamheld.com · 10/07/2025
It seems (at a minimum) like they post-trained on the virulently racist content from this thread. Musk framed this as a request for training data... and the top post is eugenics. Seems unlikely to be coincidence that the post uses the same phrasing as the prompt they later removed...
020
Will Held @williamheld.com · 03/07/2025
Btw, all of this is very nice for something that was a quick 15 line addition to Levanter. github.com/stanford-crf...
000
Will Held @williamheld.com · 03/07/2025
Have an optimizer you want to prove works better than AdamC/Muon/etc? Submit a speedrun to Marin! marin.readthedocs.io/en/latest/tu... For PRs with promising results, we're lucky to be able to help test at scale on compute generously provided by the TPU Research Cloud!
marin.readthedocs.io
Adding an Optimizer for Speedrun - Marin Documentation
Documentation for the Marin project
100
Will Held @williamheld.com · 03/07/2025
In our most similar setting to the original work (130M model), we don't see AdamC's benefits but - We use a smaller WD (0.01) identified from sweeps v.s. what is used in the paper (0.05). - We only train to Chnichilla optimal (2B tokens) whereas the original paper was at 200B.
100
Will Held @williamheld.com · 03/07/2025
We see the same pattern at 300m and 500m! Remember, everything else in these experiments is held constant by Levanter & Marin (data order, model init. etc.) Experiment files here: github.com/marin-commun...
100
Will Held @williamheld.com · 03/07/2025
As a side note, Kaiyue Wen found that weight decay also causes slower loss decrease at the start of training in wandb.ai/marin-commun... Similar to the end of training, this is likely because LR warmup also impacts the LR/WD ratio. AdamC seems to mitigate this too.
100
Will Held @williamheld.com · 03/07/2025
TL;DR: 3/4 of our scales we find the AdamC results to reproduce out of the box! When compared to AdamW with all other factors held constant, AdamC mitigates the gradient ascent at the end of training and leads to an overall lower loss (-0.04)!
100
Will Held @williamheld.com · 03/07/2025
A while ago I mentioned that for marin.community project, this gradient increase led to problematic loss ascent which we patched with Z-loss. I was curious, does AdamC just work? So over the weekend, I ran 4 experiments—130M to 1.4B params—all at ~compute-optimal token counts...🧵
marin.community
Marin
141
Will Held @williamheld.com · 03/07/2025
kyutai.org/next/unmute has built in turn-detection on the ASR and full I/O streaming for the TTS. Solves the latency issues that I think are 90% of why people use end-to-end speech models in the first place! From the details, you can @kyutai-labs.bsky.social is focused on real-world utility.
unmute.sh
Unmute by Kyutai
Make LLMs listen and speak.
010
Reposted by Will Held
Haley L. @haleyhaala.bsky.social · 21/06/2025
Flattered and shocked for our paper to receive the #facct2025 best paper award.
1103
Will Held @williamheld.com · 17/06/2025
As far as I can tell, the models aren't good enough right now that they can replace VFX at any high quality commercial scale. They are exactly good enough to generate fake viral videos for ad revenue on TikTok/Instagram & spread misinformation. Is there any serious argument for their safe release??
000
Will Held @williamheld.com · 17/06/2025
I don't really see an argument for releasing such models with photorealistic generation capabilities. What valid & frequent business use case is there for photorealistic video & voice generation like Veo 3 offers?
110
Will Held @williamheld.com · 17/06/2025
I've only seen Veo 3 (or any other video generation model) used to produce viral videos. The fake videos seem to successfully trick the majority of commenters and have no visible watermark or disclosure of AI use.
110
Reposted by Will Held
Brendan Nyhan @brendannyhan.bsky.social · 12/06/2025
What would you say if you saw it in another country? A senator from a coequal branch of government dragged away by security from asking a question of a Cabinet official
26479142
Reposted by Will Held
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
🚨 70 million US workers are about to face their biggest workplace transmission due to AI agents. But nobody’s asking them what they want. While AI R&D races to automate everything, we took a different approach: auditing what workers want vs. what AI can deliver across the US workforce.🧵
1227
Will Held @williamheld.com · 06/06/2025
Really cool to see theory connect to practice! We observed this phenomenon when trying to do deeper WSD cooldowns of our 8B model in the marin.community project! We Z-Lossed our way through the pain, but cool to see some stronger theory: marin.readthedocs.io/en/latest/re...
marin.community
Marin
0101
Will Held @williamheld.com · 05/06/2025
Now, I wouldn't do research on LLMs if I thought that was true in the long term! But I think it's reasonable for skeptics to question whether advances in inference efficiency, hardware efficiency, and even core energy infrastructure will happen soon enough for current companies to capitalize.
000
Will Held @williamheld.com · 05/06/2025
The underlying assumption being that they can (a la Uber/Lyft) eventually increase prices once the core customers are fundamentally reliant on AI. The real question then is "what is demand once you start charging the true unit costs?". Personally, I found this article sobering but well reasoned.
wheresyoured.at
The Subprime AI Crisis
None of what I write in this newsletter is about sowing doubt or "hating," but a sober evaluation of where we are today and where we may end up on the current path. I believe that the artificial intel...
110
Will Held @williamheld.com · 05/06/2025
Without knowing all the model details or with transparent financials, it's hard to say but I would naively suspect most AI companies are in the red both on a cost per query basis (for API services) and on a cost per user basis (for subscription services).
100
Will Held @williamheld.com · 05/06/2025
I haven't seen people mocking the revenue forecasts, but I agree with your take w.r.t. demand. The bigger question is whether demand is the constraint? Unlike standard software or even manufacturing businesses, I'm not sure the economies of scale look great if you factor in cost per query.
120
Reposted by Will Held
Jameel Jaffer @jameeljaffer.bsky.social · 05/06/2025
What foreign power could do as much damage to the United States as Trump is doing to it right now? www.whitehouse.gov/presidential...
whitehouse.gov
Enhancing National Security by Addressing Risks at Harvard University
BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION Admission into the United States to attend, conduct research, or teach at our
712633
Will Held @williamheld.com · 02/06/2025
Based on current administration policies, China is about to have an influx of returning talent and a accelerated advantage in research investments. You need to be both sinophobic and irrational to expect the US to continue as the global scientific powerhouse with these policy own-goals.
https://www.nature.com/articles/d41586-020-00084-7
130
Will Held @williamheld.com · 29/05/2025
Given that they published the same work in both the ICLR workshop and ACL... I am skeptical of the claim that "The current version of Zochi represents a substantial advancement over our earlier systems that published workshop papers at ICLR 2025" 😂
110
Will Held @williamheld.com · 29/05/2025
Looks like they simultaneously submitted the same paper to an ICLR workshop: openreview.net/forum?id=rDC...
openreview.net
Siege: Multi-Turn Jailbreaking of Large Language Models with Tree...
We introduce Siege, a multi-turn adversarial framework that models the gradual erosion of Large Language Model (LLM) safety through a tree search perspective. Unlike single-turn jailbreaks that...
120
Reposted by Will Held
Kate Starbird @katestarbird.bsky.social · 25/05/2025
"“From time-to-time instances will arise in which the society, or segments of it, threaten the very mission of the university & its values... In such a crisis, it becomes the obligation of the university as an institution to oppose such measures & actively to defend its interests and its values.”
427872
Reposted by Will Held
David Hall @dlwh.bsky.social · 19/05/2025
Super excited Marin is finally out! Come see what we've been building! Code/platform for training fully reproducible models end-to-end, from data to evals. Plus a new high quality 8B base model. Percy did a good job explaining it on the other place. marin.community x.com/percyliang/s...
x.com
Percy Liang on X: "What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3" / X
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3
1196
Will Held @williamheld.com · 19/05/2025
Learn more about the project in Percy's blog post: marin.community/blog/2025/05... And about the Models we are releasing in @dlwh.bsky.social's training retro: marin.readthedocs.io/en/latest/re...
marin.community
Introducing Marin: An Open Lab for Building Foundation Models
Open-source software is a success story: It powers the world’s digital infrastructure. It allows anyone in the world to contribute based on merit. It leads to greater innovation, collaboration, and se...
001