Will Held @williamheld.com · 11/05/2026The development process and findings above are described in more detail in a blog I wrote: openathena.ai/blog/delphi/ There's interactive figures for all of the above, plus some additional results on estimating the cost of overtraining and on random seed variance at this scale! 060
Will Held @williamheld.com · 11/05/2026Beyond pretraining loss, we find that downstream tasks also scale predictably if you do things carefully! The important trick is that you need to use soft metrics, e.g. NLL or BPB, to predict performance, following Rylan Schaeffer's "Emergence is a Mirage" work. 130
Will Held @williamheld.com · 11/05/2026At this point, we pre-registered the new launch on Github and later on Twitter (sorry Bsky). Stressful to pre-register but important to avoid survivorship bias on scaling law extrapolation findings!! Thankfully, the new recipe scaled predictably for held-out PPL loss. 120
Will Held @williamheld.com · 11/05/2026Before risking another run blowing up, I sanity checked that this new recipe seemed reasonably close to the empirical optimal across two hidden dimensions, four batch sizes, and three token counts. Complete(d)P is pretty impressive and generalized well to our setting! 120
Will Held @williamheld.com · 11/05/2026Last fall, I naively thought that I had a decent scaling recipe built on top of muP and a few other works on hyperparameter transfer! But when we scaled to larger (8B+) scales with compute support from the Google TPU Research Cloud, things quickly went wrong! 130
Will Held @williamheld.com · 11/05/2026Predictable scaling laws are a foundational tool even for compute-rich organizations. They let you iterate faster and make decisions with less compute. What scaling laws don't tell you, though, is how to train the model at each compute scale! 140
Will Held @williamheld.com · 11/05/2026To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error. Getting there took some work 🧵 2389
Will Held @williamheld.com · 28/07/2025The SALT Lab is at #ACL2025 with our genius leader @diyiyang.bsky.social. Come see work from @yanzhe.bsky.social, @dorazhao.bsky.social @oshaikh.bsky.social, @michaelryan207.bsky.social, and myself at any of the talks and posters below! 030
Will Held @williamheld.com · 10/07/2025It seems (at a minimum) like they post-trained on the virulently racist content from this thread. Musk framed this as a request for training data... and the top post is eugenics. Seems unlikely to be coincidence that the post uses the same phrasing as the prompt they later removed... 020
Will Held @williamheld.com · 03/07/2025In our most similar setting to the original work (130M model), we don't see AdamC's benefits but - We use a smaller WD (0.01) identified from sweeps v.s. what is used in the paper (0.05). - We only train to Chnichilla optimal (2B tokens) whereas the original paper was at 200B. 100
Will Held @williamheld.com · 03/07/2025We see the same pattern at 300m and 500m! Remember, everything else in these experiments is held constant by Levanter & Marin (data order, model init. etc.) Experiment files here: github.com/marin-commun... 100
Will Held @williamheld.com · 03/07/2025As a side note, Kaiyue Wen found that weight decay also causes slower loss decrease at the start of training in wandb.ai/marin-commun... Similar to the end of training, this is likely because LR warmup also impacts the LR/WD ratio. AdamC seems to mitigate this too. 100
Will Held @williamheld.com · 03/07/2025TL;DR: 3/4 of our scales we find the AdamC results to reproduce out of the box! When compared to AdamW with all other factors held constant, AdamC mitigates the gradient ascent at the end of training and leads to an overall lower loss (-0.04)! 100
Will Held @williamheld.com · 02/06/2025Based on current administration policies, China is about to have an influx of returning talent and a accelerated advantage in research investments. You need to be both sinophobic and irrational to expect the US to continue as the global scientific powerhouse with these policy own-goals. 130
Will Held @williamheld.com · 19/05/2025Marin repurposes GitHub, which has been successful for open-source *software*, for AI: 1. Preregister an experiment as a GitHub issue 2. Submit a PR, which implements the experiment in code 3. PR is reviewed by experts in the community 4. Watch the execution of the experiment live! 100
Will Held @williamheld.com · 19/05/2025How much faster would the science of large-scale AI advance if we could open-source the *process* of building a frontier model? Not just the final models/code/data, but also negative results, toy experiments, and even spontaneous discussions. That's what we're trying @ marin.community 194
Will Held @williamheld.com · 15/05/2025It feels worth conference organizers running a study to see if this significantly impacts reviewer scores. I hope things like this are placebos, but if not we need to seriously consider whether existing peer-review processes for big ML conferences are providing value. 040
Will Held @williamheld.com · 07/05/2025Results? We tested ✅ GPT-4o (end-to-end audio) ✅ GPT pipeline (transcribe + text + TTS) ✅ Gemini 2.0 Flash ✅ Gemini 2.5 Pro We find GPT-4o shines on latency & tone while Gemini 2.5 leads in safety & prompt adherence. No model wins everything. (3/5) 100
Will Held @williamheld.com · 10/04/2025The Model Context Protocol is cool because it gives external developers a way to add meaningful functionality on top of LLM platforms. To limit test this, I made a "Realtime Voice" MCP using free STT, VAD, and TTS systems. The result is a janky, but makes me me excited about the ecosystem to come! 141
Will Held @williamheld.com · 17/12/2024Update: Gemini 2.0 Flash now supported in Talk Arena! Come try the new Gemini and determine how strong it is at Speech & Audio compared to DiVA Llama 3, Qwen 2 Audio, and GPT 4o Advanced Voice at talkarena.org 012
Will Held @williamheld.com · 10/12/2024Testing models on 18 commonly used static evaluation benchmarks, we find that none produce the same rankings as our interactive user evaluation. This suggests common interaction areas might be missing in existing static benchmarks used for Large Audio Models! (3/5) 120
Will Held @williamheld.com · 10/12/2024Before releasing Talk Arena, we collected votes from over 350 paid participants on Prolific comparing five popular models’ text responses. The initial standings show 🏅DiVA, 🥈GPT4o, 🥉Gemini, 4️⃣ Qwen2 Audio, 5️⃣ Typhoon. (2/5) 120
Will Held @williamheld.com · 10/12/2024With an increasing number of Large *Audio* Models 🔊, which one do users like the most? Introducing talkarena.org — an open platform where users speak to LAMs and receive text responses. Through open interaction, we focus on rankings based on user preferences rather than static benchmarks. 🧵 (1/5) 3318
Will Held @williamheld.com · 06/12/2024RIP to the first product I got to do real "Big Data" work on 🫡🫡🫡 100
Will Held @williamheld.com · 05/12/2024I believe you are thinking of the Llama 3.2 models which I think are not covered in the paper, but are pruned then refined with distillation! huggingface.co/meta-llama/L... 120
Will Held @williamheld.com · 17/11/2024Trying to break into a closet, realizing he's been caught, and feigning innocence!! 050
Will Held @williamheld.com · 14/11/2024I'll be at the Google Theory and Practice of Foundation Models Workshop today and tomorrow! FOMO for EMNLP, but excited to chat more casually at a smaller non-archival workshop 😅 I am presenting at the Lightning Talks tomorrow at 1:30 PM on our Distilled Voice Assistant model if you're around! 081
Will Held @williamheld.com · 07/11/2024Of course! How you do sampling and packing is one of those things that matters a lot in practice, but often gets shoved to appendices because it's not exciting. For example, this non-trivial solution from DeepMind which isn't referenced in the main text. 140
Will Held @williamheld.com · 07/11/2024More recent works take the "don't attend across documents" even further and manually mask out the attention across documents. This gives you the compute-efficiency without the "it's weird to attend across documents at all" aspect. 3110
Will Held @williamheld.com · 07/11/2024Yes! Token packing has been the standard since RoBERTa. Excerpt below! The intuition is that the model quickly learns to not attend across [SEP] boundaries and packing avoids "wasting" compute on padding tokens required to make the variable batch size consistent. 2182