Sign in

Benjamin Warner

@benjaminwarner.dev
355 followers 164 following 41 posts

Research at sophont.med, previously answer.ai Vaccines save lives.

PostsRepliesMedia
Reposted by Benjamin Warner
Ted Underwood @tedunderwood.com · 08/08/2026
initially seems good, but the more you use a tool like this, the more you lose your own ability to forecast cylones
20717113
Benjamin Warner @benjaminwarner.dev · 29/03/2026
and serve with folded in lora layers in bf16 as it skips the quantization aware training step to requant those layers
100
Benjamin Warner @benjaminwarner.dev · 29/03/2026
Thinking Machines (and others) found that RL on low rank LoRA can match full-finetuning RL thinkingmachines.ai/blog/lora/, without Blackwell GPUs it would be easiest to upcast the mxfp4 MoE layers to bf16 to train
110
Benjamin Warner @benjaminwarner.dev · 27/10/2025
Some personal news: I've joined sophont.med to help build the next generation of open medical foundation models. We've relaunched medarc.ai, our open science research community. Join us if you want to help advance open medical AI. And we are hiring.
131
Reposted by Benjamin Warner
mr. TIM @timkellogg.me · 13/09/2025
counterpoint: GPT-5 does this, it says it doesn’t know rather than hallucinate, the world hasn’t fallen apart
3204
Reposted by Benjamin Warner
Tom Aarsen @tomaarsen.com · 09/09/2025
ModernBERT goes MULTILINGUAL! One of the most requested models I've seen, @jhuclsp.bsky.social has trained state-of-the-art massively multilingual encoders using the ModernBERT architecture: mmBERT. Stronger than an existing models at their sizes, while also much faster! Details in 🧵
1146
Benjamin Warner @benjaminwarner.dev · 06/09/2025
ChatGPT has been the best technical search engine since o4-mini. Thinking Mini still makes for a good faster search if you don’t need the extra reasoning ability.
000
Benjamin Warner @benjaminwarner.dev · 24/08/2025
GPT 5 Thinking (the smartest one) ignored the low quality sources and only cited the high quality and reliable sources.
000
Benjamin Warner @benjaminwarner.dev · 24/08/2025
A modern example: When attempting to trick GPT 5 + search with a question on the health benefits of raw milk, GPT 5 Fast (the less smart one) started out by citing the raw milk institute before eventually concluding there aren’t any benefits and citing high quality sources.
100
Benjamin Warner @benjaminwarner.dev · 24/08/2025
Good LLMs do know and/or can reason about these things. Small, cheap, and often free LLMs are the models which cannot. Remember the glue on pizza Reddit post that the subpar Google AI cited uncritically? Bing’s then integration of GPT 3.5 recognized the Reddit post as sarcasm.
100
Reposted by Benjamin Warner
Sung Kim @sungkim.bsky.social · 24/08/2025
Writing Speed-of-Light Flash Attention for 5090 in CUDA C++ by Thien Tran He walkthrough how he learned to implement Flash Attention for 5090 in CUDA C++. The main objective is to learn writing attention in CUDA C++,
1133
Reposted by Benjamin Warner
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/08/2025
Microsoft made a useful LLM copilot tool that could summarize text in spreadsheets. They provided clear instructions about how to use it and not to use it. In response, journalists are now mocking them for doing exactly the right thing and showing how to use and not use the tools.
1111120
Benjamin Warner @benjaminwarner.dev · 18/07/2025
Reports of AI eating entry level jobs are greatly exaggerated. My guess is current and near-future LLMs are more likely to increase the demand for programmers, not decrease demand (Jevons Paradox).
110
Benjamin Warner @benjaminwarner.dev · 20/02/2025
There isn't a canonical version, but there are retrieval models from GTE and Nomic which might work for your task. GTE: huggingface.co/Alibaba-NLP/... Nomic: huggingface.co/nomic-ai/mod...
010
Benjamin Warner @benjaminwarner.dev · 10/02/2025
For more details, including our simple training method, see Benjamin Clavié's twitter announcement, our model, blog post, and paper. Twitter: x.com/bclavie/stat... Model: huggingface.co/answerdotai/... Blog: www.answer.ai/posts/2025-0... Paper: arxiv.org/abs/2502.03793
010
Benjamin Warner @benjaminwarner.dev · 10/02/2025
Can all encoders be instruction-tuned? Can we replicate ModernBERT's results with an older model like RoBERTa or peer model like GTE-en-MLM? No. And it's not close.
220
Benjamin Warner @benjaminwarner.dev · 10/02/2025
When we finetune ModernBERT-Large-Instruct on task specific datasets, the generative MLM head is better or nearly equal to standard classification heads.
100
Benjamin Warner @benjaminwarner.dev · 10/02/2025
After instruction tuning on Flan, ModernBERT-Large-Instruct outperforms similarly sized LLMs on MMLU & MMLU-Pro, and achieves ~90 percent of Llama 3.2 1B's performance with ~65 percent fewer parameters.
110
Benjamin Warner @benjaminwarner.dev · 10/02/2025
With @bclavie.bsky.social and @ncoop57.bsky.social, we tried to answer two questions: - Can an instruction-tuned ModernBERT zero-shot tasks using the MLM-head? - Could we then fine-tune instruction-tuned ModernBERT to complete any task? Detailed answers: arxiv.org/abs/2502.03793
arxiv.org
It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers
While encoder-only models such as BERT and ModernBERT are ubiquitous in real-world NLP applications, their conventional reliance on task-specific classification heads can limit their applicability com...
141
Benjamin Warner @benjaminwarner.dev · 10/02/2025
One of the questions we debated while training ModernBERT was whether a modern trained encoder would unlock zero-shot reasoning using only it's generative head? Spoilers: the answer is yes.
from transformers import pipeline

model_name = "answerdotai/ModernBERT-Large-Instruct"
fill_mask = pipeline("fill-mask", model=model_name, tokenizer=model_name)

text = """You will be given a question and options. Select the right answer.
QUESTION: If (G, .) is a group such that (ab)^-1 = a^-1b^-1, for all a, b in G, then G is a/an
CHOICES:
- A: commutative semi group
- B: abelian group
- C: non-abelian group
- D: None of these
ANSWER: [unused0] [MASK]"""

results = fill_mask(text)
answer = results[0]["token_str"].strip()
print(f"Predicted answer: {answer}")  # Answer: B
3286
Reposted by Benjamin Warner
Simon Willison @simonwillison.net · 05/02/2025
o3-mini is really good at writing internal documentation - feed it a codebase, get back a detailed explanation of how specific aspects of it work simonwillison.net/2025/Feb/5/o...
simonwillison.net
o3-mini is really good at writing internal documentation
I wanted to refresh my knowledge of how the Datasette permissions system works today. I already have [extensive hand-written documentation](https://docs.datasette.io/en/latest/authentication.html) for...
618216
Reposted by Benjamin Warner
Maria Antoniak @mariaa.bsky.social · 27/01/2025
If you want to quickly catch up on all the open modeling things (DeepSeek, ModernBERT, etc.), this was a great overview, by @natolambert.bsky.social. I somehow got into an argument last week with someone who was insisting that all models are industrial blackboxes... and I wish I'd had this on hand.
interconnects.ai
The latest open artifacts (#6): Reasoning models, China's lead in open-source, and a growing multimodal space
Artifacts log 6 The open LM ecosystem yet again accelerates.
05310
Benjamin Warner @benjaminwarner.dev · 23/01/2025
You can find the models on Hugging Face here: - gte-modernbert-base: huggingface.co/Alibaba-NLP/... - gte-reranker-modernbert-base: huggingface.co/Alibaba-NLP/...
huggingface.co
Alibaba-NLP/gte-modernbert-base · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Benjamin Warner @benjaminwarner.dev · 23/01/2025
In addition to being the best retrieval model under 300M params on METB (without extra work), and top 10 for under 1B, here's a fun tidbit from Alibaba's GTE ModernBERT model card: gte-modernbert-base beats gte-qwen1.5-7b on LoCo long context retrieval with 7B less parameters.
130
Reposted by Benjamin Warner
Tom Aarsen @tomaarsen.com · 14/01/2025
The newest extremely strong embedding model based on ModernBERT-base is out: `cde-small-v2`. Both faster and stronger than its predecessor, this one tops the MTEB leaderboard for its tiny size! Details in 🧵
1317
Reposted by Benjamin Warner
Antoine Chaffin @nohtow.bsky.social · 14/01/2025
ModernBERT-embed-base is awesome because it allows to use ModernBERT-base for various tasks out-of-the-box But the large variant of ModernBERT is also awesome... So today, @lightonai.bsky.social is releasing ModernBERT-embed-large, the larger and more capable iteration of ModernBERT-embed!
1122
Benjamin Warner @benjaminwarner.dev · 10/01/2025
What's ModernBERT? It's a drop-in replacement for existing BERT models, but smarter, faster, and supports longer context. Check out our announcement post for more details: huggingface.co/blog/modernb...
huggingface.co
Finally, a Replacement for BERT: Introducing ModernBERT
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Benjamin Warner @benjaminwarner.dev · 10/01/2025
ModernBERT is officially released on Transformers v4.48.0. You no longer need to install from git to use. If you are plugging ModernBERT into an existing encoder finetuning pipeline, try increasing the learning rate. We've found that ModernBERT tends to prefer a higher LR than older models.
Transformers v4.48.0: ModernBERT, Aria, TimmWrapper, ColPali, Falcon3, Bamba, VitPose, DinoV2 w/ Registers, Emu3, Cohere v2, TextNet, DiffLlama, PixtralLarge, Moonshine
1113
Benjamin Warner @benjaminwarner.dev · 07/01/2025
*Actually, that’s good compared to the 4090’s PCIe 4 without NVLink
000
Benjamin Warner @benjaminwarner.dev · 07/01/2025
The good: 32GB The bad: $2,000 The Ugly*: PCIe 5 without NVLink
100
Reposted by Benjamin Warner
John West @johnwest.bsky.social · 01/01/2025
Via @simonwillison.net's excellent blog, I found this great quote about AI models, from @benjaminwarner.dev et al. www.answer.ai/posts/2024-1... It seems to me that AI will be most relevant in people's lives because the Honda Civic is ubiquitous, not so much because everyone is driving a Ferrari.
Basically, a frontier model like OpenAI’s O1 is like a Ferrari SF-23. It’s an obvious triumph of engineering, designed to win races, and that’s why we talk about it. But it takes a special pit crew just to change the tires and you can’t buy one for yourself. In contrast, a BERT model is like a Honda Civic. It’s also an engineering triumph, but more subtly, since it is engineered to be affordable, fuel-efficient, reliable, and extremely useful. And that’s why they’re absolutely everywhere.
121
Reposted by Benjamin Warner
Tom Aarsen @tomaarsen.com · 31/12/2024
That didn't take long! Nomic AI has finetuned the new ModernBERT-base encoder model into a strong embedding model for search, classification, clustering and more! Details in 🧵
23710
Benjamin Warner @benjaminwarner.dev · 24/12/2024
ModernBERT is a “foundation model” so you’ll either need to finetune it for entailment/NLI or wait for someone else to finetune it. I suspect it would be good at NLI once finetuned.
230
Benjamin Warner @benjaminwarner.dev · 24/12/2024
We evaluated ModernBERT on MLDR using ColBERT-style retrieval using that code. That process was smaller scale than a full ColBERT finetune, which would need additional contrastive training, likely use multiple teacher models, etc as detailed here by @bclavie.bsky.social www.answer.ai/posts/2024-0...
210
Benjamin Warner @benjaminwarner.dev · 24/12/2024
Thanks. ModernBERT is a base model. It’ll need additional contrastive pretraining to really shine as a retrieval model, but our early results in the paper look promising. Hopefully there will be multiple open source retrieval tuned models to choose from early next year, including ColBERT finetunes.
220
Benjamin Warner @benjaminwarner.dev · 22/12/2024
Thanks for the kind words. We tried to fit as much information within our page limit as possible and have a comprehensive appendix. As far as the name goes, all I’ll say is be careful not to use an overly strong code name.
030
Benjamin Warner @benjaminwarner.dev · 22/12/2024
(early results in our paper)
010
Benjamin Warner @benjaminwarner.dev · 22/12/2024
Thanks. It’ll need additional contrastive pretraining to really shine as a retrieval model, but our early results look promising. Hopefully there will be multiple open source retrieval tuned models to choose from early next year.
120
Benjamin Warner @benjaminwarner.dev · 22/12/2024
PS: BlueSky needs to make their really long account tags not count against the character limit.
150
Benjamin Warner @benjaminwarner.dev · 22/12/2024
I'm looking forward to seeing what you all will build with a modern encoder.
110
Benjamin Warner @benjaminwarner.dev · 22/12/2024
A big thanks to Iacopo Poli and @lightonai.bsky.social for sponsoring the compute to train ModernBERT, @bclavie.bsky.social for organizing the ModernBERT project, and to everyone who offered assistance and advice along the way. Also h/t to Johno Whitaker for the illustrations.
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
Thanks to my two co-leads: @nohtow.bsky.social , @bclavie.bsky.social , & the rest of our stacked author cast: @orionweller.bsky.social, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, @tomaarsen.com , @ncoop57.bsky.social , Griffin Adams, @howard.fm , & Iacopo Poli
161
Benjamin Warner @benjaminwarner.dev · 22/12/2024
For all the model design, training, and evaluation details, check out our Arxiv preprint: arxiv.org/abs/2412.13663
arxiv.org
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Encoder-only transformer models such as BERT offer a great performance-size tradeoff for retrieval and classification tasks with respect to larger decoder-only models. Despite being the workhorse of n...
140
Benjamin Warner @benjaminwarner.dev · 22/12/2024
Last, we trained ModernBERT on variety of data sources, including web docs, code, & scientific articles, for a total of 2 trillion tokens of English text & code. 1.7 trillion tokens at a short 1024 sequence length, followed by 300 billion tokens at a long 8192 sequence length.
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
Second, we carefully designed ModernBERT's architecture run to efficiently across most common GPUs. Many common older models don't consider the hardware they will run on and are slower than they should be. Not so with ModernBERT. (Full model sequence packing illustrated below)
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
How did we do it? First, we brought all the modern LLM architectural improvements to encoders, including alternating global & local attention, RoPE, and GeGLU layers, and added full model unpadding using Flash Attention for maximum performance (illustrated in the next post).
150
Benjamin Warner @benjaminwarner.dev · 22/12/2024
ModernBERT was designed from the ground up for speed and memory efficiency. ModernBERT is both faster and more memory efficient than every major encoder released since the original BERT.
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
ModernBERT-base is the first encoder to beat DeBERTaV3-base on GLUE. ModernBERT is also competitive or top scoring on single vector retrieval, ColBERT retrieval, and programming benchmarks.
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
ModernBERT is available to use today on Transformers (pip install from main). More details in our announcement post. huggingface.co/blog/modernb...
huggingface.co
Finally, a Replacement for BERT: Introducing ModernBERT
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
130
Benjamin Warner @benjaminwarner.dev · 22/12/2024
This week we released ModernBERT, the first encoder to reach SOTA on most common benchmarks across language understanding, retrieval, and code, while running twice as fast as DeBERTaV3 on short context and three times faster than NomicBERT & GTE on long context.
27415