Sign in

Mechanical Dirk

@mechanicaldirk.bsky.social
522 followers 241 following 63 posts

Training big models at @ai2.bsky.social.

PostsRepliesMedia
Mechanical Dirk @mechanicaldirk.bsky.social · 24/01/2026
Ownership is a scam invented by big thing to sell more stuff.
000
Mechanical Dirk @mechanicaldirk.bsky.social · 18/01/2026
Severed hand?!? Just go to an embassy and do it there!
100
Mechanical Dirk @mechanicaldirk.bsky.social · 10/01/2026
Phone, as well as email, are a way in which anyone in the world can get a piece of my attention whenever they want. As we are figuring out increasingly effective ways of monetizing attention, exploitation of this resource expands.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 17/12/2025
Oof
010
Mechanical Dirk @mechanicaldirk.bsky.social · 08/12/2025
Be sure to make that point when the funding for the other stuff disappears all the same.
000
Reposted by Mechanical Dirk
Nathan Lambert @natolambert.bsky.social · 20/11/2025
Happy Olmo day to all who celebrate. Sorry to all who delayed releases today to get out of our way. We're hiring.
0322
Reposted by Mechanical Dirk
Ai2 @ai2.bsky.social · 20/11/2025
Announcing Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use, and an open model flow—not just the final weights, but the entire training journey. Best fully open 32B reasoning model & best 32B base model. 🧵
16917
Reposted by Mechanical Dirk
Kyle Lo @ COLM2026 @kylelo.bsky.social · 20/11/2025
we released Olmo 3! lot of exciting stuff but wanna focus on: 🐟Olmo 3 32B Base, the best fully-open base model to-date, near Qwen 2.5 & Gemma 3 on diverse evals 🐠Olmo 3 32B Think, first fully-open reasoning model approaching Qwen 3 levels 🐡12 training datasets corresp to different staged training
1417
Reposted by Mechanical Dirk
Nathan Lambert @natolambert.bsky.social · 14/11/2025
I'm excited to announce my RLHF Book is now in pre-order for the @manning.com Early Access Program (MEAP), and for this milestone it's 50% off. Excited to land in print in early 2026! Lots of improvements coming soon. Thanks for the support! hubs.la/Q03Tc37Q0
4474
Mechanical Dirk @mechanicaldirk.bsky.social · 14/11/2025
@ananyahjha93.bsky.social knows the pain. We once got reviews to a paper that said "Please do further experiments [which would cost $2M]", and _also_ a review that said "This work is too expensive to be relevant to anyone in the field.". In the same paper!
060
Mechanical Dirk @mechanicaldirk.bsky.social · 12/11/2025
Incredible work by Apple's UX department, enabling three different corner radii at the same time 🙈
020
Reposted by Mechanical Dirk
Daniel Buschek @dbuschek.bsky.social · 23/10/2025
While reviewing for #CHI2026, I've noticed four new writing issues in #HCI papers, likely due to an increased use of #LLMs / #AI. I describe them here - and how to fix them: dbuschek.medium.com/when-llms-wr...
dbuschek.medium.com
When LLMs Write Our Papers
Four writing issues I notice as a reviewer — and how to fix them
2285
Mechanical Dirk @mechanicaldirk.bsky.social · 08/10/2025
You presumably have something you want to run, and it doesn't run. Or it runs on CPU. Spin up a Claude Code console and tell it about the problem. Tell it the command line that produces the wrong result. It'll start suggesting ways of fixing your environment, down to modifying installed drivers.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 08/10/2025
Let Claude Code sort out your environment.
200
Mechanical Dirk @mechanicaldirk.bsky.social · 24/09/2025
Are humans allowed to attend?
010
Mechanical Dirk @mechanicaldirk.bsky.social · 24/09/2025
What field/area is like this now?
200
Reposted by Mechanical Dirk
Ai2 @ai2.bsky.social · 18/08/2025
We’re releasing early pre-training checkpoints for OLMo-2-1B to help study how LLM capabilities emerge. They’re fine-grained snapshots intended for analysis, reproduction, and comparison. 🧵
1276
Mechanical Dirk @mechanicaldirk.bsky.social · 18/08/2025
Mein Dreijähriger: "Ich will den Lerns Geschichte Podcast hören!" Was ist denn "Lerns Geschichte"? Zwei Minuten später im Radio: "Lernen's a bissel @geschichte.fm, dann ..." 😲
000
Mechanical Dirk @mechanicaldirk.bsky.social · 28/07/2025
Almost all post-training is "dusting off capable base models"
110
Mechanical Dirk @mechanicaldirk.bsky.social · 18/07/2025
Unverified second hand information: In the US, all fish has to be flash-frozen before being served raw. In Canada, it does not.
000
Mechanical Dirk @mechanicaldirk.bsky.social · 24/06/2025
I think for the moment we're competing on a different axis. They do quite well on impact per GPU hour. We do well on impact per person hour.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 08/06/2025
In ML, you can get surprisingly far without ever looking at your training data, and yet you'll always be limited. Thus, in ML, "look at the data" means, "Don't just stir the pot of linear algebra, find out what's really happening."
040
Mechanical Dirk @mechanicaldirk.bsky.social · 03/06/2025
This project is a perfect model of an OLMo contribution. Well scoped, practical, sound theoretical underpinnings, and @lambdaviking.bsky.social submitted the paper 24h before the deadline 😍. It's integrated into the OLMo trainer here: github.com/allenai/OLMo...
020
Mechanical Dirk @mechanicaldirk.bsky.social · 13/05/2025
Meanwhile, OLMo is now the citation for QK norm, which we definitely didn't invent? You win some, you lose some.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 01/05/2025
Finally, OLMo 1B. This is the most commonly requested OLMo feature l, and it's finally here.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 23/04/2025
After ICML, I decided all conferences should be in Vienna from now on.
000
Reposted by Mechanical Dirk
Jacob Morrison @jacobcares.bsky.social · 23/04/2025
I'm in Singapore for @iclr-conf.bsky.social ! Come check out our spotlight paper on the environmental impact of training OLMo (link in next tweet) during the Saturday morning poster session from 10-12:30 -- happy to chat about this or anything else! DMs should be open, email works too
1105
Mechanical Dirk @mechanicaldirk.bsky.social · 18/04/2025
Came across arxiv.org/pdf/2504.05058 today. What a cool example of work you can do when LLM training data is open!
arxiv.org
170
Reposted by Mechanical Dirk
Ai2 @ai2.bsky.social · 15/04/2025
Ever wonder how LLM developers choose their pretraining data? It’s not guesswork— all AI labs create small-scale models as experiments, but the models and their data are rarely shared. DataDecide opens up the process: 1,050 models, 30k checkpoints, 25 datasets & 10 benchmarks 🧵
Plot shows the relationship between compute used to predict a ranking of datasets and how accurately that ranking reflects performance at the target (1B) scale of models pretrained from scratch on those datasets.
15211
Reposted by Mechanical Dirk
Jiacheng Liu @liujch1998.bsky.social · 09/04/2025
Today we're unveiling OLMoTrace, a tool that enables everyone to understand the outputs of LLMs by connecting to their training data. We do this on unprecedented scale and in real time: finding matching text between model outputs and 4 trillion training tokens within seconds. ✨
1415
Mechanical Dirk @mechanicaldirk.bsky.social · 07/04/2025
The fact that my Bsky feed is all tariffs and none Llama 4 means the platform is pretty much cooked for research purposes.
110
Reposted by Mechanical Dirk
Alisa Liu @alisawuffles.bsky.social · 21/03/2025
We created SuperBPE🚀, a *superword* tokenizer that includes tokens spanning multiple words. When pretraining at 8B scale, SuperBPE models consistently outperform the BPE baseline on 30 downstream tasks (+8% MMLU), while also being 27% more efficient at inference time.🧵
Segmentation of the sentence "By the way, I am a fan of the Milky Way" under BPE and SuperBPE.
38316
Mechanical Dirk @mechanicaldirk.bsky.social · 16/03/2025
It costs $90k. The $1000 are just a down payment.
110
Mechanical Dirk @mechanicaldirk.bsky.social · 13/03/2025
Error bars! @hails.computer will be so proud!
020
Mechanical Dirk @mechanicaldirk.bsky.social · 13/03/2025
Biggest one yet! Best one yet! Plus, some fun training stories at the bottom of the blog post (allenai.org/blog/olmo2-32B).
allenai.org
OLMo 2 32B: First fully open model to outperform GPT 3.5 and GPT 4o mini | Ai2
Introducing OLMo 2 32B, the most capable and largest model in the OLMo 2 family.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 10/03/2025
When I played Civilization, I always named my religion "PDF", so I can convert cities to PDF.
020
Reposted by Mechanical Dirk
Ai2 @ai2.bsky.social · 25/02/2025
Introducing olmOCR, our open-source tool to extract clean plain text from PDFs! Built for scale, olmOCR handles many document types with high throughput. Run it on your own GPU for free—at over 3000 token/s, equivalent to $190 per million pages, or 1/32 the cost of GPT-4o!
38213
Mechanical Dirk @mechanicaldirk.bsky.social · 19/02/2025
That seems like a completely normal sleep schedule for a 3 month old. Source: My kid.
010
Reposted by Mechanical Dirk
Ai2 @ai2.bsky.social · 11/02/2025
We took our most efficient model and made an open-source iOS app📱but why? As phones get faster, more AI will happen on device. With OLMoE, researchers, developers, and users can get a feel for this future: fully private LLMs, available anytime. Learn more from @soldaini.net👇 youtu.be/rEK_FZE5rqQ
youtu.be
Ai2 OLMoE: Fully open source, running entirely on-device
YouTube video by Ai2
23014
Mechanical Dirk @mechanicaldirk.bsky.social · 10/02/2025
#humblebrag
040
Mechanical Dirk @mechanicaldirk.bsky.social · 26/01/2025
I used Thinkmate. I want to roughly pick my own specs while knowing nothing about compatibility. I don't like RGB lights everywhere. And I want the thing to be reliable. No regrets.
000
Mechanical Dirk @mechanicaldirk.bsky.social · 26/01/2025
14.8T tokens in 2.8M hours is about 1500 tokens per second. That's a very good number for 37B active parameters, but by no means unbelievable.
000
Mechanical Dirk @mechanicaldirk.bsky.social · 26/01/2025
You posted about AI.
010
Mechanical Dirk @mechanicaldirk.bsky.social · 25/01/2025
I haven't read it. But I did listen to an AI generated conversation about its contents...
020
Reposted by Mechanical Dirk
Nathan Lambert @natolambert.bsky.social · 22/01/2025
Behind the scenes with what its like to build language models and pursue (hopefully) cutting edge AI research Interviewing OLMo 2 leads: Open secrets of training language models What we have learned and are going to do next. YouTube: buff.ly/40IlSFF Podcast / notes:
buff.ly
Interviewing OLMo 2 leads: Open secrets of training language models
What we have learned and are going to do next.
1338
Mechanical Dirk @mechanicaldirk.bsky.social · 19/01/2025
In November, every post here was about NLP. Now it's all about TikTok. We're doing the Twitter speed run.
020
Mechanical Dirk @mechanicaldirk.bsky.social · 06/01/2025
A few days ago, we did finally release the OLMo 2 tech report: arxiv.org/pdf/2501.00656. There is a lot of good stuff in there, but the stability work we did over the summer makes me particularly proud.
arxiv.org
010
Reposted by Mechanical Dirk
Nathan Lambert @natolambert.bsky.social · 03/01/2025
Everyone wants open-source language models but no one wants to lift these heavy ass weights. We just released our paper "2 OLMo 2 Furious" Can't stop us in 2025. Links below.
65610
Mechanical Dirk @mechanicaldirk.bsky.social · 31/12/2024
Maybe Nvidia could have given us $699M worth of GPUs and we give them Beaker?
020
Mechanical Dirk @mechanicaldirk.bsky.social · 30/12/2024
If there was any hope that these kids would become super intelligent within a year, the money would flow.
000