Sign in

kylelwiggers.bsky.social

@kylelwiggers.bsky.social
3.6K followers 60 following 802 posts

Ai2 Comms Lead | kylew@allenai.org | Pronouns: he/him

PostsRepliesMedia
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 29/09/2026
Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...
developers.googleblog.com
Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
1405
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 17/09/2026
it's pretty amazing when you see work like this and realize how little we know about llms still
010
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 14/09/2026
really cool to see how students used autodiscovery to figure out the tool's strengths and where it can be improved
010
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 09/09/2026
such a cool use case—thanks to goodfire ai for showcasing the power of our fully open model stack!
010
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 08/09/2026
We're at #ECCV2026 with papers & talks across the conference. Come say hello and learn about our latest research!
051
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 04/09/2026
ML emulators mimic climate processes faster than traditional models. Next is coupling atmosphere & ocean emulators so their predictions feed into each other as the simulation runs. With the E3SM team, we built a system that does that: SamudrACE-E3SMv3. 🧵 buff.ly/9ShSTCM
161
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 01/09/2026
a story as old as time but... some new evidence that llm benchmarks aren't measuring what you think they are
010
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 27/08/2026
our ai autodiscovery tool contributed to a promising cancer finding—and it's exactly the kind of rigorous scientific work that we hope to enable more of. bravo to all the researcher teams involved
030
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 26/08/2026
great to see our open tools used to build better models for the world
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 21/08/2026
really cool to see which kinds of training text drove particular olmo skills - the benefit of full oppeness 💪
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 18/08/2026
fascinating work showing how LLMs understand – or don't! – drugs and what they do, made possible by olmo
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 07/08/2026
are chatbots today good tutors? turns out not so much because they're overly helpful and do a lot of the hard work for you—cool new research from us
010
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 06/08/2026
huggingface has been an amazing partner, and we're looking forward to collaborating more closely as we train highly capable new models 👀
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 05/08/2026
thx everyone for coming! had great convos with incredibly inspiring ppl, was a privilege!
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 01/08/2026
really cool use of our open source infini-gram engine to study llm regurgitation
050
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 28/07/2026
learn about the truly impressive infra behind our olmoearth platform 👇
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 24/07/2026
fully open models are important, and we wholeheartedly support their development to advance science for everyone 🚀
001
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 21/07/2026
our incredible completely free service gets even better 🚀
000
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 17/07/2026
i'll be in seattle for this—come say hi!
000
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 15/07/2026
We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵
295
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 13/07/2026
Our @skylightmarine.bsky.social team is rolling out Shippy, an AI agent for the people protecting our ocean in real time. A wrong answer at these stakes can send a patrol vessel miles off course. Here's the architecture we built so analysts can trust it. 🧵
142
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 10/07/2026
olmOCR 2 is now in the Ai2 Playground—our home for our fully open text, video, and image understanding models. 🧵
1102
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 02/07/2026
The Danish Foundation Models (DFM) project is adapting our modular FlexOlmo architecture into a lighter-weight system that runs on commodity hardware—putting collaborative model building within reach of smaller research groups & organizations. 🧵
1329
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 29/06/2026
Today we're releasing OlmoEarth v1.2, the latest in our family of open foundation models for Earth observation. 🌍 We've switched to rotary positional embeddings (RoPE), which reduces artifacts in the embeddings & gives a small performance boost. 🧵
2123
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 25/06/2026
Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
1266
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 18/06/2026
Learn how @thinkaisquared.bsky.social & Domyn used Olmo, our family of fully open language models, to build their own models for regulated industries like finance, healthcare, & the public sector. 🧵
121
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 17/06/2026
We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
3277
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 12/06/2026
Building an LLM means evaluating it over & over as it changes. Tweak a hyperparameter or scale the model up, & every new checkpoint sends you back through the same benchmarking loop. We're releasing olmo-eval, a workbench built for this kind of iterative model development. 🧵
183
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 11/06/2026
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵
15511
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 10/06/2026
𝗔𝗖𝗘𝟮𝗦-𝗦𝗛𝗶𝗘𝗟𝗗+, our new climate emulator that learns to separate the effects of sea surface temperature & CO2, is now on @hf.co—check it out → huggingface.co/allenai/ACE2...
huggingface.co
allenai/ACE2S-SHiELD-plus · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
092
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 19/05/2026
Today we’re releasing OlmoEarth v1.1. It’s 3x cheaper to run than v1 while delivering the same state-of-the-art performance—and fully open. 🧵
1577
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 21/05/2026
Brendan Works is a product manager focused on paratransit services in Seattle. See how he built PointCheck, a website accessibility checker powered by our open Molmo, MolmoWeb, & Olmo 3 models. 👇
1152
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 12/05/2026
Now available in AstaLabs in limited research preview: MyScholarQA, a personalized version of ScholarQA for scientific deep research. ScholarQA helps synthesize evidence from 12M+ open-access papers. MyScholarQA adds user profiles to tailor that synthesis to you. 🧵
1161
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 11/05/2026
Artificial Analysis relies on our IFBench eval to test how closely models follow user prompts. Most evals in their Intelligence Index saturate within months. IFBench hasn't because it measures what others miss—and what frontier models still struggle with. 🧵
1142
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 08/05/2026
Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵
217223
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 05/05/2026
Robotics models often struggle outside controlled environments. Ours is built to work in real ones. Today we're launching MolmoAct 2, which can assist with a host of chores & lab tasks, plus the MolmoAct 2-Bimanual YAM dataset—the largest open robotics dataset of its kind. 🧵
26918
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 01/05/2026
Today we published a Q&A with Interim CEO Peter Clark on what’s next for Ai2, from advancing truly open AI systems to applying AI in areas like scientific discovery & the planet. 👇
161
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 30/04/2026
New AstaBench results show frontier models making progress on scientific research, but the benchmark remains far from solved. Claude Opus 4.7 leads overall at 58.0%, while GPT-5.5 comes within 5.1 points at less than half the measured cost per problem. 🧵
2132
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 20/04/2026
Last year, we introduced FlexOlmo, a novel way to train parts of a model independently then combine them later. BAR builds on that idea for a harder problem: how to keep improving a model without having to retrain each time. 🧵
13710
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 10/04/2026
You can now train, adapt, and eval web agents on your own tasks. We're releasing the full MolmoWeb codebase—the training code, eval harness, annotation tooling, synthetic data pipeline, & client-side code for our demo. 🧵
1184
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 07/04/2026
Today we're releasing WildDet3D—an open model for monocular 3D object detection in the wild. It works with text, clicks, or 2D boxes, and on zero-shot evals it nearly doubles the best prior scores. 🧵
1256
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 27/03/2026
MolmoBot, our open robotic manipulation suite trained entirely in simulation, now has code, training data, a data generation pipeline, & evals all available. This puts our robotics models within reach of any research lab—no extensive real-world data collection required. 🧵
3267
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 24/03/2026
Today we're releasing MolmoWeb, an open source agent that can navigate + complete tasks in a browser on your behalf. Built on Molmo 2 in 4B & 8B sizes, it sets a new open-weight SOTA across four major web-agent benchmarks & even surpasses agents built on proprietary models. 🧵
2468
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 05/03/2026
Introducing Olmo Hybrid, a 7B fully open model combining transformer and linear RNN layers. It decisively outperforms Olmo 3 7B across evals, w/ new theory & scaling experiments explaining why. 🧵
1315
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 02/03/2026
In just a few weeks, researchers used AutoDiscovery to generate 20K+ hypotheses across oncology, climate science, marine ecology, entomology, cybersecurity, music cognition, social sciences, & more. Now we're extending access for three more months—and refreshing credits. 👇
1123
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 27/02/2026
We analyzed 250K+ queries & 430K+ clickstream interactions from Asta, our AI-powered research assistant—and today we're releasing the full dataset. How do researchers actually use AI science tools? Here's what we found. 🧵
1236
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 25/02/2026
Can AI predict what scientists will do next—not just one piece, but the whole research process? PreScience is our new model eval for forecasting how science unfolds end-to-end, from how research teams form to a paper's eventual impact. Built with UChicago, supported by NSF.
153
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 18/02/2026
We've released a Chrome extension for Asta—a faster way to go from finding a paper to asking questions about it while you read. 🧵
1135
Reposted by @kylelwiggers.bsky.social
Ai2 @ai2.bsky.social · 13/02/2026
Data mixing – determining how much web text, code, math, etc., you need for LM development – is a first-order lever on model quality. Introducing Olmix: a framework for configuring mixing methods at the start of dev & efficiently updating as data changes throughout. 🧵
1236