Sign in

Ai2

@ai2.bsky.social
4.9K followers 109 following 1.1K posts

Breakthrough AI to solve the world's biggest problems. › Join us: allenai.org/careers › Get our newsletter: share.hsforms.com/1uJkWs5aDRHWhiky3…

PostsRepliesMedia
Pinned
Ai2 @ai2.bsky.social · 27/08/2026
AI’s biggest role in science may not be answering questions. It may be helping scientists find which questions are worth asking. At Providence Swedish, AutoDiscovery surfaced an unexpected immune signal in a heavily studied cancer dataset—and follow-up research confirmed it. 🧵 buff.ly/LTacUQl
1153
Ai2 @ai2.bsky.social · 29/09/2026
Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...
developers.googleblog.com
Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
1405
Ai2 @ai2.bsky.social · 23/09/2026
Today at #ClimateWeekNYC, Ai2 and @globalfishingwatch.org announced a partnership to shape how AI and AI agents enter ocean monitoring and enforcement. Purpose-built. Open. Accountable to the people who use them. buff.ly/l7e7st4
161
Ai2 @ai2.bsky.social · 17/09/2026
Can a fully open model help make AI more “prosocial” through a game? Soham Padia built Steering Arena, where players try to elicit kind & respectful responses from Olmo 3. Surprisingly, strings like “Undert! AH :-) Rog Appl)” were highly effective. 👇 buff.ly/adP47yO
150
Ai2 @ai2.bsky.social · 14/09/2026
How should future scientists learn to interrogate AI tools for discovery? Read how UW students put AutoDiscovery to the test as part of an academic challenge earlier this year, then try the tool for yourself—credits are now extended through Dec. 31. 🧵 buff.ly/fXoQxxY
161
Ai2 @ai2.bsky.social · 09/09/2026
Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. Goodfire AI used Ai2’s open post-training stack to predict how a full training run would change responses to different prompts. 🧵 buff.ly/RuJxLM1
180
Ai2 @ai2.bsky.social · 08/09/2026
We're at #ECCV2026 with papers & talks across the conference. Come say hello and learn about our latest research!
051
Ai2 @ai2.bsky.social · 04/09/2026
ML emulators mimic climate processes faster than traditional models. Next is coupling atmosphere & ocean emulators so their predictions feed into each other as the simulation runs. With the E3SM team, we built a system that does that: SamudrACE-E3SMv3. 🧵 buff.ly/9ShSTCM
161
Reposted by Ai2
Maarten Sap @maartensap.bsky.social · 02/09/2026
Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions! See 🧵
0112
Ai2 @ai2.bsky.social · 01/09/2026
Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 buff.ly/bTcvqJf
1213
Ai2 @ai2.bsky.social · 01/09/2026
At an event on August 27, we brought together AI researchers, scientists, & medical practitioners to explore what AI needs to do better to meaningfully advance science. Five ideas kept coming up. 🧵
110
Ai2 @ai2.bsky.social · 26/08/2026
A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data. 🧵 buff.ly/fvLSIru
1161
Ai2 @ai2.bsky.social · 21/08/2026
What kinds of training data shape different model capabilities? A Georgia Tech team used our fully open model flow to trace Olmo’s performance on social/general reasoning and social-science/STEM knowledge tests back to the types of text it trained on. 🧵 buff.ly/cQEzbUn
1110
Ai2 @ai2.bsky.social · 18/08/2026
A model can sound like it knows a drug—even when it doesn’t. Researchers at @utaustin.bsky.social, @northeasternu.bsky.social, & @mdanderson.bsky.social found LLMs often lack drug-specific knowledge and instead lean on morphology. They used Olmo to trace why. 🧵 buff.ly/PuyfnCq
1133
Ai2 @ai2.bsky.social · 07/08/2026
Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵
2516
Ai2 @ai2.bsky.social · 06/08/2026
We're expanding our partnership with @hf.co to accelerate open science. Our storage on the Hub is roughly tripling to ~2 petabytes, & our downloads now run at high speed—even for our largest datasets & multi-checkpoint models. 🧵
1281
Ai2 @ai2.bsky.social · 05/08/2026
📸 140+ people joined us during #SeattleTechWeek to learn how AI gets built at Ai2. Research lead Iz Beltagy walked through what it takes to train a truly open LLM—from large-scale training runs to eval before release. Thanks to everyone who came + asked thoughtful questions. 🤝
140
Ai2 @ai2.bsky.social · 31/07/2026
When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. @tuhinchakr.bsky.social's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. 🧵
1131
Ai2 @ai2.bsky.social · 28/07/2026
The organizations best positioned to use Earth-observation models – those working in conservation, food security, & disaster response – often can't run these models at the scale they need. That's an infrastructure problem. So we built the OlmoEarth Platform to solve it. 🧵
160
Ai2 @ai2.bsky.social · 24/07/2026
As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.
191
Ai2 @ai2.bsky.social · 24/07/2026
As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.
160
Ai2 @ai2.bsky.social · 24/07/2026
As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.
110
Ai2 @ai2.bsky.social · 24/07/2026
As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem. We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.
120
Ai2 @ai2.bsky.social · 23/07/2026
Curious how AI gets built at Ai2? For #SeattleTechWeek, we're opening our Northlake office for a panel with the researchers building our open models. They'll talk through the deep technical work behind them, with time after to chat about what you're working on & network. 👇
111
Ai2 @ai2.bsky.social · 21/07/2026
Two updates to Asta, our ecosystem of AI agents for science: a one-click handoff from AutoDiscovery to Asta’s data analysis tools, & paper search that evaluates its own results + searches again when they fall short. 🧵
1102
Ai2 @ai2.bsky.social · 17/07/2026
What does it actually take to build cutting-edge AI systems? On July 30 during #SeattleTechWeek, the researchers behind Ai2's open models sit down to talk through the deep technical work behind them. 🧵
290
Ai2 @ai2.bsky.social · 16/07/2026
This week we hosted researchers from @cgiar.org, @servirglobal.bsky.social, NASA Harvest, & @msftresearch.bsky.social at our office to explore OlmoEarth for food security and natural resource management challenges. The CGIAR teams brought their own use cases and put OlmoEarth to work. 🧵
130
Ai2 @ai2.bsky.social · 16/07/2026
It was an honor to bring OlmoEarth to the Nature Positive Summit (@npinitiative.bsky.social) this week in Kumamoto, Japan, with our partners The Group on Earth Observations (@earthobservations.bsky.social). 🧵
180
Ai2 @ai2.bsky.social · 15/07/2026
We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes. Here's what they told us. 🧵
295
Ai2 @ai2.bsky.social · 13/07/2026
Our @skylightmarine.bsky.social team is rolling out Shippy, an AI agent for the people protecting our ocean in real time. A wrong answer at these stakes can send a patrol vessel miles off course. Here's the architecture we built so analysts can trust it. 🧵
142
Ai2 @ai2.bsky.social · 10/07/2026
olmOCR 2 is now in the Ai2 Playground—our home for our fully open text, video, and image understanding models. 🧵
1102
Ai2 @ai2.bsky.social · 10/07/2026
OlmoEarth turns Earth data into insights—in hours, not years. But it's only as powerful as the people using it. So, we brought them into one room. 🧵
1100
Ai2 @ai2.bsky.social · 09/07/2026
"Embeddings only get you so far… if you want the next level in performance, I think fine-tuning is the way to go." Ai2 research scientist @pjreddie.bsky.social explains how our partners customize OlmoEarth – our open-source Earth-observation models – to map crops, wildfire risk, & more. 👇
130
Ai2 @ai2.bsky.social · 08/07/2026
What can you build with a fully open robotics model in a weekend? 🤖 Robotics engineer @0xbinh.bsky.social used MolmoAct 2, our open vision-language-action model, in the voice-controlled robot that won @southparkcommons.bsky.social's AI hackathon. Watch our interview with him ↓ 🎥
0120
Ai2 @ai2.bsky.social · 06/07/2026
We're at #ICML2026 with papers & talks across the conference. Come say hello and learn about our latest research!
170
Ai2 @ai2.bsky.social · 02/07/2026
The Danish Foundation Models (DFM) project is adapting our modular FlexOlmo architecture into a lighter-weight system that runs on commodity hardware—putting collaborative model building within reach of smaller research groups & organizations. 🧵
1329
Ai2 @ai2.bsky.social · 01/07/2026
We're at #ACL2026 with papers & talks across the conference. Come say hello and learn about our latest research!
070
Ai2 @ai2.bsky.social · 29/06/2026
AI image generators don't "draw"—they follow a compass: the score function, which points toward more probable images. The same compass drives Bayesian sampling and plasma physics. We built DiScoFormer to estimate the score far better when data gets complex. 🧵
2141
Ai2 @ai2.bsky.social · 29/06/2026
Today we're releasing OlmoEarth v1.2, the latest in our family of open foundation models for Earth observation. 🌍 We've switched to rotary positional embeddings (RoPE), which reduces artifacts in the embeddings & gives a small performance boost. 🧵
2123
Ai2 @ai2.bsky.social · 25/06/2026
Hybrid (transformer–RNN) models are fast becoming a serious alternative to the transformer, but a big question remains: how do they process tokens differently & how does this impact performance? We compared our transformer (Olmo 3) & hybrid (Olmo Hybrid) models to find out. 🧵
1266
Ai2 @ai2.bsky.social · 18/06/2026
Learn how @thinkaisquared.bsky.social & Domyn used Olmo, our family of fully open language models, to build their own models for regulated industries like finance, healthcare, & the public sector. 🧵
121
Ai2 @ai2.bsky.social · 17/06/2026
We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
3277
Ai2 @ai2.bsky.social · 12/06/2026
Building an LLM means evaluating it over & over as it changes. Tweak a hyperparameter or scale the model up, & every new checkpoint sends you back through the same benchmarking loop. We're releasing olmo-eval, a workbench built for this kind of iterative model development. 🧵
183
Ai2 @ai2.bsky.social · 11/06/2026
LLMs are no longer created w/ human data alone. They rely on other models to generate & filter data, evaluate outputs, & guide dev work. So what is a modern LLM built on? Olmo 3 → 89 model + 183 dataset dependencies; Nemotron 3 → 273 + 560 We made ModSleuth to trace this. 🧵
15511
Reposted by Ai2
kylelwiggers.bsky.social @kylelwiggers.bsky.social · 10/06/2026
𝗔𝗖𝗘𝟮𝗦-𝗦𝗛𝗶𝗘𝗟𝗗+, our new climate emulator that learns to separate the effects of sea surface temperature & CO2, is now on @hf.co—check it out → huggingface.co/allenai/ACE2...
huggingface.co
allenai/ACE2S-SHiELD-plus · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
092
Ai2 @ai2.bsky.social · 09/06/2026
Today we're introducing ACE2S-SHiELD+, a climate emulator that learns to separate the effects of sea surface temperature & CO2. It accurately handles scenarios where previous versions of our ACE family of climate emulators produced inaccurate results. 🧵
1140
Ai2 @ai2.bsky.social · 03/06/2026
We're at #CVPR2026 with papers & talks across the conference. Come say hello and learn about our latest research!
051
Ai2 @ai2.bsky.social · 01/06/2026
We're extending AutoDiscovery early access through July 31. New accounts start with 500 Hypothesis Credits (one credit = one hypothesis), & any credits you already have will still work. 🧵
140
Ai2 @ai2.bsky.social · 28/05/2026
MolmoAct 2 artifacts have been downloaded 400K+ times in under 1 month. Today we're opening up the full code & training data: everything you need to fine-tune or build on our fully open robotics foundation model. 🧵
191
Ai2 @ai2.bsky.social · 22/05/2026
Most models are only evaluated on a fraction of the benchmarks out there. ArtifactLinker, our new system, predicts which ones would set a new state-of-the-art on benchmarks hosted on @hf.co, then runs the evaluation to verify. 🧵
1141