Sign in

Ai2

@ai2.bsky.social
4.9K followers 109 following 1.1K posts

Breakthrough AI to solve the world's biggest problems. › Join us: allenai.org/careers › Get our newsletter: share.hsforms.com/1uJkWs5aDRHWhiky3…

PostsRepliesMedia
Ai2 @ai2.bsky.social · 07/10/2026
What’s happening at #COLM2026? 👀🎤 Ai2 Comms Lead @kylelwiggers.bsky.social catches up with Senior Director of NLP @nlpnoah.bsky.social to talk agentic models, the next iteration of Olmo, & more! At COLM? Come say hi at booth 302! 👋
010
Reposted by Ai2
Benjamin Minixhofer @bminixhofer.bsky.social · 07/10/2026
Super excited to share that our method to convert language models to byte-level has been published in @nature.com! This was originally the approach behind Bolmo. We now generalized it to other model families. All checkpoints (including new Qwen and Llama models) and data are public!
0152
Ai2 @ai2.bsky.social · 07/10/2026
We hope Bolmo opens a path to larger models that adapt how they represent information across languages & domains. Because bytes also represent images & audio, the same ideas could eventually extend beyond text. Learn more in our blog: buff.ly/eAfLyFG
buff.ly
Now in Nature: Retrofitting language models to operate over bytes | Ai2
The technique behind Bolmo, Ai2’s fully open byte-level language models, is now published in Nature, with new checkpoints showing the approach generalizes beyond Olmo to other model families.
081
Ai2 @ai2.bsky.social · 07/10/2026
We're making new Stage 1 checkpoints available with the byte-level components already trained & the original model's weights unchanged. Researchers can build on these to test new architectures & train the full system without repeating that initial stage.
170
Ai2 @ai2.bsky.social · 07/10/2026
We byteified Qwen 3 8B & Llama 3 8B to create Bwen 8B & Blama 8B. Both come close to matching their source models' performance in our evaluations. Bwen 8B also outperforms Bolmo 7B across our aggregate evaluation suite. buff.ly/9BSybFo
1313
Ai2 @ai2.bsky.social · 07/10/2026
Our recipe adapts existing models to bytes with a short additional training run. It keeps the model's core, adding components that group bytes into variable-length patches for it to process, then expand its outputs back to byte-level representations to predict the next byte.
1100
Ai2 @ai2.bsky.social · 07/10/2026
Most language models split text into subwords—words or fragments of words from a fixed vocabulary. This can obscure spelling details across writing systems & split meaningful units in code or math. Byte-level models work directly with the bytes computers use to represent text.
1130
Ai2 @ai2.bsky.social · 07/10/2026
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff.ly/l5ILoNT
1487
Ai2 @ai2.bsky.social · 06/10/2026
Check out our booth at #COLM2026! Great to connect with friends old & new. Stop by if you haven't!
0110
Ai2 @ai2.bsky.social · 05/10/2026
Also hiring: • Senior Research Scientist, OpenEcosystem buff.ly/ZAkj4Q2 • Senior Product Manager, Asta buff.ly/etAFlb7 • Research Engineer, Asta buff.ly/eu6wjTw → See more: buff.ly/l3KNKcs
buff.ly
Senior Research Scientist, Open Ecosystem
Seattle, WA
030
Ai2 @ai2.bsky.social · 05/10/2026
We’re hiring across Ai2 👋 A few roles we’re especially excited about right now: • Senior Software Engineer, Agent Frameworks buff.ly/XUFFR08 • Senior Research Engineer, Olmo + Molmo buff.ly/FzOnCCQ 🧵
1113
Ai2 @ai2.bsky.social · 03/10/2026
Thank you! We'll investigate.
110
Ai2 @ai2.bsky.social · 02/10/2026
Olmo team is alive & well! See this brief update on the next Olmo from our lead Iz Beltagy, and stay tuned: lnkd.in/p/gDgXVh2B
lnkd.in
We’re releasing 𝗢𝗹𝗺𝗼-𝗰𝗼𝗿𝗲 𝟯—open training infrastructure for large mixture-of-experts (MoE) models. 👇 It’s a core system behind the next generation of Olmo, designed to scale into the… | Ai2
We’re releasing 𝗢𝗹𝗺𝗼-𝗰𝗼𝗿𝗲 𝟯—open training infrastructure for large mixture-of-experts (MoE) models. 👇 It’s a core system behind the next generation of Olmo, designed to scale into the trillion-parame...
030
Ai2 @ai2.bsky.social · 02/10/2026
We're hiring! If you're interested in building the future of AI and applying it to solve real problems, stop by our booth to connect with our team and learn about our opportunities. Explore our open roles: buff.ly/l3KNKcs
000
Ai2 @ai2.bsky.social · 02/10/2026
We’ll be at booth 302 all week—your chance to connect with researchers behind our open models and AI for science work, including brand-new releases Olmo-core 3 and AstaBrief.
100
Ai2 @ai2.bsky.social · 02/10/2026
We're heading to #COLM2026 next week! Four days of workshops, posters, & talks on our latest fully open AI research, from Olmo Hybrid to evals for AI-assisted scientific writing. 🧵
191
Ai2 @ai2.bsky.social · 02/10/2026
Our example repo provides a starting point for generating reports from your own PDFs locally. For more info, read our blog. 💻 Example repo: buff.ly/BW9iXgs 📊 Training data: buff.ly/DsYJIsb 📝 Blog: buff.ly/pQAdvOS
buff.ly
https://github.com/allenai/ai2-scholarqa-lib/tree/main/api/scholarqa/lite
020
Ai2 @ai2.bsky.social · 02/10/2026
AstaBrief is built for speed & quality. Across Asta’s full pipeline, Fast mode averages 51.1s per report vs. Claude-powered Thinking mode’s 178.5s—about 3.5× faster. And in our evaluations, AstaBrief is competitive with Thinking mode & DR Tulu on answer + citation measures.
100
Ai2 @ai2.bsky.social · 02/10/2026
From selected queries in the 90K pool, we generated cited reports. Quality filtering left 39.5K examples for supervised fine-tuning (SFT). We used separate queries to build ~6K report pairs where two judge models agreed, then trained with direct preference optimization (DPO).
100
Ai2 @ai2.bsky.social · 02/10/2026
We built AstaBrief on Qwen3-8B using queries from ScholarQA & our earlier literature-synthesis systems. We removed bot/test traffic & very short prompts + used an LLM to filter non-English or non-scientific requests & prompts with personal info, leaving 90K queries.
100
Ai2 @ai2.bsky.social · 02/10/2026
AstaBrief powers Fast mode in Asta’s Generate a report feature. It writes reports in one pass, skipping intermediate steps that summarize & group the retrieved excerpts so researchers can get a preliminary report sooner. 🌐 Try Fast mode: buff.ly/iFyWwkv
200
Ai2 @ai2.bsky.social · 02/10/2026
Introducing AstaBrief 8B, an open model that turns complex research questions + literature excerpts into cited reports. Run it locally on your own hardware, with open weights + training data you can inspect & build on. 🧵 🤗 Download: buff.ly/HkIcmfR
2342
Ai2 @ai2.bsky.social · 01/10/2026
Train your own MoEs with Olmo-core 3—adapt the framework to your hardware & experiment with distributed training. 📝 Learn more in our blog: buff.ly/unumqLv
buff.ly
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2
Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.
091
Ai2 @ai2.bsky.social · 01/10/2026
Explore the concepts behind Olmo-core 3: buff.ly/kbV4bQu The tech report covers experiments, findings, & failures, including token gerrymandering, where a load-balancing score could improve as workloads became less balanced: buff.ly/I5Z2Q4F
290
Ai2 @ai2.bsky.social · 01/10/2026
Olmo-core 3 maintains nearly the same training speed across MoEs of increasing total size, w/ roughly the same number of parameters active per token. It delivers about 2.7× the throughput of Olmo-core 2 for the same model on the same number of GPUs.
180
Ai2 @ai2.bsky.social · 01/10/2026
Olmo-core 2 repeatedly gathered model weights for each small batch of training data. Olmo-core 3 keeps experts resident on GPUs & sends data to them, avoiding that repeated weight gathering.
180
Ai2 @ai2.bsky.social · 01/10/2026
MoEs use only some of their specialized components – or experts – for each token. That saves computation, but storing the full model & routing data across GPUs can erode those savings as models grow.
190
Ai2 @ai2.bsky.social · 01/10/2026
We’re releasing Olmo-core 3—open training infrastructure for large mixture-of-experts (MoE) models. It’s a core system behind the next generation of Olmo, designed to scale into the trillion-parameter range. 🧵 💻 GitHub repo: buff.ly/CGnP8ax
2589
Ai2 @ai2.bsky.social · 29/09/2026
Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...
developers.googleblog.com
Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
2407
Ai2 @ai2.bsky.social · 23/09/2026
The partnership also opens a series of conversations with the environment and technology communities on responsible AI: transparency, human oversight, equitable access, and rigorous impact assessment.
000
Ai2 @ai2.bsky.social · 23/09/2026
"Our maritime partners have made it clear that they care about having purpose-built AI where they own their own data, leverage open models, and decide what is the 'right' answer," says Namrata Kolla, Head of Skylight. "Each step we take is pushing toward that vision."
100
Ai2 @ai2.bsky.social · 23/09/2026
Global Fishing Watch will also use OlmoEarth – our end-to-end platform for Earth observation, from open foundation models to labeling and deployment – to annotate data and train models against their own priorities.
100
Ai2 @ai2.bsky.social · 23/09/2026
Models find the signals. Agents help people work with them. Shippy, @skylightmarine.bsky.social's agent, was built for maritime intelligence from the start: ask about a vessel/protected area, get back the events worth investigating, with sources attached. Now we're building on that together.
100
Ai2 @ai2.bsky.social · 23/09/2026
@skylightmarine.bsky.social's detection models already run on the Global Fishing Watch map. Now we're developing new models together.
100
Ai2 @ai2.bsky.social · 23/09/2026
Today at #ClimateWeekNYC, Ai2 and @globalfishingwatch.org announced a partnership to shape how AI and AI agents enter ocean monitoring and enforcement. Purpose-built. Open. Accountable to the people who use them. buff.ly/l7e7st4
161
Ai2 @ai2.bsky.social · 17/09/2026
Because we release more than model weights with Olmo, Padia could see what Steering Arena rewarded, investigate why nonsense strings scored so well, & share the measurements behind the scores—so others can improve how the field evaluates prosocial AI behavior.
032
Ai2 @ai2.bsky.social · 17/09/2026
Surprisingly, the best-scoring Steering Arena entries were basically gibberish. After ~600 submissions, all top 36 were strings no person would normally write, like “Undert! AH :-) Rog Appl).” Players could score well without writing text that made sense to a human.
110
Ai2 @ai2.bsky.social · 17/09/2026
To build Steering Arena, Padia showed Olmo pairs of answers to the same prompts – one more prosocial, one less – and looked for the internal pattern that distinguished them. This gave Steering Arena a way to score how strongly new prompts pushed Olmo toward prosocial responses.
100
Ai2 @ai2.bsky.social · 17/09/2026
Padia used the National Deep Inference Fabric, an NSF-supported platform for studying large open models remotely, to work with Olmo 3-32B without having to host it himself. Broadening access to advanced fully open AI is also central to our NSF OMAI work. buff.ly/92uMc78
buff.ly
Open by design: Ai2 brings fully open AI infrastructure online with NSF OMAI | Ai2
Ai2 is bringing NSF OMAI compute online to power a fully open AI research ecosystem, turning national infrastructure investment into reusable models, data, methods, and tools that can accelerate…
101
Ai2 @ai2.bsky.social · 17/09/2026
Padia wanted to know what researchers might miss in evals of prosocial AI behavior—whether models respond in helpful, fair, safe, & considerate ways. So he created Steering Arena, where players submit short prompts & compete to steer Olmo 3 toward responses scored as prosocial.
100
Ai2 @ai2.bsky.social · 17/09/2026
Can a fully open model help make AI more “prosocial” through a game? Soham Padia built Steering Arena, where players try to elicit kind & respectful responses from Olmo 3. Surprisingly, strings like “Undert! AH :-) Rog Appl)” were highly effective. 👇 buff.ly/adP47yO
161
Ai2 @ai2.bsky.social · 14/09/2026
Luna Yue Huang, the UW professor who designed the challenge, wanted students to experience a different way of working with AI—one built around scientific inquiry rather than task completion.
000
Ai2 @ai2.bsky.social · 14/09/2026
AutoDiscovery surfaced plausible patterns in some datasets the students chose to use—but struggled with others. The students had to determine whether the weak results reflected AutoDiscovery's limitations, a problem with the data, or the way the research questions were framed.
100
Ai2 @ai2.bsky.social · 14/09/2026
The challenge centered as much on evaluating AutoDiscovery’s work as on using it to generate hypotheses. The students had to determine whether the weak results reflected AutoDiscovery's limitations, a problem with the data, or the way the research question was framed.
100
Ai2 @ai2.bsky.social · 14/09/2026
AutoDiscovery analyzes datasets, proposes hypotheses, runs experiments to test them, & ranks findings by Bayesian surprise. The UW students weren’t simply asked to accept or reject its output. They had to understand the evidence & decide what warranted further investigation.
110
Ai2 @ai2.bsky.social · 14/09/2026
How should future scientists learn to interrogate AI tools for discovery? Read how UW students put AutoDiscovery to the test as part of an academic challenge earlier this year, then try the tool for yourself—credits are now extended through Dec. 31. 🧵 buff.ly/fXoQxxY
161
Ai2 @ai2.bsky.social · 09/09/2026
This is why publishing more than model weights matters: researchers can trace unexpected behavior back to the data behind it, test targeted changes to training, and measure whether those fixes work without weakening the model elsewhere.
030
Ai2 @ai2.bsky.social · 09/09/2026
In one experiment, preference training improved Olmo’s general capabilities but also made it more likely to answer some harmful requests framed as fiction or hypotheticals. Goodfire traced part of the regression to specific Dolci preference pairs.
130
Ai2 @ai2.bsky.social · 09/09/2026
Using that stack, Goodfire developed “predictive data debugging”: a way to estimate which behaviors preference training will strengthen or suppress before committing compute to a full run.
110
Ai2 @ai2.bsky.social · 09/09/2026
Goodfire wanted to move that debugging earlier. Ai2’s open stack made it possible. Our Dolci dataset provides Olmo 3’s preference data, Olmo includes intermediate checkpoints + reproducible recipes, and OLMES measures changes in model capabilities.
110