Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 08/05/2026Today we’re releasing EMO, a new mixture-of-experts (MoE) model trained so modular structure emerges directly from data without human-defined priors. EMO can use a small subset of its experts for a given task while keeping near full-model performance. 🧵 217223
Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 20/11/2025Announcing Olmo 3, a leading fully open LM suite built for reasoning, chat, & tool use, and an open model flow—not just the final weights, but the entire training journey. Best fully open 32B reasoning model & best 32B base model. 🧵 16917
Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 01/10/2025The Cancer AI Alliance (CAIA) is already prototyping Asta DataVoyager in a federated, multi-institution setup for cancer studies—keeping clinical data local and secure. Read more about CAIA here: buff.ly/ACpxLNT 131
Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 09/07/2025Introducing FlexOlmo, a new paradigm for language model training that enables the co-development of AI through data collaboration. 🧵 1156
Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 13/03/2025Announcing OLMo 2 32B: the first fully open model to beat GPT 3.5 & GPT-4o mini on a suite of popular, multi-skill benchmarks. Comparable to best open-weight models, but a fraction of training compute. When you have a good recipe, ✨ magical things happen when you scale it up! 35814
Akshita Bhagia @akshitab.bsky.social · 12/02/2025I caught myself wanting to respond similarly to Claude and then told myself that it will be wasteful inference. But now I also mentally thank it each time because what if I lose that instinct with humans.. I'm already impatient with smart speakers. 030
Reposted by Akshita BhagiaLuca Soldaini 🎀 @soldaini.net · 11/02/2025They made me do video 😬 but for a good reason! We are launching an iOS app–it runs OLMoE locally 📱 We're gonna see more on-device AI in 2025, and wanted to offer a simple way to prototype with it App: apps.apple.com/us/app/ai2-o... Code: github.com/allenai/OLMo... Blog: allenai.org/blog/olmoe-app 54913
Reposted by Akshita BhagiaKyle Lo @ COLM2026 @kylelo.bsky.social · 03/01/2025kicking off 2025 with our OLMo 2 tech report while payin homage to the sequelest of sequels 🫡 🚗 2 OLMo 2 Furious 🔥 is everythin we learned since OLMo 1, with deep dives into: 🚖 stable pretrain recipe 🚔 lr anneal 🤝 data curricula 🤝 soups 🚘 tulu post-train recipe 🚜 compute infra setup 👇🧵 26917
Reposted by Akshita BhagiaJiacheng Liu @liujch1998.bsky.social · 09/12/2024Want to predict the task performance of LMs before pretraining them? We develop task scaling laws and model ladders, which predict the accuracy on individual tasks by OLMo 2 7B & 13B models within 2 points of absolute error. The cost is 1% of the compute used to pretrain them. 23314
Reposted by Akshita BhagiaAi2 @ai2.bsky.social · 26/11/2024Meet OLMo 2, the best fully open language model to date, including a family of 7B and 13B models trained up to 5T tokens. OLMo 2 outperforms other fully open models and competes with open-weight models like Llama 3.1 8B — As always, we released our data, code, recipes and more 🎁 515235
Reposted by Akshita BhagiaLuca Soldaini 🎀 @soldaini.net · 01/02/2024release day release day 🥳 OLMo 1b +7b out today and 65b soon... OLMo accelerates the study of LMs. We release *everything*, from toolkit for creating data (Dolma) to train/inf code blog blog.allenai.org/olmo-open-la... olmo paper allenai.org/olmo/olmo-pa... dolma paper allenai.org/olmo/dolma-p...blog.allenai.orgOLMo: Open Language ModelA State-Of-The-Art, Truly Open LLM and Framework 12814
Reposted by Akshita BhagiaIan Magnusson @ianmagnusson.bsky.social · 20/12/2023LMs are used to process text from many topics, styles, dialects, etc., but how well do they do? 📈 Evaluating perplexity on just one corpus like C4 doesn't tell the whole story 📉 ✨📃✨ We introduce Paloma, a benchmark of 585 domains from NY Times to r/depression on Reddit. 1177