Kyle Lo @ COLM2026 @kylelo.bsky.social · 06/10/2026Olmo Hybrid: arxiv.org/abs/2604.03444 Cracks in foundation: arxiv.org/abs/2608.10296 000
Kyle Lo @ COLM2026 @kylelo.bsky.social · 06/10/2026I'm at #colm2026, Mon-Thurs Supporting two papers from @ai2.bsky.social days 🐟Olmo Hybrid, 7B hybrid model 🐠Cracks in the Foundation, long context recipes Hoping to have fun chats w folks about scaling LM data & evals, ☺️ unbiased on pretrain, posttrain, all the 🚂s 1160
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026How2Everything blog: allenai.org/blog/how2eve... Extended paper: arxiv.org/abs/2602.088...allenai.orgHow2Everything: Mining the web to evaluate and improve LLMs on real-world procedures | Ai2How2Everything is an open framework for evaluating and improving how well LLMs generate step-by-step procedures. 030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026Olmix blog: allenai.org/blog/olmix Extended paper: arxiv.org/abs/2602.12237allenai.orgOlmix: A framework for data mixing throughout LM development | Ai2Olmix is a framework for language model data mixing that provides empirically grounded defaults and efficient reuse techniques. 120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026excited to see frens at #icml2026 & present 🐟 Olmix: efficient data mixing under token constraints & evolving data domains 🐡 How2Everything: mining the web for diverse procedural tasks for train & eval 🐠 happy to chat data & evals, both pre & post-training 191
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/07/2026audiobooks have rlly improved my new commute. discovered Libby & dunno why anyone would pay for Audible 171
Reposted by Kyle Lo @ COLM2026Melanie Walsh @mellymeldubs.bsky.social · 24/06/2026Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks. 515645
Reposted by Kyle Lo @ COLM2026Anna Rogers @annarogers.bsky.social · 26/06/2026I'm really sorry to miss all the fun at @facct.bsky.social this year! But @nlp-amelie.bsky.social and Mattes Ruckdeschel, the first two authors of this work ⬇️, are around. 0122
Reposted by Kyle Lo @ COLM2026Maria Antoniak @mariaa.bsky.social · 19/06/2026New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization! 18015
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/05/2026it could also go the other way.. 🚨 Hot take: R2 is massively underestimating how impressive our results are. A few things that feel obvious but aren’t: 👉 L23-45 explains that contrary to what R2 thinks — our idea is novel 👉 Table 2 shows we indeed implemented the baseline R2 completely missed 050
Kyle Lo @ COLM2026 @kylelo.bsky.social · 01/05/2026during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs 05811
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/04/2026oh lol ppl have been submitting wout reviewing forever, TIL it was boycotting all along 120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/04/2026kinda out of the loop, ppl are submitting to neurips but not reviewing? 110
Kyle Lo @ COLM2026 @kylelo.bsky.social · 28/03/2026thanks maria! glad got to share a fun office and collaborate during s2 days! appreciate can both chat abt difficult research problems but also peak taste tv shows w u 😆 will be in touch!! 020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 28/03/2026Today I'm saying farewell to @ai2.bsky.social. I'm so proud of our team & grateful to have shared fully-open Olmo, Dolma, olmOCR, Molmo, etc with the world I know the team is more committed than ever to advancing open-source & open-science. Forever rooting for my dear friends 🫶 3541
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026cs peer review atm feels like im in a user study that forgot to get irb review 140
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026lololol I subscribe to the @mariaa.bsky.social school of cozy figures 130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026for figs/diagrams, ive been found nano banana generates images a bit too cringe-tech for me, have had some success w committing to images all in matplotlib code, one script per fig 120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026nice post! will need to check out reveal. some of my colleagues and i have a similar workflows using markdown instead of html, but the idea of some structured doc that is in-distribution for LMs seems the right path 120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/03/2026big congrats to @lambdaviking.bsky.social for leading this project & core contributors Yanghong Li @tylerromero.bsky.social @anejsvete.bsky.social Caia Costello blog: allenai.org/blog/olmohyb... paper: allenai.org/papers/olmo-... hf collection: huggingface.co/collections/...allenai.orgIntroducing Olmo Hybrid: Combining transformers and linear RNNs for superior scaling | Ai2Ai2, a non-profit research institute founded by Paul Allen, is committed to breakthrough AI to solve the world’s biggest problems. 031
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/03/2026our new Olmo Hybrid model combines attention with linear RNN layers 🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks as always: weights, data, ckpts, training code, etc. all fully open 1354
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/03/2026DrawEduMath is our benchmark testing VLM understanding of K-12 student math work, which is prerequisite for their use in educational contexts one year after, while VLMs are strong math solvers today, they still underperform on our bench, esp for students who need the most help 030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026this work was led by our intern Mayee Chen and was one of the new ideas we adopted into Olmo 3! blog post: allenai.org/blog/olmix arxiv paper: arxiv.org/abs/2602.12237allenai.orgOlmix: A framework for data mixing throughout LM development | Ai2Olmix is a framework for language model data mixing that provides empirically grounded defaults and efficient reuse techniques. 020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026one of my favorite topics is dealing with data constraints! what if your proposed mix is 30% code but you don't have enough code? we can repeat our data until we hit target proportions, but too much is risky we view data mixing as (data) constrained optimization 130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all! 1294
Kyle Lo @ COLM2026 @kylelo.bsky.social · 10/02/2026learning how to do something is a first-order use case for LMs, the development bottleneck has been collecting data covering a wide diversity of topics, until now ✌🏻 020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 10/02/2026incredibly fun project led by our intern yapei chang we mined the web for thousands of real-world “how to do X” step by step instructions and turned it into a dataset, synth data training procedure, eval suite, etc. 1283
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/02/2026lol rip 😮💨 It’s like a score calculated against gold reference citations in generated lit review, so even humans don’t score high. i think the eval is saturated cuz so much subjectivity in what counts as appropriate citation. better phrasing is maybe that the citations are sensible up to some X 050
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/02/2026they’re separate poorly named systems lol 😂 Separate projects approaching same problem from different angles. Scholar QA approach from agentic system design, use whatever model. Ope Scholar approach from model-first, very light on system. The teams are working together to fuse ideas 010
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/02/2026our open model proving out specialized rag LMs over scientific literature has been published in nature ✌🏻 congrats to our lead @akariasai.bsky.social & team of students and Ai2 researchers/engineers www.nature.com/articles/s41... 24310
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/02/20260 days since last mixup of eval results between "copa" (choice of plausible alternatives) & "coqa" (conversational QA) tasks 😐 040
Kyle Lo @ COLM2026 @kylelo.bsky.social · 27/01/2026The 5th Generation, Evaluation, and Metrics (GEM) Workshop will be at #ACL2026! Call for papers is out. Topics include: 🐟 LMs as evaluators 🐠 Living benchmarks 🍣 Eval with humans and more New for 2026: Opinion & Statement Papers! Full CFP: gem-workshop.com/call-for-pap... 0227
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026mm yea i think that's always the case w productivity tools. imo ability to adopt new tools is core part of the job. just like transition from plain text editors to IDEs, from sending files via FPT to using git for collab, from ad hoc Makefiles to package managers, etc. AI is just the latest thing 030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026my concern is the growing pool of "unknown unknowns" as i interact less with code directly. imo probably why i subconsciously have been leaning toward cursor over claude code or similar agents, even if the latter has a higher code-to-keystrokes ratio 070
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026i dont feel worse at this even if im not writing papers from-scratch as much as during early career but coding feels different due to mismatch between what i express to the system (english) and what the system returns (code). i've already realized some gaps in libraries I used to know well. 180
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026whether my ability to review code will degrade as I offload increasingly larger workloads to AI of course, this shift is present in other forms of generation, like paper writing, where my role has shifted to reviewing/editing (student's) drafts. 160
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026some thoughts about skill degradation w/ AI coding im onboard w views that "english is the new programming language" & "software engineering", translating ambiguous goals to technical specs/execution, is still a skill. im more concerned w shift from my role as a writer to a reviewer and 2150
Kyle Lo @ COLM2026 @kylelo.bsky.social · 18/01/2026lucky to chat w sen. patty murray about olmo & importance of fully open AI 2501
Kyle Lo @ COLM2026 @kylelo.bsky.social · 17/01/2026using opus to extract research topics from papers & it was giving me useless words like "training", "datasets", and "evaluation" kept prompting it w examples of more informative topics and it ended up with "LLM training", "LLM datasets", and "LLM evaluation" thx 3130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 16/01/2026just realized ive had food on my face all day & nobody at office told me, thx ai2 frens 😫 060
Kyle Lo @ COLM2026 @kylelo.bsky.social · 15/01/2026u gotta shitpost more maria, ur content too informative 😆 370
Kyle Lo @ COLM2026 @kylelo.bsky.social · 15/01/2026i appreciate bsky has less AI product advertising; i do want to see more memes/shitposting/fun stuff and insights from industry/open source sphere, even if they dont have an attached paper 170
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026ive been clicking around in UI but i cant find it 😭 pls help 100
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026bsky wish list i like the idea of different feeds but i actually want my subscription to select feeds to be taken as a preference signal ("more like this") that informs a "home/default" feed. i really dislike the UX of having to tab through each subscribed feed, esp when there's also post overlap 381
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026some notion of 'views/impressions'? it kinda sucks to post and only see a couple of likes & no replies. if there's some intermediate signal that shows people at least read the post, that'd incentivize more imo 130