Sign in

Kyle Lo @ COLM2026

@kylelo.bsky.social
6.8K followers 600 following 585 posts

language models, data & evals, prev co-lead of Olmo @ai2.bsky.social, nlp @uwcse, statistics @uw, open science, tabletop, seattle, he/him,🧋 kyleclo.com

PostsRepliesMedia
Kyle Lo @ COLM2026 @kylelo.bsky.social · 06/10/2026
Olmo Hybrid: arxiv.org/abs/2604.03444 Cracks in foundation: arxiv.org/abs/2608.10296
000
Kyle Lo @ COLM2026 @kylelo.bsky.social · 06/10/2026
I'm at #colm2026, Mon-Thurs Supporting two papers from @ai2.bsky.social days 🐟Olmo Hybrid, 7B hybrid model 🐠Cracks in the Foundation, long context recipes Hoping to have fun chats w folks about scaling LM data & evals, ☺️ unbiased on pretrain, posttrain, all the 🚂s
1160
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026
How2Everything blog: allenai.org/blog/how2eve... Extended paper: arxiv.org/abs/2602.088...
allenai.org
How2Everything: Mining the web to evaluate and improve LLMs on real-world procedures | Ai2
How2Everything is an open framework for evaluating and improving how well LLMs generate step-by-step procedures.
030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026
Olmix blog: allenai.org/blog/olmix Extended paper: arxiv.org/abs/2602.12237
allenai.org
Olmix: A framework for data mixing throughout LM development | Ai2
Olmix is a framework for language model data mixing that provides empirically grounded defaults and efficient reuse techniques.
120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/07/2026
excited to see frens at #icml2026 & present 🐟 Olmix: efficient data mixing under token constraints & evolving data domains 🐡 How2Everything: mining the web for diverse procedural tasks for train & eval 🐠 happy to chat data & evals, both pre & post-training
191
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/07/2026
audiobooks have rlly improved my new commute. discovered Libby & dunno why anyone would pay for Audible
171
Reposted by Kyle Lo @ COLM2026
Melanie Walsh @mellymeldubs.bsky.social · 24/06/2026
Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks.
Screenshot of paper abstract that reads: 

AI FICTION IN THE WILD Neel Gupta  Maria Antoniak  Melanie Walsh

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, English-language ChatGPT-user conversations (Zhao et al.), we find that more than one third of the conversations involve some form of fiction generation—including original stories, roleplay, fanfiction, and erotica. This AI-generated fiction is notably dominated by power users. We identify common fiction generation patterns and profiles among these users, including what we call infinite story demanders, who repeatedly request and revise variations of the same or similar narratives over extended periods of time. We show that users especially gravitate toward fanfiction and erotica, and that they are broadly drawn to generic forms, repetition, immediacy, and niche combinations of story elements. Our findings motivate two theoretical provocations. First, we argue that AI technologies may lead to a shift in the conventional relationship between the author and reader, potentially producing what we call a solipsistic reader-writer, who both generates and consumes fiction within a closed conversational loop, interacting with a machine rather than a human other. Second, we note that LLMs enable interactivity, play, and permutation in ways that are seemingly pleasurable for users, raising questions about where AI will fit into contemporary storytelling and entertainment ecosystems. We situate these developments within broader transformations in literature and media, including self-publishing, fanfiction, and pornography, and suggest that AI-generated fiction shares structural affinities with on-demand, personalized, and repetitive cultural forms.
515645
Reposted by Kyle Lo @ COLM2026
Anna Rogers @annarogers.bsky.social · 26/06/2026
I'm really sorry to miss all the fun at @facct.bsky.social this year! But @nlp-amelie.bsky.social and Mattes Ruckdeschel, the first two authors of this work ⬇️, are around.
0122
Reposted by Kyle Lo @ COLM2026
Maria Antoniak @mariaa.bsky.social · 19/06/2026
New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!
Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
18015
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/05/2026
it could also go the other way.. 🚨 Hot take: R2 is massively underestimating how impressive our results are. A few things that feel obvious but aren’t: 👉 L23-45 explains that contrary to what R2 thinks — our idea is novel 👉 Table 2 shows we indeed implemented the baseline R2 completely missed
050
Kyle Lo @ COLM2026 @kylelo.bsky.social · 01/05/2026
wholesome science 😊
020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 01/05/2026
during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs
05811
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/04/2026
oh lol ppl have been submitting wout reviewing forever, TIL it was boycotting all along
120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/04/2026
kinda out of the loop, ppl are submitting to neurips but not reviewing?
110
Kyle Lo @ COLM2026 @kylelo.bsky.social · 28/03/2026
thanks for the support!
010
Kyle Lo @ COLM2026 @kylelo.bsky.social · 28/03/2026
thanks maria! glad got to share a fun office and collaborate during s2 days! appreciate can both chat abt difficult research problems but also peak taste tv shows w u 😆 will be in touch!!
020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 28/03/2026
Today I'm saying farewell to @ai2.bsky.social. I'm so proud of our team & grateful to have shared fully-open Olmo, Dolma, olmOCR, Molmo, etc with the world I know the team is more committed than ever to advancing open-source & open-science. Forever rooting for my dear friends 🫶
3541
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026
cs peer review atm feels like im in a user study that forgot to get irb review
140
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026
lololol I subscribe to the @mariaa.bsky.social school of cozy figures
130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026
for figs/diagrams, ive been found nano banana generates images a bit too cringe-tech for me, have had some success w committing to images all in matplotlib code, one script per fig
120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 26/03/2026
nice post! will need to check out reveal. some of my colleagues and i have a similar workflows using markdown instead of html, but the idea of some structured doc that is in-distribution for LMs seems the right path
120
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/03/2026
big congrats to @lambdaviking.bsky.social for leading this project & core contributors Yanghong Li @tylerromero.bsky.social @anejsvete.bsky.social Caia Costello blog: allenai.org/blog/olmohyb... paper: allenai.org/papers/olmo-... hf collection: huggingface.co/collections/...
allenai.org
Introducing Olmo Hybrid: Combining transformers and linear RNNs for superior scaling | Ai2
Ai2, a non-profit research institute founded by Paul Allen, is committed to breakthrough AI to solve the world’s biggest problems.
031
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/03/2026
our new Olmo Hybrid model combines attention with linear RNN layers 🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks as always: weights, data, ckpts, training code, etc. all fully open
1354
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/03/2026
DrawEduMath is our benchmark testing VLM understanding of K-12 student math work, which is prerequisite for their use in educational contexts one year after, while VLMs are strong math solvers today, they still underperform on our bench, esp for students who need the most help
030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026
this work was led by our intern Mayee Chen and was one of the new ideas we adopted into Olmo 3! blog post: allenai.org/blog/olmix arxiv paper: arxiv.org/abs/2602.12237
allenai.org
Olmix: A framework for data mixing throughout LM development | Ai2
Olmix is a framework for language model data mixing that provides empirically grounded defaults and efficient reuse techniques.
020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026
one of my favorite topics is dealing with data constraints! what if your proposed mix is 30% code but you don't have enough code? we can repeat our data until we hit target proportions, but too much is risky we view data mixing as (data) constrained optimization
130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 13/02/2026
our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all!
1294
Kyle Lo @ COLM2026 @kylelo.bsky.social · 11/02/2026
literally all the time 😮‍💨 this was yesterday
030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 10/02/2026
learning how to do something is a first-order use case for LMs, the development bottleneck has been collecting data covering a wide diversity of topics, until now ✌🏻
020
Kyle Lo @ COLM2026 @kylelo.bsky.social · 10/02/2026
incredibly fun project led by our intern yapei chang we mined the web for thousands of real-world “how to do X” step by step instructions and turned it into a dataset, synth data training procedure, eval suite, etc.
1283
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/02/2026
lol rip 😮‍💨 It’s like a score calculated against gold reference citations in generated lit review, so even humans don’t score high. i think the eval is saturated cuz so much subjectivity in what counts as appropriate citation. better phrasing is maybe that the citations are sensible up to some X
050
Kyle Lo @ COLM2026 @kylelo.bsky.social · 05/02/2026
they’re separate poorly named systems lol 😂 Separate projects approaching same problem from different angles. Scholar QA approach from agentic system design, use whatever model. Ope Scholar approach from model-first, very light on system. The teams are working together to fuse ideas
010
Kyle Lo @ COLM2026 @kylelo.bsky.social · 04/02/2026
our open model proving out specialized rag LMs over scientific literature has been published in nature ✌🏻 congrats to our lead @akariasai.bsky.social & team of students and Ai2 researchers/engineers www.nature.com/articles/s41...
24310
Kyle Lo @ COLM2026 @kylelo.bsky.social · 03/02/2026
0 days since last mixup of eval results between "copa" (choice of plausible alternatives) & "coqa" (conversational QA) tasks 😐
040
Kyle Lo @ COLM2026 @kylelo.bsky.social · 27/01/2026
The 5th Generation, Evaluation, and Metrics (GEM) Workshop will be at #ACL2026! Call for papers is out. Topics include: 🐟 LMs as evaluators 🐠 Living benchmarks 🍣 Eval with humans and more New for 2026: Opinion & Statement Papers! Full CFP: gem-workshop.com/call-for-pap...
0227
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026
mm yea i think that's always the case w productivity tools. imo ability to adopt new tools is core part of the job. just like transition from plain text editors to IDEs, from sending files via FPT to using git for collab, from ad hoc Makefiles to package managers, etc. AI is just the latest thing
030
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026
my concern is the growing pool of "unknown unknowns" as i interact less with code directly. imo probably why i subconsciously have been leaning toward cursor over claude code or similar agents, even if the latter has a higher code-to-keystrokes ratio
070
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026
i dont feel worse at this even if im not writing papers from-scratch as much as during early career but coding feels different due to mismatch between what i express to the system (english) and what the system returns (code). i've already realized some gaps in libraries I used to know well.
180
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026
whether my ability to review code will degrade as I offload increasingly larger workloads to AI of course, this shift is present in other forms of generation, like paper writing, where my role has shifted to reviewing/editing (student's) drafts.
160
Kyle Lo @ COLM2026 @kylelo.bsky.social · 21/01/2026
some thoughts about skill degradation w/ AI coding im onboard w views that "english is the new programming language" & "software engineering", translating ambiguous goals to technical specs/execution, is still a skill. im more concerned w shift from my role as a writer to a reviewer and
2150
Kyle Lo @ COLM2026 @kylelo.bsky.social · 18/01/2026
lucky to chat w sen. patty murray about olmo & importance of fully open AI
2501
Kyle Lo @ COLM2026 @kylelo.bsky.social · 17/01/2026
using opus to extract research topics from papers & it was giving me useless words like "training", "datasets", and "evaluation" kept prompting it w examples of more informative topics and it ended up with "LLM training", "LLM datasets", and "LLM evaluation" thx
3130
Kyle Lo @ COLM2026 @kylelo.bsky.social · 16/01/2026
yo endorse me for python skills
100
Kyle Lo @ COLM2026 @kylelo.bsky.social · 16/01/2026
just realized ive had food on my face all day & nobody at office told me, thx ai2 frens 😫
060
Kyle Lo @ COLM2026 @kylelo.bsky.social · 15/01/2026
u gotta shitpost more maria, ur content too informative 😆
370
Kyle Lo @ COLM2026 @kylelo.bsky.social · 15/01/2026
i appreciate bsky has less AI product advertising; i do want to see more memes/shitposting/fun stuff and insights from industry/open source sphere, even if they dont have an attached paper
170
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026
amaazinggg thxx 🙏🙏🙏
000
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026
ive been clicking around in UI but i cant find it 😭 pls help
100
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026
bsky wish list i like the idea of different feeds but i actually want my subscription to select feeds to be taken as a preference signal ("more like this") that informs a "home/default" feed. i really dislike the UX of having to tab through each subscribed feed, esp when there's also post overlap
381
Kyle Lo @ COLM2026 @kylelo.bsky.social · 14/01/2026
some notion of 'views/impressions'? it kinda sucks to post and only see a couple of likes & no replies. if there's some intermediate signal that shows people at least read the post, that'd incentivize more imo
130