Sign in

Kyle Lo @ ICML2026 🇰🇷

@kylelo.bsky.social
6.8K followers 600 following 583 posts

language models, data & evals, prev co-lead of Olmo @ai2.bsky.social, nlp @uwcse, statistics @uw, open science, tabletop, seattle, he/him,🧋 kyleclo.com

PostsRepliesMedia
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 05/07/2026
excited to see frens at #icml2026 & present 🐟 Olmix: efficient data mixing under token constraints & evolving data domains 🐡 How2Everything: mining the web for diverse procedural tasks for train & eval 🐠 happy to chat data & evals, both pre & post-training
191
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 04/07/2026
audiobooks have rlly improved my new commute. discovered Libby & dunno why anyone would pay for Audible
171
Reposted by Kyle Lo @ ICML2026 🇰🇷
Melanie Walsh @mellymeldubs.bsky.social · 24/06/2026
Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks.
Screenshot of paper abstract that reads: 

AI FICTION IN THE WILD Neel Gupta  Maria Antoniak  Melanie Walsh

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, English-language ChatGPT-user conversations (Zhao et al.), we find that more than one third of the conversations involve some form of fiction generation—including original stories, roleplay, fanfiction, and erotica. This AI-generated fiction is notably dominated by power users. We identify common fiction generation patterns and profiles among these users, including what we call infinite story demanders, who repeatedly request and revise variations of the same or similar narratives over extended periods of time. We show that users especially gravitate toward fanfiction and erotica, and that they are broadly drawn to generic forms, repetition, immediacy, and niche combinations of story elements. Our findings motivate two theoretical provocations. First, we argue that AI technologies may lead to a shift in the conventional relationship between the author and reader, potentially producing what we call a solipsistic reader-writer, who both generates and consumes fiction within a closed conversational loop, interacting with a machine rather than a human other. Second, we note that LLMs enable interactivity, play, and permutation in ways that are seemingly pleasurable for users, raising questions about where AI will fit into contemporary storytelling and entertainment ecosystems. We situate these developments within broader transformations in literature and media, including self-publishing, fanfiction, and pornography, and suggest that AI-generated fiction shares structural affinities with on-demand, personalized, and repetitive cultural forms.
515645
Reposted by Kyle Lo @ ICML2026 🇰🇷
Anna Rogers @annarogers.bsky.social · 26/06/2026
I'm really sorry to miss all the fun at @facct.bsky.social this year! But @nlp-amelie.bsky.social and Mattes Ruckdeschel, the first two authors of this work ⬇️, are around.
0122
Reposted by Kyle Lo @ ICML2026 🇰🇷
Maria Antoniak @mariaa.bsky.social · 19/06/2026
New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!
Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
18015
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 01/05/2026
during in Olmo 3 we thought long context is just finding good data nope! model architecture matters & it's hard to recover if mess it up led by @abertsch.bsky.social, we release many pretrain runs w/ small arch changes and show huge long context performance diffs
05811
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 28/03/2026
Today I'm saying farewell to @ai2.bsky.social. I'm so proud of our team & grateful to have shared fully-open Olmo, Dolma, olmOCR, Molmo, etc with the world I know the team is more committed than ever to advancing open-source & open-science. Forever rooting for my dear friends 🫶
3541
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 05/03/2026
our new Olmo Hybrid model combines attention with linear RNN layers 🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks as always: weights, data, ckpts, training code, etc. all fully open
1354
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 03/03/2026
DrawEduMath is our benchmark testing VLM understanding of K-12 student math work, which is prerequisite for their use in educational contexts one year after, while VLMs are strong math solvers today, they still underperform on our bench, esp for students who need the most help
030
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 13/02/2026
our paper on data mixing for LMs is out! while building Olmo 3, we saw gaps between data mixing literature and real practice 🐠choosing proxy size, # runs, sampling, regression, constraints.. 🐟data shifts during LM dev: can we reuse past experiments? Olmix tackles them all!
1294
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 10/02/2026
incredibly fun project led by our intern yapei chang we mined the web for thousands of real-world “how to do X” step by step instructions and turned it into a dataset, synth data training procedure, eval suite, etc.
1283
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 04/02/2026
our open model proving out specialized rag LMs over scientific literature has been published in nature ✌🏻 congrats to our lead @akariasai.bsky.social & team of students and Ai2 researchers/engineers www.nature.com/articles/s41...
24310
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 03/02/2026
0 days since last mixup of eval results between "copa" (choice of plausible alternatives) & "coqa" (conversational QA) tasks 😐
040
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 27/01/2026
The 5th Generation, Evaluation, and Metrics (GEM) Workshop will be at #ACL2026! Call for papers is out. Topics include: 🐟 LMs as evaluators 🐠 Living benchmarks 🍣 Eval with humans and more New for 2026: Opinion & Statement Papers! Full CFP: gem-workshop.com/call-for-pap...
0227
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 21/01/2026
some thoughts about skill degradation w/ AI coding im onboard w views that "english is the new programming language" & "software engineering", translating ambiguous goals to technical specs/execution, is still a skill. im more concerned w shift from my role as a writer to a reviewer and
2150
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 18/01/2026
lucky to chat w sen. patty murray about olmo & importance of fully open AI
2501
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 17/01/2026
using opus to extract research topics from papers & it was giving me useless words like "training", "datasets", and "evaluation" kept prompting it w examples of more informative topics and it ended up with "LLM training", "LLM datasets", and "LLM evaluation" thx
3130
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 16/01/2026
just realized ive had food on my face all day & nobody at office told me, thx ai2 frens 😫
060
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 14/01/2026
bsky wish list i like the idea of different feeds but i actually want my subscription to select feeds to be taken as a preference signal ("more like this") that informs a "home/default" feed. i really dislike the UX of having to tab through each subscribed feed, esp when there's also post overlap
381
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 07/01/2026
just in case it wasn’t clear which room this is
120
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 14/12/2025
just had hechalou’s yin yang milk tea and i think i’ve transcended 🤤
000
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 12/12/2025
during neurips, we kept the RL run going & model kept getting better 😂 Olmo 3.1 is a.. 🐡 32B Thinking, still best fully-open model to-date 🐠 32B Instruct, for ppl who hate long yapping, as good as qwen3 we added 10 more pages to the paper! thx for community feedback from convos at neurips
1181
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 01/12/2025
I'll be at #NeurIPS2025 from Tues-Sat! Come say hi 👋 if you wanna chat about 🦈 olmo 3 stories 🐟 pretraining data & evals 🍣 midtraining shouldnt exist 🐠 model specialization 🐡 AI for education 🍥 tabletop games
1192
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 24/11/2025
fml 🤦🏻‍♂️
010
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 20/11/2025
we released Olmo 3! lot of exciting stuff but wanna focus on: 🐟Olmo 3 32B Base, the best fully-open base model to-date, near Qwen 2.5 & Gemma 3 on diverse evals 🐠Olmo 3 32B Think, first fully-open reasoning model approaching Qwen 3 levels 🐡12 training datasets corresp to different staged training
1417
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 19/11/2025
going live with a mukbang tmr 🍱
020
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 14/11/2025
not happy abt gpt 5.1 update. it's making way more mistakes compared to gpt 5 on basic stuff latex table formatting errors (straight up missing "&" so columns misaligned, or dropping a whole column, or shifting values by 1 position), feels unusable imo 😒
040
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 12/11/2025
picking between 3 checkpoints w/ same benchmark scores but what if one of them is agi
1110
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 05/11/2025
why intern at Ai2? 🐟interns own major parts of our model development, sometimes even leading whole projects 🐡we're committed to open science & actively help our interns publish their work reach out if u wanna build open language models together 🤝 links 👇
2278
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 04/11/2025
congrats to our olmo earth team 🌎 small multimodal foundation language models + system for finetuning for important uses like agriculture, wildfire management, conservation & more 🌿
0100
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 22/10/2025
woah guess VLMs for OCR the hottest research topic this week😆 since the first olmOCR, we've been.. 🔥training our VLM using RLVR with binary unit test rewards🔥 it's incredibly effective & unit test creation easy to scale w synthetic data pipelines check it out at olmocr.allen.ai
0213
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 11/10/2025
bye #colm2025 big fan of the montreal bagels 🥯 hot take I like them better than
0120
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 09/10/2025
come say hi at posters this morning for OLMo 2 and fluid benchmarking posters 👋 and dont miss @valentinhofmann.bsky.social's talk in morning #colm2025 @ai2.bsky.social vry proud of my gifs
270
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 08/10/2025
@josephc.bsky.social @mariaa.bsky.social and I are at poster #21 findings from large scale survey of 800 researchers on how they use LMs in their research #colm2025
0163
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 06/10/2025
flyin to #colm2025 along w bunch of the @ai2.bsky.social team come chat w me about pretraining horror stories, data & evals, what we're cookin for next olmo, etc made a 🔥 poster for thursday sess, come say hi
0111
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 06/10/2025
5 am airport for the only direct flight from seattle to montreal #colm2025
1120
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 02/10/2025
not my project but I rlly like it working w cancer research center to analyze clinical data, but private data cant leave the center. so the team developed a tool that generates code for remote execution by the cancer center, developed on synthetic data, and now tested for realsies 🤩
data.at
Home - D.A.T.A.
Der komplette Workflow einer Radiologie, abgebildet in einer modularen Software, erstellt von Spezialisten. Das ist die XR RadiologySuite®
140
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 19/09/2025
had to explain to first time submitter why AC recommended accept ended up as reject 😮‍💨 been publishing long enough that i get why such things happen but can be rough
080
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 17/09/2025
LM benchmark design requires 3 decisions, how to: 🐟 select test cases 🐠 score LM on each test 🦈 aggregate scores to estimate perf fluid benchmarking is simple: 🍣 find max informative test cases 🍥 estimate 'ability', not simple avg perf why care? turn ur grey noisy benchmarks to red ones!
052
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 12/09/2025
scathing takedown of recent K2 Think model "evaluates on data it was trained on, relies on an external model and additional samples for its claimed performance gains, and artificially reduces the scores of compared models" www.sri.inf.ethz.ch/blog/k2think
sri.inf.ethz.ch
Debunking the Claims of K2-Think
K2-Think is a recently released LLM that claims performance on par with GPT-OSS 120B and DeepSeek v3.1, despite having fewer parameters. As we discuss below, the reported gains are overstated, relying...
061
Reposted by Kyle Lo @ ICML2026 🇰🇷
Nathan Lambert @natolambert.bsky.social · 04/09/2025
COLM is coming up! Very excited. I'm starting to figure out two things: 1. A small invite-only dinner for Interconnects AI (Ai2 event news later). 2. Various research chats and catchups. Fill out the form below or email me if you're interested :) 🍁🇨🇦 Interest form: buff.ly/9nWBxZ9
071
Reposted by Kyle Lo @ ICML2026 🇰🇷
Ai2 @ai2.bsky.social · 28/08/2025
🎙️ Say hello to OLMoASR—our fully open, from-scratch speech-to-text (STT) model. Trained on a curated audio-text set, it boosts zero-shot ASR and now powers STT in the Ai2 Playground. 👇
1196
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 20/08/2025
"Out of 13,048 reviewers..only 69 were deemed highly irresponsible..and enforcement was applied solely in those cases...These reviewers were contacted multiple times...as well as being personally contacted by the area chairs and senior area chairs, but still failed to fulfill them." 🫡🫡🫡
0115
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 19/08/2025
my favorite figure from work by @davidheineman.com if you're frustrated by LM evals, not knowing if results are real or noise, it's useful to decompose sources of variance: 🐠is there enough meaningful spread (signal) among compared models 🐟do scores vary between intermediate checkpoints (noise)
020
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 19/08/2025
very nice work by @datologyai.com folks on synth data for pretraining very nice results over nemotron synth, which we've generally been impressed by rephrasing the web (arxiv.org/abs/2401.16380) seems very powerful & good demonstration of how to push it further
arxiv.org
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance o...
1133
Reposted by Kyle Lo @ ICML2026 🇰🇷
Ai2 @ai2.bsky.social · 18/08/2025
We’re releasing early pre-training checkpoints for OLMo-2-1B to help study how LLM capabilities emerge. They’re fine-grained snapshots intended for analysis, reproduction, and comparison. 🧵
1276
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 14/08/2025
huge thx to the NSF & NVIDIA for supporting our work on fully open AI model science & development 🤩
2120
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 08/08/2025
⚠️ AI-generated content may be inaccurate. Verify important information independently.
040
Reposted by Kyle Lo @ ICML2026 🇰🇷
Marzena Karpinska @markar.bsky.social · 08/08/2025
GPT-5 lands first place on NoCha, our long-context book understanding benchmark. That said, this is a tiny improvement (~1%) over o1-preview, which was released almost one year ago. Have long-context models hit a wall? Accuracy of human readers is >97%... Long way to go!
Screenshot of benchmark with gpt-5 on top with 68.46% accuracy.
1186
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 08/08/2025
uniting the internet w chart crimes lol
0120