Sign in

Jacob Morrison

@jacobcares.bsky.social
575 followers 389 following 21 posts

PhD student @ UW

PostsRepliesMedia
Reposted by Jacob Morrison
Ryan Packer @typewriteralley.bsky.social · 20/09/2026
Effective government is a progressive value
1523
Reposted by Jacob Morrison
Northwest Progressive Institute @nwprogressive.org · 17/08/2026
Give Ballard a Sounder stop while it waits for light rail Give Ballard a Sounder stop while it waits for light rail "With Sound Transit's failure to deliver light rail to Ballard by 2035 as originally promised, now is the moment for Sound Transit and Seattle to build a no-frills, quick-build…
nwprogressive.org
Give Ballard a Sounder stop while it waits for light rail
Give Ballard a Sounder stop while it waits for light rail "With Sound Transit's failure to deliver light rail to Ballard by 2035 as originally promised, now is the moment for Sound Transit and Seattle to build a no-frills, quick-build station near the Ballard Locks to connect people to Sounder commuter rail service that’s already up and running," Matthew Trecha writes. Trecha is a transit enthusiast who has lived in Europe and Canada.
1406
Jacob Morrison @jacobcares.bsky.social · 13/06/2026
Interesting! Thanks for the citations, I hadn’t seen those before. I agree; I’m quite frustrated at how opaque it is, so I appreciate it. In general it was quite messy to try to get *any* estimate (even ignoring that the actual useful lifespan of these GPUs is seemingly much longer than we expected)
000
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Full per-stage breakdown, methodology, and discussion in the paper: arxiv.org/abs/2605.01158.
arxiv.org
The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining
Modern language model development extends far beyond pretraining, yet environmental reporting remains narrowly focused on the cost of training a single final model. In this work, we provide the first ...
230
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Across the full pipeline we estimate ~4,251 tCO2eq and ~15,887 kL of water for the Olmo 3 series, which is equal to nearly 850 US homes' annual electricity, or 140 years of water use for the average person in the US.
100
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
We also estimated our total water use: ~15,887 kL, none of which was from datacenter cooling. Our cluster uses closed-loop cooling, so all of it came from power generation. Just changing the grid would more than double it, and evaporative cooling would nearly double it again:
100
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Specifically, from *generating rollouts*. RL trains on long traces (up to 32k tokens, avg >10k) across many iterations. The generator runs at near-peak power; the trainer idles ~75% of the time waiting for rollouts, so 87% of Think post-training energy goes to generation.
100
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Reasoning models are far more expensive to post-train. For our 32B model, post-training our Think model takes 17x more datacenter energy than post-training the Instruct variant, and almost all of that gap is reinforcement learning.
100
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Every stage has its own experimentation cycle, with dev fractions running from 69% (pretraining) to >95% (mid-training, DPO). Concurrent work from Epoch AI estimates frontier labs at 77–90%: epoch.ai/gradient-upd...
110
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
That 82% is up from the ~50% we previously reported for Olmo 1 and 2 pretraining: arxiv.org/abs/2503.05804
arxiv.org
Holistically Evaluating the Environmental Impact of Creating Language Models
As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the po...
110
Jacob Morrison @jacobcares.bsky.social · 07/05/2026
Most LLM environmental reporting covers only the final pretraining runs. For Olmo 3, we measured every stage across all four variants: 7B and 32B, instruct and reasoning, and found that 82% of the compute went to development, all before the final runs 😱
2443
Reposted by Jacob Morrison
Kunal Jha @kjha02.bsky.social · 03/10/2025
Forget modeling every belief and goal! What if we represented people as following simple scripts instead (i.e "cross the crosswalk")? Our new paper shows AI which models others’ minds as Python code 💻 can quickly and accurately predict human behavior! shorturl.at/siUYI%F0%9F%...
33814
Reposted by Jacob Morrison
Saumya Malik @saumyamalik.bsky.social · 03/06/2025
Thank you to co-authors @natolambert.bsky.social, @valentinapy.bsky.social, @jacobcares.bsky.social, Sander Land, @nlpnoah.bsky.social, @hanna-nlp.bsky.social! Read more in the paper here (ArXiv soon!): github.com/allenai/rewa... Dataset, leaderboard, and models here: huggingface.co/collections/...
huggingface.co
Reward Bench 2 - a allenai Collection
Datasets, spaces, and models for Reward Bench 2 benchmark and paper!
021
Reposted by Jacob Morrison
Ai2 @ai2.bsky.social · 02/06/2025
RewardBench 2 is here! We took a long time to learn from our first reward model evaluation tool to make one that is substantially harder and more correlated with both downstream RLHF and inference-time scaling.
The RewardBench 2 Leaderboard on HuggingFace.
1208
Reposted by Jacob Morrison
Nathan Lambert @natolambert.bsky.social · 29/04/2025
Heading to NAACL? With "verification being the key to AI" you should go to the poster session Friday, 9-10:30am to chat with my star colleagues @valentinapy.bsky.social + @jacobcares.bsky.social about RewardBench (and really RewardBench 2, evaluation, and reward models in post-training).
0142
Jacob Morrison @jacobcares.bsky.social · 28/04/2025
Valentina and I will be presenting RewardBench at NAACL! Come say hi at the poster session on Friday and we can chat about reward models, staying up for 30 hours straight to rapidly reset from Singapore time, and more 🏜️
053
Reposted by Jacob Morrison
Valentina Pyatkin @valentinapy.bsky.social · 27/04/2025
I'll be at #NAACL2025: 🖇️To present my paper "Superlatives in Context", showing how the interpretation of superlatives is very context dependent and often implicit, and how LLMs handle such semantic underspecification 🖇️And we will present RewardBench on Friday Reach out if you want to chat!
1285
Jacob Morrison @jacobcares.bsky.social · 26/04/2025
what a flattering picture lol
120
Jacob Morrison @jacobcares.bsky.social · 23/04/2025
📜Paper: arxiv.org/abs/2503.05804 ✍️Thanks to my illustrious coauthors @clarana.bsky.social @jaredfern.bsky.social timdettmers.com @strubell.bsky.social @jessedodge.bsky.social, t'was a fun project 🌏
arxiv.org
Holistically Evaluating the Environmental Impact of Creating Language Models
As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the po...
094
Jacob Morrison @jacobcares.bsky.social · 23/04/2025
I'm in Singapore for @iclr-conf.bsky.social ! Come check out our spotlight paper on the environmental impact of training OLMo (link in next tweet) during the Saturday morning poster session from 10-12:30 -- happy to chat about this or anything else! DMs should be open, email works too
1105
Reposted by Jacob Morrison
Ai2 @ai2.bsky.social · 13/03/2025
Announcing OLMo 2 32B: the first fully open model to beat GPT 3.5 & GPT-4o mini on a suite of popular, multi-skill benchmarks. Comparable to best open-weight models, but a fraction of training compute. When you have a good recipe, ✨ magical things happen when you scale it up!
35814
Reposted by Jacob Morrison
Cats of Yore @catsofyore.bsky.social · 26/02/2025
There are no proven benefits to raw feeding yet plenty of serious, well-documented risks. It has never been a good idea but ESPECIALLY NOW. washingtonstatestandard.com/briefs/two-w...
washingtonstatestandard.com
Two Washington cats infected with bird flu • Washington State Standard
Two domestic cats in Washington state have been infected with bird flu after eating raw pet food, according to the department of agriculture.
166821646
Jacob Morrison @jacobcares.bsky.social · 30/01/2025
also some other tülu contributors are on the market: @ljvmiranda.bsky.social (ljvmiranda921.github.io) and Xinxi Lyu (alrope123.github.io) are also applying to phd programs, and @valentinapy.bsky.social (valentinapy.github.io) is on the faculty market, hire them all!!
011
Jacob Morrison @jacobcares.bsky.social · 30/01/2025
check out the updated paper here: arxiv.org/pdf/2411.15124 (with a beautiful new template!) and the model here: huggingface.co/allenai/Llam... and on the ai2 playground: playground.allenai.org
100
Jacob Morrison @jacobcares.bsky.social · 30/01/2025
big tülu is here! can't wait for everyone to try it, it's been a lot of fun seeing how RL performs at this scale thanks to @hamishivi.bsky.social and @vwxyzjn.bsky.social, and preference data from @ljvmiranda.bsky.social on an unrelated note, I'm applying to phd programs this year 👀
150
Reposted by Jacob Morrison
Ai2 @ai2.bsky.social · 30/01/2025
Here is Tülu 3 405B 🐫 our open-source post-training model that surpasses the performance of DeepSeek-V3! It demonstrates that our recipe, which includes RVLR scales to 405B - with performance on par with GPT-4o, & surpassing prior open-weight post-trained models of the same size including Llama 3.1.
The logo for Tülu 405B.
29221
Reposted by Jacob Morrison
Hamish Ivison @hamishivi.bsky.social · 08/01/2025
Excited to see Tulu 3 sits in between Llama 3.1 and 3.3 instruct on the chatbot arena leaderboard right now! Particularly happy it is top 20 for Math and Multi-turn prompts :) All the details and data on how to train a model this good are right here: arxiv.org/abs/2411.15124
0153
Reposted by Jacob Morrison
Nathan Lambert @natolambert.bsky.social · 08/01/2025
Very pleased to see Tulu 3 70B more or less tied with Llama 3.1 70B Instruct on style controlled ChatBotArena. The only model anywhere close to that with open code and data for post-training! Lots of stuff people can build on. Next looking for OLMo 2 numbers.
0243
Reposted by Jacob Morrison
Costa Huang @vwxyzjn.bsky.social · 06/01/2025
We released the OLMo 2 report! Ready for some more RL curves? 😏 This time, we applied RLVR iteratively! Our initial RLVR checkpoint on the RLVR dataset mix shows a low GSM8K score, so we did another RLVR on GSM8K only and another on MATH only 😆. And it works! A thread 🧵 1/N
1125
Reposted by Jacob Morrison
Kyle Lo @kylelo.bsky.social · 03/01/2025
kicking off 2025 with our OLMo 2 tech report while payin homage to the sequelest of sequels 🫡 🚗 2 OLMo 2 Furious 🔥 is everythin we learned since OLMo 1, with deep dives into: 🚖 stable pretrain recipe 🚔 lr anneal 🤝 data curricula 🤝 soups 🚘 tulu post-train recipe 🚜 compute infra setup 👇🧵
26917
Reposted by Jacob Morrison
Jiacheng Liu @liujch1998.bsky.social · 09/12/2024
Want to predict the task performance of LMs before pretraining them? We develop task scaling laws and model ladders, which predict the accuracy on individual tasks by OLMo 2 7B & 13B models within 2 points of absolute error. The cost is 1% of the compute used to pretrain them.
23314
Reposted by Jacob Morrison
derek guy @dieworkwear.bsky.social · 27/11/2024
Why is Tokyo so fashionable? Some theories. 🧵
Saagar Enjeti tweets: "Probably a cold take but IMO Tokyo is the male fashion capital of the world: whether it’s western wear, suits, street wear the aesthetic is refined to the highest possible level

From the salaryman to the rebel teen they are impeccably dressed

It also helps no one is fat"
21988851709
Reposted by Jacob Morrison
Luca Soldaini 🎀 @soldaini.net · 26/11/2024
OLMo 2 is out 🥳 7B and 13B trained on 5T tokens, and meticulousy instruction tuned using Tulu 3 recipe. Simply the best fully open models yet. Really proud of the work & the amazing team at @ai2.bsky.social
926044
Jacob Morrison @jacobcares.bsky.social · 26/11/2024
🍲
1182
Reposted by Jacob Morrison
Ai2 @ai2.bsky.social · 26/11/2024
Meet OLMo 2, the best fully open language model to date, including a family of 7B and 13B models trained up to 5T tokens. OLMo 2 outperforms other fully open models and competes with open-weight models like Llama 3.1 8B — As always, we released our data, code, recipes and more 🎁
The OLMo 2 models sit at the Pareto frontier of training FLOPs vs model average performance.
515235
Jacob Morrison @jacobcares.bsky.social · 21/11/2024
Thanks Tyler, great to hear from you!!
000
Reposted by Jacob Morrison
Luca Soldaini 🎀 @soldaini.net · 21/11/2024
yeah language models are great, but which Tulu 3 are you - brat tulu, a @jacobcares.bsky.social favorite - PNW tulu, don’t forget where @ai2.bsky.social is from - dank tulu 💪 - tulu at tulu, bc tulu means sunrise in farsi
191
Reposted by Jacob Morrison
Ai2 @ai2.bsky.social · 21/11/2024
Meet Tülu 3, a set of state-of-the-art instruct models with fully open data, eval code, and training algorithms. We invented new methods for fine-tuning language models with RL and built upon best practices to scale synthetic instruction and preference data. Demo, GitHub, paper, and models 👇
211131
Jacob Morrison @jacobcares.bsky.social · 21/11/2024
Thanks to everybody that worked on this, our most fun project so far 🥳 Can't wait to see what we do next!
100
Jacob Morrison @jacobcares.bsky.social · 21/11/2024
Models: huggingface.co/collections/... Training data: huggingface.co/collections/... Paper: allenai.org/papers/tulu-... Blog: allenai.org/blog/tulu-3 Technical blog: allenai.org/blog/tulu-3-... Training code: github.com/allenai/open... Eval code: github.com/allenai/olmes
100
Jacob Morrison @jacobcares.bsky.social · 21/11/2024
I'm so excited that we're finally releasing Tülu 3, our new post-training recipe! We're releasing models built on top of Llama 3.1 base (OLMo coming soon!), all of our datasets, a (73 page!) paper, new evaluations, and all of our code.
1100
Reposted by Jacob Morrison
Nathan Lambert @natolambert.bsky.social · 21/11/2024
I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread.
821343