Sign in

Hamish Ivison

@hamishivi.bsky.social
1.2K followers 378 following 63 posts

I (try to) do NLP research. Antipodean abroad. currently doing PhD @uwcse, prev @usyd @ai2 🇦🇺🇨🇦🇬🇧 ivison.id.au

PostsRepliesMedia
Reposted by Hamish Ivison
Michael Noukhovitch @mnoukhov.bsky.social · 15/09/2026
Is RL actually making your LLM better? Gains from RL are mostly on easy questions🤯 We're calling this the Matthew Effect for RL on LLMs. We then leverage async RL to solve harder problems by Never Giving Up! arxiv.org/abs/2609.13443 and mnoukhov.github.io/posts/ngu/ and check out thread below 🧵👇
1267
Hamish Ivison @hamishivi.bsky.social · 04/02/2026
Wrote up some results around reproducing length control from the L1 paper in RL. ivison.id.au/2026/02/02/r...
010
Reposted by Hamish Ivison
Ai2 @ai2.bsky.social · 08/05/2025
We’re live on Reddit! Ask us Anything about our OLMo family of models. We have six of our researchers on hand to answer all your questions.
192
Hamish Ivison @hamishivi.bsky.social · 07/05/2025
I’ll be around for this! Come ask us questions about olmo and tulu :)
041
Hamish Ivison @hamishivi.bsky.social · 27/03/2025
Excited to be back home in Australia (Syd/Melb) for most of April! Email or DM if you want to grab a coffee :)
030
Reposted by Hamish Ivison
Nathan Lambert @natolambert.bsky.social · 17/03/2025
@vwxyzjn.bsky.social and @hamishivi.bsky.social have uploaded intermediate checkpoints for our recent RL models at Ai2. Folks should do research into how RL finetuning is impacting the weights! Models with it: OLMo 2 7B, 13B, 32B Instruct; Tulu 3, 3.1 8B; Tulu 3 405b
0122
Hamish Ivison @hamishivi.bsky.social · 04/03/2025
How well do data-selection methods work for instruction-tuning at scale? Turns out, when you look at large, varied data pools, lots of recent methods lag behind simple baselines, and a simple embedding-based method (RDS) does best! More below ⬇️ (1/8)
1134
Hamish Ivison @hamishivi.bsky.social · 20/02/2025
(1/8) Excited to share some new work: TESS 2! TESS 2 is an instruction-tuned diffusion LM that can perform close to AR counterparts for general QA tasks, trained by adapting from an existing pretrained AR model. 📜 Paper: arxiv.org/abs/2502.13917 🤖 Demo: huggingface.co/spaces/hamis... More below ⬇️
141
Hamish Ivison @hamishivi.bsky.social · 12/02/2025
GRPO makes everything better 😌
020
Hamish Ivison @hamishivi.bsky.social · 12/02/2025
020
Reposted by Hamish Ivison
Ai2 @ai2.bsky.social · 11/02/2025
We took our most efficient model and made an open-source iOS app📱but why? As phones get faster, more AI will happen on device. With OLMoE, researchers, developers, and users can get a feel for this future: fully private LLMs, available anytime. Learn more from @soldaini.net👇 youtu.be/rEK_FZE5rqQ
youtu.be
Ai2 OLMoE: Fully open source, running entirely on-device
YouTube video by Ai2
23014
Hamish Ivison @hamishivi.bsky.social · 30/01/2025
li'l holiday project from the tulu team :) Scaling up the Tulu recipe to 405B works pretty well! We mainly see this as confirmation that open-instruct scales to large-scale training -- more exciting and ambitious things to come!
1141
Hamish Ivison @hamishivi.bsky.social · 20/01/2025
Seems like a good time to share this: a poster from a class project diving a little more into Tulu 3's RLVR. Deepseek R1 release today shows that scaling this sort of approach up can be very very effective!
110
Hamish Ivison @hamishivi.bsky.social · 08/01/2025
Excited to see Tulu 3 sits in between Llama 3.1 and 3.3 instruct on the chatbot arena leaderboard right now! Particularly happy it is top 20 for Math and Multi-turn prompts :) All the details and data on how to train a model this good are right here: arxiv.org/abs/2411.15124
0153
Reposted by Hamish Ivison
Costa Huang @vwxyzjn.bsky.social · 06/01/2025
We released the OLMo 2 report! Ready for some more RL curves? 😏 This time, we applied RLVR iteratively! Our initial RLVR checkpoint on the RLVR dataset mix shows a low GSM8K score, so we did another RLVR on GSM8K only and another on MATH only 😆. And it works! A thread 🧵 1/N
1125
Hamish Ivison @hamishivi.bsky.social · 04/01/2025
More OLMo! More performance! More details! We applied Tulu post-training to OLMo 2 as well, so you can get strong model performance AND see what your model was actually trained on.
060
Reposted by Hamish Ivison
Natasha Jaques @natashajaques.bsky.social · 18/12/2024
UW News put out a Q&A about our recent work on Variational Preference Learning, a technique for personalizing Reinforcement Learning from Human Feedback (RLHF) washington.edu/news/2024/12...
washington.edu
Q&A: New AI training method lets systems better adjust to users’ values
University of Washington researchers created a method for training AI systems — both for large language models like ChatGPT and for robots — that can better reflect users’ diverse values. It...
1308
Reposted by Hamish Ivison
Jiacheng Liu @liujch1998.bsky.social · 09/12/2024
Want to predict the task performance of LMs before pretraining them? We develop task scaling laws and model ladders, which predict the accuracy on individual tasks by OLMo 2 7B & 13B models within 2 points of absolute error. The cost is 1% of the compute used to pretrain them.
23314
Hamish Ivison @hamishivi.bsky.social · 06/12/2024
New OpenAI RL finetuning API reminds me a lot of RLVR, which we used for Tülu 3 (arxiv.org/abs/2411.15124). Using RL to train against labels is a simple idea, but very effective (>10pt gains just using GSM8k train set). It's implemented for you to use in Open-Instruct 😉: github.com/allenai/open...
Test accuracy, train rewards, kl divergence, and response lenght training curves when training Tulu 3 SFT and Tulu 3 DPO on the MATH or GSM8k train sets, and evaluating on MATH/GSM8k using RLVR. Performance significantly improves in both cases.
171
Reposted by Hamish Ivison
Nathan Lambert @natolambert.bsky.social · 06/12/2024
OpenAI announced a new RL finetuning API. You can do this on open models w the repo we used to train Tulu 3. Expanding reinforcement learning with verifiable rewards to more domains and with better answer extraction and to more domains in our near roadmap. buff.ly/3V4JEIJ
1479
Reposted by Hamish Ivison
Matthew Finlayson @mattf.nl · 06/12/2024
Curious about all this inference-time scaling hype? Attend our NeurIPS tutorial: Beyond Decoding: Meta-Generation Algorithms for LLMs (Tue. 1:30)! We have a top-notch panelist lineup. Our website: cmu-l3.github.io/neurips2024-...
Panelist photos: Rishabh Agarwal (Google, McGill), Noam Brown (OpenAl), Beidi Chen (CMU), Nouha Dziri (AI2), Jakob Foerster (Oxford, Meta)
1273
Reposted by Hamish Ivison
Akari Asai @akariasai.bsky.social · 04/12/2024
I’m on the academic job market this year! I’m completing my @uwcse.bsky.social @uwnlp.bsky.social Ph.D. (2025), focusing on overcoming LLM limitations like hallucinations, by building new LMs. My Ph.D. work focuses on Retrieval-Augmented LMs to create more reliable AI systems 🧵
37117
Reposted by Hamish Ivison
Nathan Lambert @natolambert.bsky.social · 03/12/2024
We're hiring another predoctoral researcher for my team at Ai2/OLMo next year. The goal of this position is to mentor and grow future academic stars of NLP/AI over 1-2 years before grad school. This ends up being people done with BS or MS who want to continue to a PhD soon. buff.ly/49nuggo
6547
Hamish Ivison @hamishivi.bsky.social · 02/12/2024
Excited to be at #NeurIPS next week in 🇨🇦! Please reach out if you want to chat about LM post-training (Tülu!), data curation, or anything else :) I'll be around all week, with two papers you should go check out (see image or next tweet):
2132
Hamish Ivison @hamishivi.bsky.social · 30/11/2024
I know it doesn't know much if anything about me but this was pretty surprisingly good!
130
Hamish Ivison @hamishivi.bsky.social · 29/11/2024
Watching RL training curves is too addictive... begging my models to yap more and get more reward 🙏
080
Reposted by Hamish Ivison
William Merrill @lambdaviking.bsky.social · 28/11/2024
🔥 Old Norse poetry gen The Vikings call, say now, OLMo 2, the ruler of languages. May your words fly over the seas, all over the world, for you are wise. Wordsmith, balanced and aligned, for you the skalds themselves sing, your soul, which hears new lifeforms, may it live long and tell a saga.
Víkingar kalla, segja þú nú,
OLMo 2, ríki málanna þinn.
Munu þínar orð fljúga hafra,
Öll um heim, því þú ert vissi.
Málsmiður, mættugur og mjúkaligr,
Fyrir þik skáldar sjálfur kveða,
Sál þíð, sem heyrir nýjan kvikendi,
Munu langt lífið og segja sagan.
2132
Reposted by Hamish Ivison
Jacob Morrison @jacobcares.bsky.social · 26/11/2024
🍲
1182
Hamish Ivison @hamishivi.bsky.social · 26/11/2024
What's that? A fully open LM competitive with Gemma and Qwen*? Happy to have helped a bit with this release (Tulu 3 recipe used here)! OLMo-2 13B actually beats Tulu 3 8B on these evals, making it a SOTA fully open LM!!! (*on the benchmarks we looked at, see tweet for more)
1101
Reposted by Hamish Ivison
fizz ☭ @fizz.allura.moe · 24/11/2024
open source tulu 3 model recreation! rivals the original sft and other models in its size range huggingface.co/allura-org/T...
huggingface.co
allura-org/Teleut-7b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1346
Reposted by Hamish Ivison
Ai2 @ai2.bsky.social · 21/11/2024
Meet Tülu 3, a set of state-of-the-art instruct models with fully open data, eval code, and training algorithms. We invented new methods for fine-tuning language models with RL and built upon best practices to scale synthetic instruction and preference data. Demo, GitHub, paper, and models 👇
211131
Hamish Ivison @hamishivi.bsky.social · 21/11/2024
We actually have all the weights.... except the LM head for this blursed checkpoint. Don't ask.
4340
Hamish Ivison @hamishivi.bsky.social · 21/11/2024
If you love instagram reels, I made a Tulu 3 video for you :)
150
Hamish Ivison @hamishivi.bsky.social · 21/11/2024
It's just a chill release.
050
Hamish Ivison @hamishivi.bsky.social · 21/11/2024
Excited to release Tulu 3! We worked hard to try and make the best open post-training recipe we could, and the results are good! I was lucky enough to work on almost every stage of the pipeline in one way or another. Some comments + highlights ⬇️
195
Hamish Ivison @hamishivi.bsky.social · 21/11/2024
I made a bluesky account for the tulu 2 release... and then left it alone. But I'm back for Tulu 3 (and will stay)!
160
Hamish Ivison @hamishivi.bsky.social · 21/11/2023
Check out the Tulu 2 suite , a set of Llama-2 models finetuned+DPO-trained on a mixture of publicly available datasets! Our best-performing models are competitive with SoTA open models on a range of benchmarks incl. AlpacaEval and MT-Bench. 📜Paper: arxiv.org/abs/2311.10702
172