Reposted by Alisa LiuAi2 @ai2.bsky.social · 07/05/2025📢We’re taking your questions now on Reddit for tomorrow’s AMA! Ask us anything about OLMo, our family of fully-open language models. Our researchers will be on hand to answer them Thursday, May 8 at 8am PST. 132
Alisa Liu @alisawuffles.bsky.social · 21/03/2025Right! "kick the bucket" is too infrequent, but there are more common idiomatic expressions like "in the long run" or "on the other hand." In general I would say non-idiomatic MWEs are more common, like uses of prepositions ("depend on") which require memorization. 040
Reposted by Alisa LiuValentin Hofmann @valentinhofmann.bsky.social · 21/03/2025Humans store thousands of multi-word expressions like "of course" in their mental lexicon, but current tokenizers don't support multi-word tokens. Enter SuperBPE, a tokenizer that lifts this restriction and brings substantial gains in efficiency and performance! 🚀 Details 👇 151
Reposted by Alisa LiuAlton B.H. Worthington @abhw.bsky.social · 21/03/2025Hell yeah superwords. (I wanna call em supertokens, but I didn't develop them.) 021
Reposted by Alisa LiuJonathan Hayase @jon.jon.ke · 21/03/2025Tokenizers govern the allocation of computation. It's a waste to spend a whole token of compute predicting the "way" in "By the way". SuperBPE redirects that compute to predict more difficult tokens, leading to wins on downstream tasks! 041
Alisa Liu @alisawuffles.bsky.social · 21/03/2025nothing beats writing papers together with co-1st @jon.jon.ke — the mention didn't work the first time! 030
Reposted by Alisa LiuNoah A. Smith @nlpnoah.bsky.social · 21/03/2025a small change to building your BPE tokenizer gets your pretrained LM 8 MMLU points (for example) and 27% inference-time efficiency boost ... 1111
Alisa Liu @alisawuffles.bsky.social · 21/03/2025Play around with our tokenizers here! superbpe.github.io 🚀 Paper: arxiv.org/abs/2503.13423 HF models & tokenizers: tinyurl.com/superbpe This work would not have been possible w/o co-1st 🌟@jon.jon.ke🌟, @valentinhofmann.bsky.social @sewoong79.bsky.social @nlpnoah.bsky.social @yejinchoinka.bsky.social 161
Alisa Liu @alisawuffles.bsky.social · 21/03/2025SuperBPE🚀 is a seamless replacement for BPE in modern LM development pipelines, requiring no changes to the model architecture or training framework. You can use it in HuggingFace right now! 260
Alisa Liu @alisawuffles.bsky.social · 21/03/2025Why does SuperBPE🚀 work? We find that loss is distributed more uniformly over tokens in SuperBPE models. They are less overfit to high-frequency, easy-to-predict tokens (e.g. “way” after “By the”), and at the same time master a much broader set of language phenomena. 160
Alisa Liu @alisawuffles.bsky.social · 21/03/2025Then we pretrain 8B models from scratch with BPE and SuperBPE🚀, fixing everything about the training setup except the tokenizer. We see +4% on avg📈 across 30 downstream tasks, and win on 25/30 of individual tasks, while also being 27% more efficient at inference time. 150
Alisa Liu @alisawuffles.bsky.social · 21/03/2025What can we gain from less restrictive tokenization? To find out, we developed SuperBPE🚀, which learns subword *and* superword tokens. SuperBPE dramatically improves encoding efficiency over BPE — at a fixed vocab size of 200k, SuperBPE reduces sequence length by 33% on average! 140
Alisa Liu @alisawuffles.bsky.social · 21/03/2025E.g. “math teacher” = “Mathelehrer” in German. At the extreme, Chinese *doesn’t use whitespace at all*, so its tokens can span many words — yet this has seemingly not hindered LMs like @deepseek_ai from learning it! 160
Alisa Liu @alisawuffles.bsky.social · 21/03/2025This started with a curiosity💡: why do all LLMs limit tokens to *parts* of whitespace-delimited words? After all, many word sequences (e.g. “by the way”) function as single units. Different languages can also express the same meaning in one or several words. 170
Alisa Liu @alisawuffles.bsky.social · 21/03/2025We created SuperBPE🚀, a *superword* tokenizer that includes tokens spanning multiple words. When pretraining at 8B scale, SuperBPE models consistently outperform the BPE baseline on 30 downstream tasks (+8% MMLU), while also being 27% more efficient at inference time.🧵 38316
Reposted by Alisa LiuChristopher Akiki @cakiki.bsky.social · 28/02/2025This is also addressed in the appendix of @alisawuffles.bsky.social and colleagues' paper on BPE mixture inference. I think it might have been discovered by @soldaini.net if I'm not mistaken. arxiv.org/abs/2407.16607 131
Alisa Liu @alisawuffles.bsky.social · 11/12/2024excited to be at #NeurIPS2024! I'll be presenting our data mixture inference attack 🗓️ Thu 4:30pm w/ @jon.jon.ke — stop by to learn what trained tokenizers reveal about LLM development (‼️) and chat about all things tokenizers. 🔗 arxiv.org/abs/2407.16607 0134
Reposted by Alisa LiuJiacheng Liu @liujch1998.bsky.social · 09/12/2024Want to predict the task performance of LMs before pretraining them? We develop task scaling laws and model ladders, which predict the accuracy on individual tasks by OLMo 2 7B & 13B models within 2 points of absolute error. The cost is 1% of the compute used to pretrain them. 23314
Reposted by Alisa LiuMichael Saxon @saxon.me · 06/12/2024🚨I too am on the job market‼️🤯 I'm searching for faculty positions/postdocs in multilingual/multicultural NLP, vision+language models, and eval for genAI! I'll be at #NeurIPS2024 presenting our work on meta-evaluation for text-to-image faithfulness! Let's chat there! Papers in🧵, see more: saxon.me 1488
Reposted by Alisa LiuMechanical Dirk @mechanicaldirk.bsky.social · 02/12/2024We just updated the OLMo repo at github.com/allenai/OLMo! There are now several training configs that together reproduce the training runs that lead to the final OLMo 2 models. In particular, all the training data is available, tokenized and shuffled exactly as we trained on it!github.comGitHub - allenai/OLMo: Modeling, training, eval, and inference code for OLMoModeling, training, eval, and inference code for OLMo - allenai/OLMo 05411
Reposted by Alisa LiuAi2 @ai2.bsky.social · 26/11/2024Meet OLMo 2, the best fully open language model to date, including a family of 7B and 13B models trained up to 5T tokens. OLMo 2 outperforms other fully open models and competes with open-weight models like Llama 3.1 8B — As always, we released our data, code, recipes and more 🎁 515235
Reposted by Alisa LiuLuca Soldaini 🎀 @soldaini.net · 26/11/2024OLMo 2 is out 🥳 7B and 13B trained on 5T tokens, and meticulousy instruction tuned using Tulu 3 recipe. Simply the best fully open models yet. Really proud of the work & the amazing team at @ai2.bsky.social 926044
Reposted by Alisa LiuChristoph Molnar @christophmolnar.bsky.social · 24/11/2024No one can explain stochastic gradient descent better than this panda.media.tenor.coma panda bear is rolling around in the grass in a zoo enclosure .Alt: a panda bear is rolling around in the grass in a zoo enclosure . 1021632
Reposted by Alisa LiuVivek Kalyan @vivekkalyan.com · 24/11/2024Reading the TÜLU 3 paper from @ai2.bsky.social. It's refreshing to see a research lab treating AI as a real science with full reports, data, code, logs, evals. Paper: allenai.org/papers/tulu-... Demo: playground.allenai.org Code: github.com/allenai/open... Eval: github.com/allenai/olmes Notesallenai.org 1255
Reposted by Alisa LiuAi2 @ai2.bsky.social · 21/11/2024Meet Tülu 3, a set of state-of-the-art instruct models with fully open data, eval code, and training algorithms. We invented new methods for fine-tuning language models with RL and built upon best practices to scale synthetic instruction and preference data. Demo, GitHub, paper, and models 👇 211131
Reposted by Alisa LiuNathan Lambert @natolambert.bsky.social · 21/11/2024I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread. 821343
Reposted by Alisa LiuAkari Asai @akariasai.bsky.social · 19/11/20241/ Introducing ᴏᴘᴇɴꜱᴄʜᴏʟᴀʀ: a retrieval-augmented LM to help scientists synthesize knowledge 📚 @uwnlp.bsky.social & Ai2 With open models & 45M-paper datastores, it outperforms proprietary systems & match human experts. Try out our demo! openscholar.allen.ai 616339