Sign in

Michael Hu

@michahu.bsky.social
69 followers 128 following 6 posts

PhD student at NYU. NLP & training data. michahu.github.io

PostsRepliesMedia
Michael Hu @michahu.bsky.social · 26/11/2024
Boavista album by Stephan Bodzin: open.spotify.com/track/7ujvbI...
open.spotify.com
Nothing Like You
Stephan Bodzin, Luna Semara · Boavista · Song · 2021
020
Michael Hu @michahu.bsky.social · 26/11/2024
Is this #1 in your Spotify wrapped 😆
010
Michael Hu @michahu.bsky.social · 19/11/2024
thanks for featuring this work!
010
Michael Hu @michahu.bsky.social · 12/11/2024
In joint work with @MayeeChen @NickLourie @kchonyc @HazyResearch, we use our optimization framework to analyze failures of existing methods. We then turn these insights into: Aioli 🧄, a fully-online data mixing algorithm! paper: arxiv.org/abs/2411.05735 code: github.com/HazyResearch...
arxiv.org
Aioli: A Unified Optimization Framework for Language Model Data Mixing
Language model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to efficiently learn mixture ...
000
Michael Hu @michahu.bsky.social · 12/11/2024
So you want a good pretraining data mix🧑‍🍳, but which data mixing algorithm do you pick? DoGE, DoReMi, Skill-it, grid searching proportions… 😵‍💫 It turns out that these algorithms are all special cases of Linear Mixing Optimization, our new data mixing framework! 🧵
100
Michael Hu @michahu.bsky.social · 10/11/2024
metropolis-hastings: 1️⃣ sample from your proposal function 2️⃣ run the sample through your filter, proportional to the desired pdf 3️⃣ use the kept samples to initialize the next round i wonder if we can connect iterative approaches to synthetic data as making specific choices in an MCMC framework...
000