Sign in

Houjun Liu

@jemoka.com
299 followers 611 following 70 posts

NLP & POMDPs; CS@Stanford; gradient descent enthusiast www: jemoka.com ac: nlp.stanford.edu/~houjun/

PostsRepliesMedia
Houjun Liu @jemoka.com · 14/09/2026
Let me know if you end up using it; I'd love to hear about it :) Its MIT licensed for any desired shenanigans. Thanks to Pratyusha Sharma, Kaden Zhang, and Tianle Yu for contributions and feedback; thanks to Róbert Csordás' tooling for inspiration. theseus.jemoka.com
theseus.jemoka.com
theseus
Train, evaluate, and dispatch language model jobs with minimal boilerplate.
000
Houjun Liu @jemoka.com · 14/09/2026
This project has been used in stuff like our paper Thoughtbubbles (arxiv.org/abs/2510.00219), and it made the whole effort a lot easier.
100
Houjun Liu @jemoka.com · 14/09/2026
To make claims like the above^, theseus logs MFU, throughput calculations and also have a click-to-profile which gives you an xprof trace on the click of a button from a website printed at the beginning of training.
110
Houjun Liu @jemoka.com · 14/09/2026
Measuring a reimplementation of Qwen2.5 3B gets ~7,186 toks/s/GPU FSDPx2 block size 512 on A100 NVLink, which depending on how you do the math is around 40% MFU which I think is pretty sick. Mileage may vary and also the topology, I've hit 20s on multi-host H100 stacks
100
Houjun Liu @jemoka.com · 14/09/2026
yes there are (lots hand written!) docs
100
Houjun Liu @jemoka.com · 14/09/2026
DDP, Zero-1, FSDP, and TP are reimplemented; models you add get them for free as long as you annotate logical axes of your parameters
100
Houjun Liu @jemoka.com · 14/09/2026
and you can take your declared job and scale it up or down to remote clusters, with one-click dispatch to SLURM, Volcano K8s, and plain SSH hosts. this command is even idempotent to resumes across hardware up to floating-point differences :)
100
Houjun Liu @jemoka.com · 14/09/2026
adding a dataset is indeed as simple as it sounds, and theseus provides you with parallel tokenization, document packing, etc. for free
100
Houjun Liu @jemoka.com · 14/09/2026
along the way the library buys you a lot of power for free, like 𝙩𝙞𝙢𝙚 𝙩𝙧𝙖𝙫𝙚𝙡 𝙙𝙚𝙗𝙪𝙜𝙜𝙞𝙣𝙜, where you can ask theseus to find and replay the exact batch, activation, and intermediate of any model you've trained and attach evaluations
100
Houjun Liu @jemoka.com · 14/09/2026
I always found it a bit tough to do the subclassing HF thing for architectures thing, so I spent the last year trying to avoid doing that. something simple like trying to change a tiny component of a GPT shouldn't take more than 4 lines of code...
100
Houjun Liu @jemoka.com · 14/09/2026
🚨 new package day!! 🚨 hi friends, I'm open-sourcing the system I use to do architectures research: github.com/Jemoka/thx
100
Reposted by Houjun Liu
Naomi Saphra @nsaphra.bsky.social · 11/11/2025
meet me at this button friends
refresh button for address bar visiting openreview.net
3181
Reposted by Houjun Liu
Claas Voelcker @cvoelcker.bsky.social · 11/11/2025
LMAO, openreview down point 9pm UCT when ICLR is supposed to be releasing. Coincidence?
041
Reposted by Houjun Liu
Houjun Liu @jemoka.com · 02/10/2025
Introducing 𝘁𝗵𝗼𝘂𝗴𝗵𝘁𝗯𝘂𝗯𝗯𝗹𝗲𝘀: a *fully unsupervised* LM for input-adaptive parallel latent reasoning ✅ Learn yourself a reasoning model with normal pretraining ✅ Better perplexity compared to fixed thinking tokens No fancy loss, no chain of thought labels 🚀
173
Houjun Liu @jemoka.com · 02/10/2025
I'm really excited about this. Because this model is trained with literally nothing but LM loss, it helps create a new reasoning paradigm where reasoning capabilities are baked right in at pretraining, unifying train and test time behaviors. Look ma, no distribution shift! 🙏
000
Houjun Liu @jemoka.com · 02/10/2025
Better yet, without us teaching the model to do this at all, it learned to allocate more compute at tokens of higher entropy (even as measured by an independently trained model of the same architecture), and use less compute where there's either too little or too much entropy. 🤯
110
Houjun Liu @jemoka.com · 02/10/2025
By just using our approach, you don't have to do any extra work to get pretraining gains! We show across scale AND computation match that our approach performs better in pretraining perplexity than both regular transformers and manually inserting non-adaptive thinking tokens. 🥳
110
Houjun Liu @jemoka.com · 02/10/2025
We design an transformer variant that uses a score-attenuated "forking" mechanism to clone useful residuals the model wants to update and attend to, thus creating a 𝗯𝘂𝗯𝗯𝗹𝗲 of latent computation for those highly-informative tokens.
130
Houjun Liu @jemoka.com · 02/10/2025
Current approaches in scaling inference-time compute require supervising with explicit chain-of-thought data, which limits thoughts to be sequential and in human language only. 😔 Wouldn't it be nice if you can do normal pretraining, and somehow get latent thinking for free? 🤔
110
Houjun Liu @jemoka.com · 02/10/2025
Joint work with my wonderful collaborators @shikharmurty.bsky.social, @robertcsordas.bsky.social, and @chrmanning.bsky.social. Paper: arxiv.org/abs/2510.00219. Code and Package: github.com/stanfordnlp/....
arxiv.org
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
Current approaches for scaling inference-time compute in transformers rely on training them to emit explicit chain-of-thought tokens before producing an answer. While these methods are powerful, they…
110
Houjun Liu @jemoka.com · 02/10/2025
Introducing 𝘁𝗵𝗼𝘂𝗴𝗵𝘁𝗯𝘂𝗯𝗯𝗹𝗲𝘀: a *fully unsupervised* LM for input-adaptive parallel latent reasoning ✅ Learn yourself a reasoning model with normal pretraining ✅ Better perplexity compared to fixed thinking tokens No fancy loss, no chain of thought labels 🚀
173
Reposted by Houjun Liu
Houjun Liu @jemoka.com · 20/08/2025
New Paper Day! For EMNLP findings—in LM red-teaming, we show you have to optimize for **both** perplexity and toxicity for high-probability, hard to filter, and natural attacks!
162
Houjun Liu @jemoka.com · 20/08/2025
Thanks to @schmidtsciences.bsky.social and Lambda Labs for generously supporting our work :)
000
Houjun Liu @jemoka.com · 20/08/2025
🤔 think this is all too much? No worries, we are also dropping a **PACKAGE** to do this for you. Check it out: github.com/sisl/astra-rl
github.com
GitHub - sisl/astra-rl: The Adaptive Stress Testing for Robust AI (ASTRA) toolbox provides tooling to support model developers and testing in the full life cycle of making more robust AI Systems through the application of adaptive stress testing and adversarial training.
The Adaptive Stress Testing for Robust AI (ASTRA) toolbox provides tooling to support model developers and testing in the full life cycle of making more robust AI Systems through the application of...
100
Houjun Liu @jemoka.com · 20/08/2025
☝️ And so.... You should optimize for **BOTH** attack success and perplexity to get the most effective attacks!
100
Houjun Liu @jemoka.com · 20/08/2025
Even across baseline methods, low-perplexity prompts result in more effective attacks, but optimizing for attack success alone results in high-perplexity prompts.
100
Houjun Liu @jemoka.com · 20/08/2025
In fact, our method allows us to discover a Pareto tradeoff (🤯) between attack success and prompt likelihood; tuning a single parameter in our method travels along the Pareto-optimal front.
100
Houjun Liu @jemoka.com · 20/08/2025
Using the Adaptive Stress Testing (AST) framework as a reward signal for an online DPO-based optimization, we present a method to discover **both** high-probability prompts that are also successful in attacks.
100
Houjun Liu @jemoka.com · 20/08/2025
Most approaches in gradient-based red-teaming result in very low-probability prompts, which previous work have shown are both easier to filter and bad negative examples for downstream hardening.
100
Houjun Liu @jemoka.com · 20/08/2025
Done at Stanford Intelligent Systems Laboratory — my joint first author Amelia Hardy, along with our wonderful collaborators Allie Griffith, @bernardlange.bsky.social, Duncan Eddy, Mykel Kochenderfer. Paper: arxiv.org/pdf/2407.09447 Python package to do this for yourself: github.com/sisl/astra-rl
arxiv.org
100
Houjun Liu @jemoka.com · 20/08/2025
New Paper Day! For EMNLP findings—in LM red-teaming, we show you have to optimize for **both** perplexity and toxicity for high-probability, hard to filter, and natural attacks!
162
Reposted by Houjun Liu
FOCS 2026 @focs2026.bsky.social · 13/07/2025
The list of accepted papers at #FOCS2025 is up! focs.computer.org/2025/accepte...
focs.computer.org
Accepted Papers – FOCS 2025
03615
Reposted by Houjun Liu
Haskell programming language @haskell.org · 21/06/2025
You're not too dumb for Haskell, you just need a reason to practice. :)
0143
Reposted by Houjun Liu
Journal of Open Source Software @joss-openjournals.bsky.social · 03/07/2025
Just published in JOSS: 'Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers' doi.org/10.21105/joss.08183
011
Reposted by Houjun Liu
Martin Kleppmann @martin.kleppmann.com · 27/06/2025
OCaml @ocaml.org is in The Economist!
4549
Reposted by Houjun Liu
TTIC @tticconnect.bsky.social · 27/06/2025
We’re proud to announce three new tenure-track assistant professors joining TTIC in Fall 2026: Yossi Gandelsman, Will Merrill, and Nick Tomlin (@nickatomlin.bsky.social). Meet them here: buff.ly/JH1DFtT
072
Reposted by Houjun Liu
Mathurin Massias @mathurinmassias.bsky.social · 18/06/2025
New paper on the generalization of Flow Matching www.arxiv.org/abs/2506.03719 🤯 Why does flow matching generalize? Did you know that the flow matching target you're trying to learn *can only generate training points*? w @quentinbertrand.bsky.social @annegnx.bsky.social @remiemonet.bsky.social 👇👇👇
25617
Reposted by Houjun Liu
Houjun Liu @jemoka.com · 02/06/2025
New Paper Day! For ACL 2025 Findings: You should **drop dropout** when you are training your LMs AND MLMs!
1114
Houjun Liu @jemoka.com · 02/06/2025
I think there's a very cool training dynamics situation going on here; if you are pretraining with webtext and driving loss down to 0, disbursed representations matter a LOT less since your training corpus already regularizes decently.
000
Houjun Liu @jemoka.com · 02/06/2025
Through a 🤏 pinch of interp, we show that model editing success gets degraded by pretraining with dropout. Dispersed representations built by dropout => less consistent representation of the world => worse models.
100
Houjun Liu @jemoka.com · 02/06/2025
BERTs and encoder models are not saved from this either, with MLM and SQuAD performance being degraded by just turning on 10% dropout.
100
Houjun Liu @jemoka.com · 02/06/2025
This stays true BOTH 1) at scale 2) with early dropout, which is supposed to be a way to stabilize convergence.
100
Houjun Liu @jemoka.com · 02/06/2025
We show that applying dropout in pretraining kneecaps the models, even with downstream finetuning.
110
Houjun Liu @jemoka.com · 02/06/2025
Though most frontier shops already skips dropout in their biggest models, but BERTs, smaller LMs, and FTing still are trained with plenty of dropout (sometimes up to 30% 😱)
100
Houjun Liu @jemoka.com · 02/06/2025
Work done with John Bauer and ‪@chrmanning.bsky.social‬ Paper: arxiv.org/pdf/2505.24788
100
Houjun Liu @jemoka.com · 02/06/2025
New Paper Day! For ACL 2025 Findings: You should **drop dropout** when you are training your LMs AND MLMs!
1114
Reposted by Houjun Liu
Kristina Gligoric @IC2S2 @gligoric.bsky.social · 30/05/2025
I'm excited to announce that I’ll be joining the Computer Science department at Johns Hopkins as an Assistant Professor this Fall! I’ll be working on large language models, computational social science, and AI & society—and will be recruiting PhD students. Apply to work with me!
14914
Reposted by Houjun Liu
Sky @sky.app · 28/05/2025
Introducing Sky for Mac. Watch this 90-second sneak peek:
55611
Houjun Liu @jemoka.com · 21/05/2025
looks like deepmind has been busy deepmind.google/models/veo/ deepmind.google/models/gemin...
000
Houjun Liu @jemoka.com · 17/05/2025
@radbuglet.bsky.social since when did you become a user of bluesky?
000