Houjun Liu @jemoka.com · 14/09/2026To make claims like the above^, theseus logs MFU, throughput calculations and also have a click-to-profile which gives you an xprof trace on the click of a button from a website printed at the beginning of training. 110
Houjun Liu @jemoka.com · 14/09/2026DDP, Zero-1, FSDP, and TP are reimplemented; models you add get them for free as long as you annotate logical axes of your parameters 100
Houjun Liu @jemoka.com · 14/09/2026and you can take your declared job and scale it up or down to remote clusters, with one-click dispatch to SLURM, Volcano K8s, and plain SSH hosts. this command is even idempotent to resumes across hardware up to floating-point differences :) 100
Houjun Liu @jemoka.com · 14/09/2026adding a dataset is indeed as simple as it sounds, and theseus provides you with parallel tokenization, document packing, etc. for free 100
Houjun Liu @jemoka.com · 14/09/2026along the way the library buys you a lot of power for free, like 𝙩𝙞𝙢𝙚 𝙩𝙧𝙖𝙫𝙚𝙡 𝙙𝙚𝙗𝙪𝙜𝙜𝙞𝙣𝙜, where you can ask theseus to find and replay the exact batch, activation, and intermediate of any model you've trained and attach evaluations 100
Houjun Liu @jemoka.com · 14/09/2026I always found it a bit tough to do the subclassing HF thing for architectures thing, so I spent the last year trying to avoid doing that. something simple like trying to change a tiny component of a GPT shouldn't take more than 4 lines of code... 100
Houjun Liu @jemoka.com · 14/09/2026🚨 new package day!! 🚨 hi friends, I'm open-sourcing the system I use to do architectures research: github.com/Jemoka/thx 100
Houjun Liu @jemoka.com · 02/10/2025Better yet, without us teaching the model to do this at all, it learned to allocate more compute at tokens of higher entropy (even as measured by an independently trained model of the same architecture), and use less compute where there's either too little or too much entropy. 🤯 110
Houjun Liu @jemoka.com · 02/10/2025By just using our approach, you don't have to do any extra work to get pretraining gains! We show across scale AND computation match that our approach performs better in pretraining perplexity than both regular transformers and manually inserting non-adaptive thinking tokens. 🥳 110
Houjun Liu @jemoka.com · 02/10/2025We design an transformer variant that uses a score-attenuated "forking" mechanism to clone useful residuals the model wants to update and attend to, thus creating a 𝗯𝘂𝗯𝗯𝗹𝗲 of latent computation for those highly-informative tokens. 130
Houjun Liu @jemoka.com · 02/10/2025Introducing 𝘁𝗵𝗼𝘂𝗴𝗵𝘁𝗯𝘂𝗯𝗯𝗹𝗲𝘀: a *fully unsupervised* LM for input-adaptive parallel latent reasoning ✅ Learn yourself a reasoning model with normal pretraining ✅ Better perplexity compared to fixed thinking tokens No fancy loss, no chain of thought labels 🚀 173
Houjun Liu @jemoka.com · 20/08/2025Even across baseline methods, low-perplexity prompts result in more effective attacks, but optimizing for attack success alone results in high-perplexity prompts. 100
Houjun Liu @jemoka.com · 20/08/2025In fact, our method allows us to discover a Pareto tradeoff (🤯) between attack success and prompt likelihood; tuning a single parameter in our method travels along the Pareto-optimal front. 100
Houjun Liu @jemoka.com · 20/08/2025Using the Adaptive Stress Testing (AST) framework as a reward signal for an online DPO-based optimization, we present a method to discover **both** high-probability prompts that are also successful in attacks. 100
Houjun Liu @jemoka.com · 20/08/2025New Paper Day! For EMNLP findings—in LM red-teaming, we show you have to optimize for **both** perplexity and toxicity for high-probability, hard to filter, and natural attacks! 162
Houjun Liu @jemoka.com · 02/06/2025Through a 🤏 pinch of interp, we show that model editing success gets degraded by pretraining with dropout. Dispersed representations built by dropout => less consistent representation of the world => worse models. 100
Houjun Liu @jemoka.com · 02/06/2025BERTs and encoder models are not saved from this either, with MLM and SQuAD performance being degraded by just turning on 10% dropout. 100
Houjun Liu @jemoka.com · 02/06/2025This stays true BOTH 1) at scale 2) with early dropout, which is supposed to be a way to stabilize convergence. 100
Houjun Liu @jemoka.com · 02/06/2025We show that applying dropout in pretraining kneecaps the models, even with downstream finetuning. 110
Houjun Liu @jemoka.com · 02/06/2025New Paper Day! For ACL 2025 Findings: You should **drop dropout** when you are training your LMs AND MLMs! 1114
Houjun Liu @jemoka.com · 14/02/2025the world if I could spell "causal interventions" correctly on the first try 120