Sign in

Cong Lu

@cong-ml.bsky.social
959 followers 463 following 11 posts

Research Scientist @ Google DeepMind, in open-ended learning, and AI for Scientific Discovery.

PostsRepliesMedia
Cong Lu @cong-ml.bsky.social · 11/06/2025
More major advantages! 🌟 COST-EFFECTIVE: StochasTok allows enhanced subword skills to be seamlessly 'retrofitted' into existing pretrained models - thus avoiding costly pretraining! ENHANCED ROBUSTNESS: Improves resilience to alternative tokenizations! (see examples) [6/]
110
Cong Lu @cong-ml.bsky.social · 11/06/2025
Empirically, we find: LANGUAGE: As hoped, StochasTok unlocks language manipulation ability! (see task examples below) MATH: Furthermore, StochasTok dramatically changes multi-digit addition, enabling grokking and even generalization to UNSEEN TOKENIZERS!🤯 [5/]
100
Cong Lu @cong-ml.bsky.social · 11/06/2025
The underlying StochasTok algorithm is extremely simple! 1️⃣ Simply tokenize text with ANY base tokenizer, 2️⃣ Then, stochastically split some of those tokens into equivalent token pairs. That’s basically it! Repeat step 2 for the desired granularity. [3/]
100
Cong Lu @cong-ml.bsky.social · 11/06/2025
🤔The problem: Standard tokenization gives distinct token IDs for each token - making it unnecessarily hard to learn, e.g., ‘book’=3092 and ‘cook’=171691 differ by a single letter. 🎉The solution: Allow LLMs to naturally 'see inside' tokens via alternative tokenizations! [2/]
100
Cong Lu @cong-ml.bsky.social · 11/06/2025
🚀Introducing “StochasTok: Improving Fine-Grained Subword Understanding in LLMs”!🚀 LLMs are incredible but still struggle disproportionately with subword tasks, e.g., for character counts, wordplay, multi-digit numbers, fixing typos… Enter StochasTok, led by Anya Sims! [1/]
132
Cong Lu @cong-ml.bsky.social · 12/12/2024
Interested in robust model-based offline RL algorithms? Come check out Anya Sims presenting our new paper investigating the edge of reach problem in offline MBRL! 📍East Exhibit Hall A-C #4603 #NeurIPS2024
010