Reposted by Tian JinGintare Karolina Dziugaite @gkdziugaite.bsky.social · 22/04/2025Excited to share our research on what matters in sparse LLM pre-training. Stop by our poster @ ICLR 🗓️ April 24th session #2. 181
Reposted by Tian JinSuvinay @suvinay.bsky.social · 22/04/2025Scaling Laws provide a valuable lens in guiding model design and computational budgets. Our recent work extends this lens to the realm of _fine-grained_ sparsity. Check out our #ICLR2025 paper, and the thread below from lead-author @tjin.bsky.social summarizing our findings. 021
Reposted by Tian JinDan Roy @roydanroy.bsky.social · 21/04/2025Tian and Karolina and team are at ICLR. Come say hi. 081
Tian Jin @tjin.bsky.social · 21/04/2025📣 The Journey Matters: Our #ICLR2025 paper shows how to pretrain sparse LLMs with half the size of dense LLMs while maintaining quality. We found that the average parameter count during sparse pre-training predicts quality, not final size. An MIT/Rice/Google/ISTA collab 🧵 1/N 140
Tian Jin @tjin.bsky.social · 27/02/2025Excited to share our work with friends from MIT/Google on Learned Asynchronous Decoding! LLM responses often contain chunks of tokens that are semantically independent. What if we can train LLMs to identify such chunks and decode them in parallel, thereby speeding up inference? 1/N 1179