Sign in

Hanlin Zhang

@hlzhang109.bsky.social
34 followers 45 following 24 posts

CS PhD candidate @Harvard hanlin-zhang.com

PostsRepliesMedia
Hanlin Zhang @hlzhang109.bsky.social · 08/07/2026
Presenting our work at the #ICML26 Oral session about Prescriptive Scaling. We study a practical question for language models: given a pre-training compute budget, what downstream performance can we reliably expect after post-training? 📍 #ICML Oral Hall C 🕧 Thu, Jul 9, 4:30 PM–4:45 PM Local Time
121
Hanlin Zhang @hlzhang109.bsky.social · 17/06/2026
Announcing our COLM 2026 workshop: Scientific Understanding of Foundation Models: we invite submissions on training dynamics, scaling laws, data and optimization, post-training, reward modeling, evaluation science, reliability, reproducibility, and theoretical understanding of foundation models.
100
Hanlin Zhang @hlzhang109.bsky.social · 18/03/2026
Learning from feedback is instrumental but human preference data can be expensive. How much reward supervision could we get from raw web text instead, without human labels? Our work, a year of incredible effort by @fjxdaisy.bsky.social, advances pure RLHF training across multiple models and tasks.
193
Hanlin Zhang @hlzhang109.bsky.social · 02/07/2025
Introducing EvoLM, a model suite with 100+ decoder-only LMs (1B/4B) trained from scratch, across four training stages — 🟦 Pre-training 🟩 Continued Pre-Training (CPT) 🟨 Supervised Fine-Tuning (SFT) 🟥 Reinforcement Learning (RL)
arxiv.org
EvoLM: In Search of Lost Language Model Training Dynamics
Modern language model (LM) training has been divided into multiple stages, making it difficult for downstream developers to evaluate the impact of design choices made at each stage. We present EvoLM, ...
131
Hanlin Zhang @hlzhang109.bsky.social · 18/06/2025
New work [JSKZ25] w/ Jikai, Vasilis, @shamkakade.bsky.social . We introduce new formulations and tools for evaluating LM capabilities, which help explain observations of post-training behaviors of Qwen-series models. More details: - hanlin-zhang.com/causal-capab... - x.com/_hanlin_zhan...
000
Hanlin Zhang @hlzhang109.bsky.social · 23/04/2025
Highlights from #ICLR2025 — a brief thread 🧵
110
Reposted by Hanlin Zhang
Andreas Kirsch @blackhc.bsky.social · 05/04/2025
I want to reshare @brandfonbrener.bsky.social's @NeurIPSConf 2024 paper on CoLoR-Filter: A simple yet powerful method for selecting high-quality data for language model pre-training! With @hlzhang109.bsky.social @schwarzjn.bsky.social @shamkakade.bsky.social
2188
Reposted by Hanlin Zhang
Sham Kakade @shamkakade.bsky.social · 22/11/2024
(1/n) 💡How can we speed up the serial runtime of long pre-training runs? Enter Critical Batch Size (CBS): the tipping point where the gains of data parallelism balance with diminishing efficiency. Doubling batch size halves the optimization steps—until we hit CBS, beyond which returns diminish.
2174
Reposted by Hanlin Zhang
Yuda Song @yus167.bsky.social · 06/12/2024
LLM self-improvement has critical implications in synthetic data, post-training and test-time inference. To understand LLMs' true capability of self-improvement, we perform large-scale experiments with multiple families of LLMs, tasks and mechanisms. Here is what we found: (1/9)
1124