Sign in

Tiancheng Hu

@tiancheng.bsky.social
1K followers 1.1K following 92 posts

PhD student @CambridgeLTL; Previously @DLAB @EPFL; Interested in NLP and CSS. Apple Scholar, Gates Scholar.

PostsRepliesMedia
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Paper: arxiv.org/abs/2601.06407 Huge congrats to @riverdong.bsky.social for leading this work, together with Zheng Hui, Caiqi Zhang, Ivan Vulić, Nigel Collier, and special thanks to Andreea Bobu. I’ll present the paper as an oral on Tuesday, July 7, 11:00–12:30, in Harbor B-C. Come say hi!
arxiv.org
Value of Information: A Framework for Human-Agent Communication
Large Language Model (LLM) agents deployed for real-world tasks face a fundamental dilemma: user requests are underspecified, yet agents must decide whether to act on incomplete information or interru...
010
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Overall, VoI provides a parameter-free, inference-time framework for adaptive agent communication. The goal is not just to build agents that execute tasks. It is to build agents that communicate thoughtfully, knowing when to ask, when to act, and how to respect user effort.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Even with imperfect calibration, VoI remains competitive with manually tuned baselines.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Why does VoI work? A key step is estimating a belief distribution over latent user states. We analyze calibration across tasks: models are reasonably calibrated in Animal Guessing, but less so in Medical Diagnosis, where symptoms are noisier and more ambiguous.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
A useful way to think about it: Clarification is not just about uncertainty. It is about whether reducing uncertainty will actually improve the final decision enough to justify interrupting the user. That is why VoI accounts for: • ambiguity • task risk • user effort rather than confidence alone.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Why does communication cost matter? When asking is cheap, agents should gather more information. When asking is expensive, they should stop earlier. Fixed-round methods often over-ask or under-ask. Confidence thresholds can work, but only after brittle manual tuning. VoI adapts to the stated cost.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
VoI matches or exceeds the best manually tuned baselines in 18/20 settings, with gains up to +1.36 utility in high-cost settings. Unlike confidence thresholding, VoI does not require task-specific threshold tuning. It adapts to the cost-benefit tradeoff automatically.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
We evaluate VoI across four human-agent communication settings: 🎮 20 Questions 🩺 Medical diagnosis ✈️ Flight recommendation 🛒 Ambiguous WebShop These tasks vary in ambiguity, risk, and user effort cost: exactly the tradeoff agents face in real deployments.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
If positive: ask. Otherwise: commit. This turns clarification into an explicit cost-benefit decision.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
Core idea: ask only when the expected information gain is worth the user effort. For each candidate clarification question, VoI compares: Expected utility if the agent acts now Expected utility after possible user answers Communication cost of asking Net value = expected improvement − user cost.
100
Tiancheng Hu @tiancheng.bsky.social · 30/06/2026
[#ACL2026 Paper Alert] Real-world requests are often underspecified. So when should an AI agent ask a clarification question? Not always. Not never. We introduce Value of Information (VoI): a decision-theoretic framework for deciding when to ask, when to act, and when to stop.
120
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
Reasoning Boosts Opinion Alignment in LLMs From Frédéric Berdoz, Yann Billeter, Yann Vonlanthen, Roger Wattenhofer
000
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
Related ICLR papers: What Do Large Language Models Know About Opinions? From Erfan Jahanparast Zhiqing Hong @serinachang5.bsky.social Benchmarking Overton Pluralism in LLMs From @elinorpd.bsky.social Jiayi Wu @taylor-sorensen.bsky.social Jiaxin Pei @mbakker.bsky.social
110
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
Check out the paper and data for details! Paper: arxiv.org/abs/2510.17516 Data: huggingface.co/datasets/pit... Website: simbench.tiancheng.hu See you Fri, Apr 24, 2026 • 3:15 PM – 5:45 PM @ Pavilion 4 P4-#5206
arxiv.org
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current ev...
110
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
Another result that really stood out to me and kept me thinking: Post-training helps LLMs predict consensus opinions, but hurts their ability to model pluralistic disagreement. So making a model a better chat assistant may make it a worse social simulator. How should we train a good simulator?
100
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
Evaluation in this area has been pretty fragmented. If we want LLM social simulation to become scientifically useful, we need a shared way to measure when, how, and why models succeed or fail. The overall picture is still far from perfect: the best model we tested at release scored 40.8 / 100.
100
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
This is the motivation behind SimBench: a benchmark for group-level human behavior simulation with LLMs. It brings together 20 datasets spanning moral dilemmas, economic games, psych assessments, and more, so this can be studied in a standardized way rather than through isolated one-off tasks.
120
Tiancheng Hu @tiancheng.bsky.social · 16/04/2026
SimBench now at #ICLR2026! Often in social simulations, the goal is not to predict what one specific person will do. It is to estimate how a group will respond, whether in pre-testing a real polling question, or in stress-testing a policy or intervention before running it in the real world.
141
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
It was a great fun working on this with Caiqi @dirkhovy.bsky.social Nigel @cambridgeltl.bsky.social @milanlp.bsky.social
010
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
7/7 As models become ever so widely deployed in the real world, we need to build models that are not just capable, but genuinely reliable. Let us know what you think! What LLM failure do you think is most underappreciated as a calibration problem? Paper: www.techrxiv.org/doi/full/10....
techrxiv.org
Position: Large Language Model Failures from Hallucination to Homogenization Are Different Facets of Miscalibration
This position paper argues that diverse failures of Large Language Models (LLMs), from confident hallucinations to collapsed diversity to brittle safety refusals, are best understood as different face...
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
Deploy — design interfaces that let uncertainty reach the humans acting on model outputs.
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
6/7 Our call to action: Measure — make calibration metrics standard in benchmark reporting. Leaderboards track accuracy; almost none track calibration. Train — base models start well-calibrated. We need alignment methods that don't destroy that.
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
5/7 The field treats each failure with a bespoke patch: RAG for hallucination, temperature for diversity, safety classifiers for over-refusal. But these work around the model's broken uncertainty rather than fixing it. That's why the same failures keep resurfacing in new forms.
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
4/7 We argue these aren't separate bugs. They're four facets of the same problem: 🔴 Probabilistic — can't match requested distributions 🟠 Semantic — confidence ≠ correctness 🔵 Distributional — output diversity collapse 🟢 Metacognitive — can't assess its own competence
121
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
3/7 It goes deeper than randomness. Ask it "what book should I read?" and it defaults to the same WEIRD-centric bestsellers. Ask a nuanced political question and it responds with near-zero variation, hallucinating consensus where none exists. It can spend 1,000 tokens "reasoning" about 2+3.
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
2/7 Try this: ask any LLM for a random number between 1 and 10. Models prefer "7" much more often than anything else. It can describe probability correctly when you ask. It just can't do it.
110
Tiancheng Hu @tiancheng.bsky.social · 02/03/2026
1/7 🧵 The GPT-4 technical report featured detailed calibration curves. Since then, not a single major model release has reported calibration. The field quietly stopped measuring whether models know what they don't know. Our new position paper argues this is a mistake. Here's why.
192
Tiancheng Hu @tiancheng.bsky.social · 04/02/2026
A privilege to represent @cambridgeltl.bsky.social @camlangsci.bsky.social @gatescambridge.bsky.social Huge thanks to the entire team, the Secretariat, the Expert Advisory Panel and all reviewers.
010
Tiancheng Hu @tiancheng.bsky.social · 04/02/2026
A crucial challenge is evaluation: existing evaluation methods do not reliably reflect how systems perform in real-world settings. Nonetheless, with companies investing hundreds of billions to scale up, we expect model capabilities to continue growing.
100
Tiancheng Hu @tiancheng.bsky.social · 04/02/2026
Yet, the frontier remains "jagged": models may still fail on simple tasks, like counting objects in an image. Their performance also tends to decline when prompted in languages other than English, which has major implications for global deployment and fairness.
100
Tiancheng Hu @tiancheng.bsky.social · 04/02/2026
I focused on "Current Capabilities." We documented rapid advances: AI now achieves Gold-medal performance at the Math Olympiad and agents are increasingly automating useful work, from software engineering to curriculum design.
102
Tiancheng Hu @tiancheng.bsky.social · 04/02/2026
Proud to contribute to the new International AI Safety Report chaired by @YoshuaBengio, with a fantastic international team! Every word was weighed to ensure a rigorous, evidence-based view of current AI capabilities and the risks they pose. A short summary of my section below.
100
Reposted by Tiancheng Hu
Yoshua Bengio @yoshuabengio.bsky.social · 25/11/2025
I’m pleased to share the Second Key Update to the International AI Safety Report, which outlines how AI developers, researchers, and policymakers are approaching technical risk management for general-purpose AI systems. (1/6)
22910
Tiancheng Hu @tiancheng.bsky.social · 31/10/2025
Personalization certainly needs boundaries and we show how that could look like!
000
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
Great fun working on this with @bminixhofer.bsky.social and Prof. Collier at @cambridgeltl.bsky.social. Special thanks to Paul Martin, and Arcee AI's Mergekit library.
010
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
TL;DR: The alignment-calibration trade-off is real, but you don't have to be stuck with the endpoints. Model merging provides a simple, powerful dial to find the perfect balance of capability and reliability for YOUR application. Paper here: arxiv.org/abs/2510.17426 (8/8)
arxiv.org
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model output...
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
Better calibration has benefits beyond accuracy scores. It helps reduce "mode collapse" in generation tasks, leading to more diverse generations (and higher utility too), as measured on NoveltyBench. It improves model performance on group-level simulation tasks too! (7/8)
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
And it gets better with scale! 📈 The benefits of merging, both the accuracy boost and the stability of the "sweet spot", become even more pronounced in larger, more capable models. This echoes prior work which shows merging bigger models are more effective and stable. (6/8)
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
The Pareto-superior frontier is a general phenomenon we observe across model families (Gemma, Qwen), sizes, and datasets, where we can consistently find a better-balanced model. We show Qwen 2.5 results on BBH and MMLU-Pro below. (5/8)
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
It's NOT a zero-sum game between base and instruct. We find a "sweet spot" merge that is Pareto-superior: it has HIGHER accuracy than both parents while substantially restoring the calibration lost during alignment. (4/8)
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
Our solution is simple and computationally cheap: model merging. By interpolating between the well-calibrated base model and its capable but overconfident instruct counterpart, we create a continuous spectrum to navigate this trade-off. No retraining needed. (3/8)
100
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
Let's start by redefining the problem. We argue the "alignment tax" MUST include the severe loss of model calibration. Instruction tuning doesn't just nudge performance; it wrecks calibration, causing a huge spike in overconfidence. (2/8)
120
Tiancheng Hu @tiancheng.bsky.social · 30/10/2025
Instruction tuning unlocks incredible skills in LLMs, but at a cost: they become dangerously overconfident. You face a choice: a well-calibrated base model or a capable but unreliable instruct model. What if you didn't have to choose? What if you could navigate the trade-off? (1/8)
121
Tiancheng Hu @tiancheng.bsky.social · 29/10/2025
River, Yinhong and I will all be in person and we look forward to the discussions!
031
Reposted by Tiancheng Hu
Ona de Gibert @onadegibert.bsky.social · 28/10/2025
See you next week at EMNLP! We will be presenting our work: Scaling Low-Resource MT via Synthetic Data Generation with LLMs 📍 Poster Session 13 📅 Fri, Nov 7, 10:30-12:00 - Hall C 📖 Check it out! arxiv.org/abs/2505.14423 @helsinki-nlp.bsky.social @cambridgenlp.bsky.social @emnlpmeeting.bsky.social
arxiv.org
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
We investigate the potential of LLM-generated synthetic data for improving low-resource Machine Translation (MT). Focusing on seven diverse target languages, we construct a document-level synthetic co...
082
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
Huge thanks to my amazing collaborators @joachimbaumann.bsky.social @Lorenzo Lupo @nigelcollier.bsky.social @dirkhovy.bsky.social and especially @paul-rottger.bsky.social @cambridgeltl.bsky.social Work partially done during my visit to @milanlp.bsky.social. Highly recommended!
020
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
Check out the paper and data for details! Paper: arxiv.org/abs/2510.17516 Data: huggingface.co/datasets/pit... Website: simbench.tiancheng.hu (9/9)
141
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
Overall, by making progress measurable, SImBench provides the foundation to build more faithful LLM simulators. Moving forward, we should work on better training strategies for improving LLM social simulators. These will most likely diverge from advances in chat / coding models. (8/9)
120
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
We find simulation ability correlates most strongly with deep, knowledge-intensive general reasoning (MMLU-Pro, r=0.94), rather than competition math (AIME, r=0.48) To simulate humans well, a model needs a broad, nuanced understanding of the world. (7/9)
120