Sign in

Felix Draxler

@drrelax.bsky.social
20 followers 5 following 9 posts

Postdoc @ UCI with Stephan Mandt. PhD @ Heidelberg University. Generative models, Parallel Token Prediction, Free-Form Flows, Normalizing Flows

PostsRepliesMedia
Felix Draxler @drrelax.bsky.social · 03/07/2026
I’m honored to present CBG at #icml! Come stop by at the morning 7/7 posters (Hall A #2701). Paper and code dgeyfman.github.io/calibrated-g.... Thanks to my amazing collaborators: Daniel Geyfman, Jan Groeneveld, Hyunsoo Lee, @tkaraletsos.bsky.social, and Stephan Mandt.
dgeyfman.github.io
000
Felix Draxler @drrelax.bsky.social · 03/07/2026
On a Bayesian inference benchmark, CBG converges to the correct calibrated distribution while other methods clearly remain biased with more compute. CBG also sets a SOTA on the InverseBench black-hole imaging task, and scales up to natural images.
110
Felix Draxler @drrelax.bsky.social · 03/07/2026
CBG takes diffusion steps using samples from a fast generative model, avoiding approximations existing methods use which introduce unavoidable bias. It comes in gradient-free and gradient-based variants that both become more calibrated with more compute.
100
Felix Draxler @drrelax.bsky.social · 03/07/2026
AI for science needs calibrated models and uncertainty, but we show standard test-time diffusion guidance methods miss the Bayesian posterior. Our Calibrated Bayesian Guidance (CBG) fixes this, enabling Bayesian inference and setting a new SOTA for black-hole imaging.
Our CBG fits a posterior right, baselines (DPS, NDTM) are biased.
110
Felix Draxler @drrelax.bsky.social · 22/04/2026
Huge thanks to amazing collaborators: Justus, Farrin, Theo, Sameer and Stephan! Meet us at #ICLR: Apr 23, morning poster P3-#608
000
Felix Draxler @drrelax.bsky.social · 22/04/2026
Parallelize the model you care about with our code: github.com/mandt-lab/ptp
github.com
GitHub - mandt-lab/ptp: Parallel Token Prediction for Language Models (ICLR 2026)
Parallel Token Prediction for Language Models (ICLR 2026) - mandt-lab/ptp
110
Felix Draxler @drrelax.bsky.social · 22/04/2026
We verify that this works in a speculative decoding experiment: We distill Vicuna-7B on conversations to predict the same output, but at greatly reduced latency. We achieve a 2.4x speedup over AR on diverse text tasks on one GPU, with 3.2x possible with optimized implementation.
110
Felix Draxler @drrelax.bsky.social · 22/04/2026
Autoregressive models produce text by predicting the histogram of the next token. Random auxiliary variables determine the token to choose. Parallel Token Prediction directly learns what external randomness maps to which token. This allows to predict many tokens at once.
110
Felix Draxler @drrelax.bsky.social · 22/04/2026
LLMs are autoregressive and slow? No! Parallel Token Prediction decodes multiple consistent tokens in one model call. PTP allows arbitrary dependencies in one call, unlike discrete diffusion. Practical: 2.4x speedup github.com/mandt-lab/ptp ICLR: Apr 23, morning poster P3-#608
1295