Sign in

Sarthak Mittal

@sarthmit.bsky.social
222 followers 24 following 11 posts
PostsRepliesMedia
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
🤯 Unexpected Finding: Continuous-time diffusion models underperform in posterior estimation—sometimes worse than simple Gaussian assumptions! This highlights the need for better model design for parameter estimation. 🚀 Open-sourced code: github.com/sarthmit/par...
020
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
For full posterior estimation, we explore forward- and reverse-KL minimization (+combining them) with various modeling choices 🔹 Gaussian approx. 🔹 Normalizing Flows 🔹 Advanced models: Diffusion, Flow-Matching, Iterated Denoising Energy Matching Surprising insight next!
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
🚀 Key Finding: In high-dimensional spaces, amortized point estimation significantly outperforms full posterior approaches! For point estimation, we use: 🔹 Maximum Likelihood (MLE) 🔹 Maximum-a-Posteriori (MAP) But what about posterior estimation?
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
We study amortized inference, where a learner estimates the underlying parameters in its forward pass, explicitly conditioned on data. Through extensive in- and out-of-distribution evaluations, we compare point estimation vs. full posterior estimation.
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
🔍 Parametric Inference: Point vs Full Posterior Estimation Two approaches: 📌 Point Estimation (MLE/MAP) – Optimizes for a single parameter value 📊 Full Posterior Estimation – Approximates the full distribution (MCMC, VI) Which is best for amortized inference? We find out! 👇
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
🚀 New Preprint! 🚀 In-Context Parametric Inference: Point or Distribution Estimators? Thrilled to share our work on inferring probabilistic model parameters explicitly conditioned on data, in collab with @yoshuabengio.bsky.social, Nikolay Malkin & @glajoie.bsky.social! 🔗 arxiv.org/abs/2502.11617
arxiv.org
In-Context Parametric Inference: Point or Distribution Estimators?
Bayesian and frequentist inference are two fundamental paradigms in statistical estimation. Bayesian methods treat hypotheses as random variables, incorporating priors and updating beliefs via Bayes' ...
142
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
It definitely took us a while to get this out but we are excited to provide a thorough and rigorous evaluation benchmark! Exploring in-context learning approaches agnostic to modality is under-explored and we are very excited about this avenue! Code: github.com/sarthmit/par...
000
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
We provide a rigorous comparison of different architecture choices, parameterizations of densities, and training objectives for learning this amortized in-context posterior estimator. Further studies on high-dimensional problems, cases of misspecification, etc. in the paper!
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
Same insight, beyond language: estimate p(parameters | dataset) for different datasets. What do you gain? 🌟 Posterior over parameters for new datasets provided in-context through just inference instead of MCMC, etc. Fun connections to learned optimizers, meta-learning, etc.
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
Diffusion: p(image | text) for different text inputs ICL in LLMs: p(ans | question, examples) for different examples Multi-task (RL or otherwise) = p(next action | environment) for different environments Key insight: Train across diverse contexts using a shared language.
100
Sarthak Mittal @sarthmit.bsky.social · 28/02/2025
🚨 New Preprint! 🚨 We explore Amortized In-Context Bayesian Posterior Estimation with Niels, @glajoie.bsky.social, Priyank Jaini & @marcusabrubaker.bsky.social ! 🔥 Amortized Conditional Modeling = key to success in large-scale models! We use it to estimate posteriors 🔑 📄 arxiv.org/abs/2502.06601
arxiv.org
Amortized In-Context Bayesian Posterior Estimation
Bayesian inference provides a natural way of incorporating prior beliefs and assigning a probability measure to the space of hypotheses. Current solutions rely on iterative routines like Markov Chain ...
121