Sign in

David Holzmüller

@dholzmueller.bsky.social
800 followers 152 following 171 posts

Postdoc in machine learning with Francis Bach & @GaelVaroquaux: neural networks, tabular data, uncertainty, active learning, atomistic ML, learning theory. dholzmueller.github.io

PostsRepliesMedia
Reposted by David Holzmüller
Gaël Varoquaux @gaelvaroquaux.bsky.social · 29/06/2026
🧑‍💻🧑‍🏫 I'm recruiting a post-doc to work on Tabular Foundation Models, one of the hotest topics in AI, where we are at the leading edge team.inria.fr/soda/files/2... This is an opportunity to develop the next-level tabular AI, blending deep learning and tables.
team.inria.fr
33320
Reposted by David Holzmüller
Olivier Grisel @ogrisel.bsky.social · 01/04/2026
Here is the recording of the webinar I gave last week on GPU support in @scikit-learn.org and comparison of a scikit-learn pipeline vs the TabICLv2 foundational model on a non-linear heteroscedastic quantile regression task. app.livestorm.co/probabl/webi...
app.livestorm.co
[Webinar] Python array API support in scikit-learn for GPU acceleration and TabICLv2 | Probabl
142
David Holzmüller @dholzmueller.bsky.social · 24/02/2026
Maybe. But both should just converge to the true posterior with infinite pretraining scale.
010
David Holzmüller @dholzmueller.bsky.social · 24/02/2026
I tried 512 and 999 in small-scale experiments and there didn't seem to be a significant difference. I don't think a smaller choice would be better for small datasets, since we have a lot of small datasets in pretraining. But you never know...
110
Reposted by David Holzmüller
Ivan Rubachev @puhsu.bsky.social · 13/02/2026
To piggy-back a bit on foundation models for structured data discussion here My colleagues at Yandex Research just updated the GraphPFN paper. It's a Graph Foundation Model that works on graph datasets with tabular features, and shows SOTA results both in ICL regimes and when fine-tuned.
arxiv.org
GraphPFN: A Prior-Data Fitted Graph Foundation Model
Graph foundation models face several fundamental challenges including transferability across datasets and data scarcity, which calls into question the very feasibility of graph foundation models. Howe...
122
David Holzmüller @dholzmueller.bsky.social · 12/02/2026
Super hyped that it's finally out!
2151
Reposted by David Holzmüller
Gaël Varoquaux @gaelvaroquaux.bsky.social · 02/01/2026
My 2025 highlights for AI research and code: ▪ Unpacking the AI scale narrative ▪ Tabular-learning research - TabICL: table foundation model - Retrieve merge predict: data lakes ▪ Better software - Skrub: machine learning with tables - Fundamentals in scikit-learn gael-varoquaux.info/science/2025...
gael-varoquaux.info
2025 highlights: AI research and code
AI is everywhere. Can you see it here? Note Some highlights about my work in 2025: progress on tabular-learning stands out, a publication on unpacking trade-off and consequences of scale in...
0214
Reposted by David Holzmüller
arxiv stat.ML @arxiv-stat-ml.bsky.social · 15/12/2025
Sacha Braun, David Holzm\"uller, Michael I. Jordan, Francis Bach Conditional Coverage Diagnostics for Conformal Prediction arxiv.org/abs/2512.11779
031
Reposted by David Holzmüller
Judith Abécassis @judithabk6.bsky.social · 15/12/2025
Let's kick off 2026 with a workshop on Survival Analysis and Foundation Models, co-organized w. Julie Alberge Linus Bleistein Clément Berenfeld Agathe Guilloux and Julie Josse on January 27th at PariSanté Campus ! Registrations and submissions are open!!! www.linusbleistein.com/ramh
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
032
Reposted by David Holzmüller
Eugene Berta @eberta.bsky.social · 13/11/2025
Still using temperature scaling? With @dholzmueller.bsky.social, Michael I. Jordan and @bachfrancis.bsky.social we argue that with well designed regularization, more expressive models like matrix scaling can outperform simpler ones across calibration set sizes, data dimensions, and applications.
152
Reposted by David Holzmüller
Skrub @skrub-data.bsky.social · 26/09/2025
⚡ Release 0.6.2 is out ⚡ github.com/skrub-data/s...
github.com
Release 0.6.2 · skrub-data/skrub
New features The DataOp.skb.full_report() now displays the time each node took to evaluate. #1596 by Jérôme Dockès. The User guide has been reworked and expanded. Changes and deprecations Ken em...
174
Reposted by David Holzmüller
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 15/08/2025
Daniel Beaglehole, David Holzm\"uller, Adityanarayanan Radhakrishnan, Mikhail Belkin: xRFM: Accurate, scalable, and interpretable feature learning models for tabular data arxiv.org/abs/2508.10053 arxiv.org/pdf/2508.10053 arxiv.org/html/2508.10053
033
David Holzmüller @dholzmueller.bsky.social · 30/07/2025
Thanks!
000
David Holzmüller @dholzmueller.bsky.social · 29/07/2025
Solution write-up with additional insights: kaggle.com/competitions... 🥈2nd place used stacking with diverse models 🥇1st place found a larger dataset
kaggle.com
Prediction interval competition II: House price
Create a regression model for a house sale price having the narrowest overall prediction intervals
030
David Holzmüller @dholzmueller.bsky.social · 29/07/2025
I got 3rd out of 691 in a tabular kaggle competition – with only neural networks! 🥉 My solution is short (48 LOC) and relatively general-purpose – I used skrub to preprocess string and date columns, and pytabkit to create an ensemble of RealMLP and TabM models. Link below👇
2112
David Holzmüller @dholzmueller.bsky.social · 24/07/2025
Excited to have co-contributed the SquashingScaler, which implements the robust numerical preprocessing from RealMLP!
084
David Holzmüller @dholzmueller.bsky.social · 23/07/2025
Is it because mathematicians think in terms of the number of assumptions that are satisfied, while physicists think in terms of the number of things that satisfy them?
110
Reposted by David Holzmüller
Gaël Varoquaux @gaelvaroquaux.bsky.social · 09/07/2025
👨‍🎓🧾✨#icml2025 Paper: TabICL, A Tabular Foundation Model for In-Context Learning on Large Data With Jingang Qu, @dholzmueller.bsky.social, and Marine Le Morvan TL;DR: a well-designed architecture and pretraining gives best tabular learner, and more scalable On top, it's 100% open source 1/9
15115
Reposted by David Holzmüller
Lennart Purucker @lennartpurucker.bsky.social · 23/06/2025
🚨What is SOTA on tabular data, really? We are excited to announce 𝗧𝗮𝗯𝗔𝗿𝗲𝗻𝗮, a living benchmark for machine learning on IID tabular data with: 📊 an online leaderboard (submit!) 📑 carefully curated datasets 📈 strong tree-based, deep learning, and foundation models 🧵
1178
Reposted by David Holzmüller
Katharina Eggensperger @keggensperger.bsky.social · 20/06/2025
Missed the school? We have uploaded recordings of most talks to our YouTube Channel www.youtube.com/@AutoML_org 🙌
youtube.com
AutoML Freiburg Hannover Tübingen
This channel features videos about automated machine learning (AutoML) from the AutoML groups at the University of Freiburg, Leibniz University Hannover and University of Tübingen. Common topics inclu...
142
Reposted by David Holzmüller
Skrub @skrub-data.bsky.social · 28/05/2025
📝 The skrub TextEncoder brings the power of HuggingFace language models to embed text features in tabular machine learning, for all those use cases that involve text-based columns.
263
David Holzmüller @dholzmueller.bsky.social · 24/04/2025
Poster: iclr.cc/virtual/2025... Hall 3 + Hall 2B #32, 10am Singapore time Paper: arxiv.org/abs/2408.01536
arxiv.org
Active Learning for Neural PDE Solvers
Solving partial differential equations (PDEs) is a fundamental problem in science and engineering. While neural PDE solvers can be more efficient than established numerical solvers, they often require...
020
David Holzmüller @dholzmueller.bsky.social · 24/04/2025
🚨ICLR poster in 1.5 hours, presented by @danielmusekamp.bsky.social : Can active learning help to generate better datasets for neural PDE solvers? We introduce a new benchmark to find out! Featuring 6 PDEs, 6 AL methods, 3 architectures and many ablations - transferability, speed, etc.!
1102
Reposted by David Holzmüller
Skrub @skrub-data.bsky.social · 23/04/2025
The Skrub TableReport is a lightweight tool that allows to get a rich overview of a table quickly and easily. ✅ Filter columns 🔎 Look at each column's distribution 📊 Get a high level view of the distributions through stats and plots, including correlated columns 🌐 Export the report as html
164
Reposted by David Holzmüller
Gaël Varoquaux @gaelvaroquaux.bsky.social · 23/04/2025
#ICLR2025 Marine Le Morvan presents "Imputation for prediction: beware of diminishing returns": poster Thu 24th arxiv.org/abs/2407.19804 Concludes 6 years of research on prediction with missing values: Imputation is useful but improvements are expensive, while better learners yield easier gains.
1488
Reposted by David Holzmüller
Madelon Hulsebos @madelonhulsebos.bsky.social · 02/04/2025
Excited to share the new monthly Table Representation Learning (TRL) Seminar under the ELLIS Amsterdam TRL research theme! To recur every 2nd Friday. Who: Marine Le Morvan, Inria (in-person) When: Friday 11 April 4-5pm (+drinks) Where: L3.36 Lab42 Science Park / Zoom trl-lab.github.io/trl-seminar/
Details about the seminar talk titled TabICL: A Tabular Foundation Model for In-Context Learning on Large Data by Marine Le Morvan
0113
Reposted by David Holzmüller
Nick Erickson @nickerickson.bsky.social · 25/03/2025
We are excited to announce #FMSD: "1st Workshop on Foundation Models for Structured Data" has been accepted to #ICML 2025! Call for Papers: icml-structured-fm-workshop.github.io
01510
Reposted by David Holzmüller
Evan Peck @peck.phd · 20/11/2024
Trying something new: A 🧵 on a topic I find many students struggle with: "why do their 📊 look more professional than my 📊?" It's *lots* of tiny decisions that aren't the defaults in many libraries, so let's break down 1 simple graph by @jburnmurdoch.bsky.social 🔗 www.ft.com/content/73a1...
921577458
Reposted by David Holzmüller
Multiscale AI @ ICLR 2025 @multiscaleai.bsky.social · 18/03/2025
🚀Continuing the spotlight series with the next @iclr-conf.bsky.social MLMP 2025 Oral presentation! 📝LOGLO-FNO: Efficient Learning of Local and Global Features in Fourier Neural Operators 📷 Join us on April 27 at #ICLR2025! #AI #ML #ICLR #AI4Science
122
David Holzmüller @dholzmueller.bsky.social · 10/03/2025
Links: www.kaggle.com/competitions... www.kaggle.com/competitions... Link to the repo: github.com/dholzmueller... PS: The newest pytabkit version now includes multiquantile regression for RealMLP and a few other improvements. bsky.app/profile/dhol...
010
David Holzmüller @dholzmueller.bsky.social · 10/03/2025
Practitioners are often sceptical of academic tabular benchmarks, so I am elated to see that our RealMLP model outperformed boosted trees in two 2nd place Kaggle solutions, for a $10,000 forecasting challenge and a research competition on survival analysis.
kaggle.com
Rohlik Sales Forecasting Challenge
Use historical product sales data to predict future sales.
171
David Holzmüller @dholzmueller.bsky.social · 04/03/2025
The benchmark is limited to classification with AUC as a metric, which is one of RealMLP’s weaker points. Datasets are from the CC-18 benchmark, and the benchmark uses nested cross-validation unlike many other benchmarks. Link: arxiv.org/abs/2402.039...
arxiv.org
Is Deep Learning finally better than Decision Trees on Tabular Data?
Tabular data is a ubiquitous data modality due to its versatility and ease of use in many real-world applications. The predominant heuristics for handling classification tasks on tabular data rely on ...
010
David Holzmüller @dholzmueller.bsky.social · 04/03/2025
A new tabular classification benchmark provides another independent evaluation of our RealMLP. RealMLP is the best classical DL model, although some other recent baselines are missing. TabPFN is better on small datasets and boosted trees on larger datasets, though.
120
David Holzmüller @dholzmueller.bsky.social · 12/02/2025
What about work on adaptive learning rates (in the sense of convergence rates, not step sizes) that studies methods with hyperparameter optimization on a holdout set to achieve optimal/good convergence rates simultaneously for different classes of functions? E.g. projecteuclid.org/journals/ann...
000
David Holzmüller @dholzmueller.bsky.social · 08/02/2025
By the way, I think an intercept in this case is necessary because the logistic regression model does not have an intercept. For more realistic models that can learn an intercept themselves, I think an intercept for TS is probably not very important.
020
David Holzmüller @dholzmueller.bsky.social · 07/02/2025
Finally, if you just want to have the best performance for a given (large) time budget, AutoGluon combines many tabular models. It does not include some of the latest models (yet), but has a very good CatBoost, for example, and will likely outperform individual models.
010
David Holzmüller @dholzmueller.bsky.social · 07/02/2025
The library offers the same for XGBoost and LightGBM. Plus, the library includes some of the best tabular DL models like RealTabR, TabR, RealMLP, and TabM that could also be interesting to try. (ModernNCA is also very good but not included.)
110
David Holzmüller @dholzmueller.bsky.social · 07/02/2025
Using my library github.com/dholzmueller... you could, for example, use CatBoost_TD_Regressor(n_cv=5), which will use better default parameters for regression, train five models in a cross-validation setup, select the best iteration for each, and ensemble them.
100
David Holzmüller @dholzmueller.bsky.social · 07/02/2025
Interesting! Would be cool to have these datasets on OpenML as well so they are easy to use in tabular benchmarks. Here are some more recommendations for stronger tabular baselines: 1. For CatBoost and XGBoost, you'd want at least early stopping to select the best iteration.
100
Reposted by David Holzmüller
Fabian Schaipp @fschaipp.bsky.social · 05/02/2025
Learning rate schedules seem mysterious? Why is the loss going down so fast during cooldown? Turns out that this behaviour can be described with a bound from *convex, nonsmooth* optimization. A short thread on our latest paper 🚞 arxiv.org/abs/2501.18965
arxiv.org
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for the constant schedul...
2316
David Holzmüller @dholzmueller.bsky.social · 05/02/2025
github.com/EFS-OpenSour... has some calibration methods like this implemented, but their temperature scaling MLE version has a bug where it doesn't optimize, so I didn't include it in our benchmark.
010
David Holzmüller @dholzmueller.bsky.social · 05/02/2025
I think Dirichlet scaling (or the binary version Beta scaling) also includes an intercept but I'm not sure. In my experience it's very slow, though, and not better than temperature scaling at least on smaller datasets (~1K-10K calibration samples).
110
David Holzmüller @dholzmueller.bsky.social · 05/02/2025
There is an adapter for Dirichlet scaling, which is basically regularized matrix scaling. (Except that matrix scaling can exploit shifts in the logits, which a true post-hoc calibration method like Dirichlet scaling can't IIUC).
110
Reposted by David Holzmüller
Eugene Berta @eberta.bsky.social · 03/02/2025
Early stopping on validation loss? This leads to suboptimal calibration and refinement errors—but you can do better! With @dholzmueller.bsky.social, Michael I. Jordan, and @bachfrancis.bsky.social, we propose a method that integrates with any model and boosts classification performance across tasks.
4189
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
In case anyone is wondering about the name RealMLP, it is motivated by the “Real MVP” meme (which probably also inspired the RealNVP method). 6/6
010
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
The benchmark: arxiv.org/abs/2407.00956 RealMLP: github.com/dholzmueller... 5/ bsky.app/profile/dhol...
110
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
It is surprising how many DL methods perform worse than the simple MLP baseline by Gorishniy, @puhsu.bsky.social et al. This highlights the benchmarking problems in the field (and potentially the difficulty in using many of these models correctly). The situation is slowly improving. 4/
100
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
When including more baselines, RealMLP’s average rank slightly improves to make it the top-performing method overall, with a fifth place on binary classification, first place on multi-class, and second place on regression. 3/
100
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
Some caveats: All DL models are trained with a batch size of 1024, while we recommend using 256 for RealMLP on medium-sized datasets. Other choices (selection of datasets, not using bagging, choice of metrics, search spaces for baselines) can of course also influence results. 2/
100
David Holzmüller @dholzmueller.bsky.social · 16/01/2025
The first independent evaluation of our RealMLP is here! On a recent 300-dataset benchmark with many baselines, RealMLP takes a shared first place overall. 🔥 Importantly, RealMLP is also relatively CPU-friendly, unlike other SOTA DL models (including TabPFNv2 and TabM). 🧵 1/
Plots from the benchmark of Ye et al. (2024)
191