Sign in

Gaël Varoquaux

@gaelvaroquaux.bsky.social
14K followers 216 following 662 posts

Research & code: Research director @inria ►Data, Health, & Computer science ►Python coder, (co)founder of scikit-learn, joblib, & @probabl.bsky.social ►Sometimes does art photography ►Physics PhD

PostsRepliesMedia
Gaël Varoquaux @gaelvaroquaux.bsky.social · 01/10/2026
Some biggies in this release: ■ The CatEncoder can be super useful ■ Caching makes exploratory work so much more productive And many places got faster (always nice to have) or easier to use. Enjoy more powerful and easier learning with dataframes :)
041
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
🧠 TextEncoder is now LLMEncoder -- clearer name, same transformer embeddings
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
📊 DataOp reports now link to your source code and show docstrings -- can be generated without executing calculations
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
⚡ The SessionEncoder is now up to 15x faster
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
🔍 describe_transformations() -- can now print a plain text explanation of what TableVectorizer and Cleaner did to each column
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
💾 Persistent caching for DataOps -- pipeline steps can now be cached across runs
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
🆕 CatEncoder -- OneHotEncoder + TargetEncoder, two encoders combined in a single column transformer to deal with columns with many infrequent classes
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
🫂13 new contributors helped with this release!
001
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
Also in this release: faster construction of deep DataOps, ToCategorical for numeric columns, and more!
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/10/2026
✨ Skrub 0.11 has been released ✨ Full changelog: skrub-data.org/stable/CHANG... Heads-up: skrub 0.11 requires Python ≥ 3.11 and scikit-learn ≥ 1.5.2. Highlights in thread ⤵️
121
Gaël Varoquaux @gaelvaroquaux.bsky.social · 01/10/2026
Tomorrow, I'm on the Vanishing Gradients podcast to talk about agentic science with @hugobowne.bsky.social, Shipra Arora from Bain & Company, and Luca Fiaschi from @pymc-labs.bsky.social Online live panel, with QA! 🗓️ Fri Oct 2, 1pm CEST / 7am EST ➡️ Register: numfocus-org.zoom.us/webinar/regi...
numfocus-org.zoom.us
Welcome! You are invited to join a webinar: The State of Agentic Data Science: From Hype to Real-World Impact. After registering, you will receive a confirmation email about joining the webinar.
Agentic AI is beginning to reshape how data science is done, from analytical workflows and experimentation to model development and decision-making. But where does the technology stand today, and what...
0112
Gaël Varoquaux @gaelvaroquaux.bsky.social · 22/09/2026
Super excited about our Survey with Lihu Chen on the Role of Small Models in the LLM Era This topic is crucial in today's world direct.mit.edu/coli/article...
direct.mit.edu
What is the Role of Small Models in the LLM Era: A Survey
Abstract. Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning tasks, which leads to the development of increasingly large models. However, scaling up model size...
0167
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
More details about those changes and other fixes in the changelog: github.com/joblib/threa...
github.com
threadpoolctl/CHANGES.md at master · joblib/threadpoolctl
Python helpers to limit the number of threads used in native libraries that handle their own internal threadpool (BLAS and OpenMP implementations) - joblib/threadpoolctl
101
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
The README now also includes a section about how to empirically assess the subtle semantic variations of various BLAS and OpenMP runtimes: github.com/joblib/threa...
github.com
threadpoolctl/README.md at master · joblib/threadpoolctl
Python helpers to limit the number of threads used in native libraries that handle their own internal threadpool (BLAS and OpenMP implementations) - joblib/threadpoolctl
101
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
The documentation in the README of the project has been updated to explain how to achieve this. github.com/joblib/threa...
github.com
threadpoolctl/README.md at master · joblib/threadpoolctl
Python helpers to limit the number of threads used in native libraries that handle their own internal threadpool (BLAS and OpenMP implementations) - joblib/threadpoolctl
101
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
Mitigating this problem is needed to unlock the full value of free-threading Python, especially for @scikit-learn.org workloads that often nest BLAS calls (via NumPy, SciPy or PyTorch) and OpenMP calls (via Cython) under Python level threads (typically via joblib).
101
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
Oversubscription problems typically happen when nesting BLAS or OpenMP calls under Python threads: naively spawning 10 Python threads that themselves spawn 10 BLAS threads each results in 100 starving threads on a 10 cores CPU.
101
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
This release includes several contributions by itamarst.hachyderm.io.ap.brid.gy from @quansight.com. in collaboration with myself & others at @probabl.ai. It provides tools to inspect the semantics of native threadpools in various environments so as to be able to mitigate oversubscription problems.
itamarst.hachyderm.io.ap.brid.gy
212
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 16/09/2026
threadpoolctl 3.7.0 is out. This library is a utility used by @scikit-learn.org and others to coordinate the levels of thread-based parallelism in native libraries such as BLAS implementations and OpenMP runtimes used by Python libraries such as NumPy, SciPy, PyTorch and scikit-learn.
2146
Gaël Varoquaux @gaelvaroquaux.bsky.social · 13/09/2026
Bretagne sur Orge
071
Reposted by Gaël Varoquaux
Sean Carroll @seanmcarroll.bsky.social · 10/09/2026
My thoughts on existential risk: * Humanity isn't going to be wiped out any time soon. * Negligible chance that AI by itself causes catastrophic damage to humanity (millions dead). * Some nontrivial chance that human beings will leverage AI to help them do something catastrophically harmful.
2930047
Gaël Varoquaux @gaelvaroquaux.bsky.social · 06/09/2026
AI Safety & Industry Thinking 📖: a personal anecdote The recent cyber-security screw-ups with AI-mediated attacks on third parties reminded me of a discussion a few years ago... Story time! 👇
171
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 01/09/2026
🐞 Bug fix release incoming 🐛 Version 0.10.1 of skrub is now live! In this release, we fixed a Cloudpickle import error that could be triggered by some updated dependencies, along with some other improvements. Release post: github.com/skrub-data/s... 1/4
github.com
Release 0.10.1 · skrub-data/skrub
🌟 Skrub 0.10.1 has been released 🌟 Main features When a cross-validation splitter has been passed to .skb.mark_as_X(), it it now possible to select a specific split from it when calling .skb.train...
134
Reposted by Gaël Varoquaux
Erin J. Wamsley 🧠📈😴 @wamsleylab.bsky.social · 01/09/2026
Oh. my. god. The biggest scientific publishers are collecting nearly $4 BILLION A YEAR in author fees. This is absurd. We've got to break free of this system. If everyone would just agree to it, we could just post all our new papers on bsky and it wouldn't cost anything. Or something. But not this.
24319
Reposted by Gaël Varoquaux
jamilahmedai.bsky.social @jamilahmedai.bsky.social · 01/09/2026
Stanford's visiting scholar forums or the SUPost classifieds are usually great for short October sublets around Palo Alto/Menlo Park. Reposting to help spread the word.
111
Gaël Varoquaux @gaelvaroquaux.bsky.social · 31/08/2026
🏠 Acomodation around Stanford: One of my very brillant students is visiting Stanford for a few weeks in October (dates to be defined). Would people have suggestions for housing for her that wouldn't be affordable? She is very reliable and I can absolutely recommend her. Thanks
372
Reposted by Gaël Varoquaux
Gaël Varoquaux @gaelvaroquaux.bsky.social · 08/08/2026
Les lois contraignant les véhicules lourds pour protéger les véhicules légers (stationnement, distance de déplacement) ne sont pas appliquées, des forces de l'ordre refusant de verbaliser. Mais on invente de nouvelles lois contraignant les véhicules légers. ⚠ aux verbalisations à sens unique
184
Gaël Varoquaux @gaelvaroquaux.bsky.social · 08/08/2026
Les lois contraignant les véhicules lourds pour protéger les véhicules légers (stationnement, distance de déplacement) ne sont pas appliquées, des forces de l'ordre refusant de verbaliser. Mais on invente de nouvelles lois contraignant les véhicules légers. ⚠ aux verbalisations à sens unique
184
Reposted by Gaël Varoquaux
Compute! Paris @computeparis.bsky.social · 07/08/2026
At #ComputeParis, Riccardo Cappuzzo will show how to build complex preprocessing pipelines with skrub's Data Ops module, showcasing it for varying examples from churn prediction over energy usage forecasting, to image processing. skrub-data.org/stable/ compute.events/paris2026 🛸
compute.events
Compute! Paris 2026
Understanding your tools and algorithms is what gives you agency in a digital world. Paris, November 25–26, 2026.
011
Reposted by Gaël Varoquaux
Compute! Paris @computeparis.bsky.social · 07/08/2026
Tabular data powers a huge share of machine learning applications, bringing with it the challenge of turning fragmented, messy datasets into reliable features. The Python library #skrub for "Machine learning with dataframes", part of the widely used scikit-learn ecosystem, tackles this challenge.
skrub-data.org
Skrub
Machine learning with dataframes
221
Reposted by Gaël Varoquaux
Christophe Michel @christopheml.fr · 06/08/2026
"Oui mais c'est bien de porter un casque", oui mais ça ne veut pas dire qu'il faut le rendre obligatoire pour autant. Ce que de nombreux pays plus en avance que nous sur le vélo ont montré, c'est que ce qui réduit les accidents ce sont les infrastructures et un meilleur partage de la voirie.
37620
Reposted by Gaël Varoquaux
Christophe Michel @christopheml.fr · 06/08/2026
100% une décision faite pour plaire aux boomers (leur autre grande marotte étant d'immatriculer les vélos et de leur imposer un permis), qui coûte 0 à l’État, permet de faire plus de répression stupide (comme l'interdiction des écouteurs) et n'améliore qu'à la marge la sécurité des cyclistes.
1415472
Reposted by Gaël Varoquaux
Ben Recht @beenwrekt.bsky.social · 24/07/2026
Now that the rich and powerful people are calling for open-weights models, it's time to pressure them to be legitimately brave and back open-corpus models.
2589
Gaël Varoquaux @gaelvaroquaux.bsky.social · 16/07/2026
Materials for my course on machine learning for health data: runnable examples of ML on health data, to get people thinking about model flexibility and biases gael-varoquaux.info/health_ml_tu...
gael-varoquaux.info
An introduction to Machine Learning for health and epidemiology
Understand concepts important to Machine Learning in Health with notebooks that run on real health data, giving the practical elements to tackle the complexity of real statistical learning question...
2266
Reposted by Gaël Varoquaux
CAD (Collectif Accès au Droit) @cad-asso.bsky.social · 04/07/2026
📢 Contre la présomption de légitime défense pour les forces de l'ordre ! La France est déjà le pays d'Europe comptant le plus grand nombre de personnes tuées par des agents de la force publique. Ce texte aggravera ce bilan. ✒️Signez la pétition : petitions.assemblee-nationale.fr/initiatives/...
petitions.assemblee-nationale.fr
Contre la présomption de légitime défense pour les forces de l'ordre. - Contre la présomption de légitime défense pour les forces de l'ordre. - Plateforme des pétitions de l’Assemblée nationale
Le 7 juillet 2026, l'Assemblée Nationale est appelée à se prononcer sur la proposition de loi n°691, portée par le député Eric Pauget (LR), visant à reconnaître une présomption de légitime défense pou...
052
Gaël Varoquaux @gaelvaroquaux.bsky.social · 02/07/2026
A fun podcast with Maxime Gabella about our dreams about AI. The next frontier of AI is building agents that use statistical learning to tackle problem-specific challenges from data. Agents come with promises, but they need a statistical harness, as @probabl.ai's www.youtube.com/watch?v=cNpv...
youtube.com
Can AI Become a Real Data Scientist? | Gaël Varoquaux on scikit-learn, Probabl & Scientific Judgment
YouTube video by Maxime Gabella
0104
Reposted by Gaël Varoquaux
ELLIS @ellis.eu · 30/06/2026
🏹 Job alert: Postdoc on Tabular Foundation Models at Inria 📍 Palaiseau 🇫🇷 ⏰ 20 August 🔗 bit.ly/4eN1j0N
075
Gaël Varoquaux @gaelvaroquaux.bsky.social · 29/06/2026
🧑‍💻🧑‍🏫 I'm recruiting a post-doc to work on Tabular Foundation Models, one of the hotest topics in AI, where we are at the leading edge team.inria.fr/soda/files/2... This is an opportunity to develop the next-level tabular AI, blending deep learning and tables.
team.inria.fr
33320
Reposted by Gaël Varoquaux
Manuel Mendoza @longchrom.bsky.social · 22/06/2026
This article in Le Monde about French research in general and the CNRS in particular has a few depressing graphs www.lemonde.fr/sciences/art...
The graph highlights a growing divergence between Germany and France:

* In the mid-1990s, France and Germany spent similar shares of GDP on R&D.
* Germany subsequently increased its R&D investment by about 1 percentage point of GDP.
* France’s R&D effort remained largely stagnant.
* By 2023, Germany spends 3.15% of GDP on R&D compared with 2.18% in France, a gap of nearly 1 percentage point of GDP.
* France is also well below the OECD average (2.93%).

The figure therefore illustrates that France has fallen behind both Germany and the OECD average in R&D intensity over the past three decades.

Sources listed on the graphic: OECD database, UNESCO, ANR, ERC; infographic by Le Monde.
46062
Gaël Varoquaux @gaelvaroquaux.bsky.social · 22/06/2026
You're only a real statisticien if you can say "Heteroscedasticity" without flinching
2151
Reposted by Gaël Varoquaux
Ultimes scories @panettonepazzo.bsky.social · 21/06/2026
Vous avez systématiquement voté contre tout projet de loi écolo. Vous avez stigmatisé les militants en « eco-terroristes ». Vous avez coupé les crédits verts des budgets. Vous avez nommé des ministres pro énergies fossiles et pesticides. Vous avez parlé «d’écologie punitive». Ne venez pas pleurer
Capture d’écran d’un post X du HuffPost avec article et avec ce texte :

« On n’a pas été assez loin et assez vite » : La canicule oblige les macronistes à un rare mea culpa »
551286645
Reposted by Gaël Varoquaux
scikit-learn @scikit-learn.org · 12/06/2026
🎉 Scikit-learn 1.9 released: ■Solid improvements to many existing estimators: faster, more stable, handling missing values, adding GPU support… ■Also, enhanced estimator displays in notebooks, ■And callbacks that enable progress bars or monitoring of convergence blog.scikit-learn.org/updates/rele...
blog.scikit-learn.org
scikit-learn release 1.9: better numerics, new core functionality
Author: Gael Varoquaux
03013
Reposted by Gaël Varoquaux
Olivier Grisel @ogrisel.bsky.social · 04/06/2026
The CfP deadline for Compute! Paris 2026 was extended to Sunday, June 7! Just a few days left to submit a proposal on Open Source scientific compute, data science, ML & AI topics. Conference dates and venue: November 25–26, 2026, Sorbonne Université · Paris compute.events/paris2026/cf...
compute.events
Call for Proposals — Compute! Paris 2026
Submit your talk proposal for Compute! Paris 2026. The Call for Proposals is open from April 15th to June 7th, 2026.
075
Gaël Varoquaux @gaelvaroquaux.bsky.social · 26/05/2026
Modern AIs tackle different questions than data science on industry or scientific applications. Statistical thinking, front and center in data science, is often hidden in AI. But it’s just as crucial. And too often, we treat data science or AI as merely a programming exercise
1279
Reposted by Gaël Varoquaux
NeurIPS Europe @neuripseurope.bsky.social · 17/05/2026
FAQ on NeurIPS Europe: NeurIPS Europe is an official NeurIPS 2026 satellite event taking place in Paris, France, alongside the main conference in Sydney and the other satellite event in Atlanta. NeurIPS authors can present their papers at any of the three locations, subject to space availability.
25218
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 06/05/2026
- It is now possible to pass arguments to the scorers in Data Ops, such as sample weights. - Diagrams for the Learner and parameter searches now include the full DataOp graph in their notebook repr. - It is now possible to find nodes by name in the DataOp graph.
121
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 06/05/2026
- fuzzy_join and Joiner now allow to choose the metric that should be used for matching. - ApplyToCols now has the exclude_cols parameter, to define which columns should not be transformed.
001
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 06/05/2026
- The TableReport now uses plot_distributions and compute_associations to control the distribution and association tabs respectively. - The cleaner now allows to control whether numeric-looking strings ("['1', '2', '3']") should be parsed to float.
101
Reposted by Gaël Varoquaux
Skrub @skrub-data.bsky.social · 06/05/2026
✨ Skrub version 0.9.0 has been released ✨ This release adds some advanced features to the Data Ops, the has_dtype() selector, as well as some clarity improvements for the Cleaner and TableReport. Release post: github.com/skrub-data/s...
github.com
Release 0.9.0 · skrub-data/skrub
✨ Skrub version 0.9.0 has been released ✨ Main changes Scorers used by Data Ops can now take additional arguments (like sample weights). By @jeromedockes in #1995 The new methods .skb.find() and ....
133
Gaël Varoquaux @gaelvaroquaux.bsky.social · 24/04/2026
#ICLR2026 paper✨️: Quantifying epistemic uncertainty of Blackbox classifiers, and link to better decisions Calibration on steroids, qualifying full prediction uncertainty with no need for Bayes, and tuning individual decisions 👇
1293