Sign in

Skrub

@skrub-data.bsky.social
648 followers 48 following 162 posts

skrub is a Python library to ease preprocessing and feature engineering for tabular machine learning. Our long-term goal is to directly connect database tables to machine learning estimators. skrub-data.org discord.gg/ABaPnm7fDC

PostsRepliesMedia
Skrub @skrub-data.bsky.social · 01/10/2026
🫂13 new contributors helped with this release!
001
Skrub @skrub-data.bsky.social · 01/10/2026
Also in this release: faster construction of deep DataOps, ToCategorical for numeric columns, and more!
101
Skrub @skrub-data.bsky.social · 01/10/2026
🧠 TextEncoder is now LLMEncoder -- clearer name, same transformer embeddings
101
Skrub @skrub-data.bsky.social · 01/10/2026
📊 DataOp reports now link to your source code and show docstrings -- can be generated without executing calculations
101
Skrub @skrub-data.bsky.social · 01/10/2026
⚡ The SessionEncoder is now up to 15x faster
101
Skrub @skrub-data.bsky.social · 01/10/2026
🔍 describe_transformations() -- can now print a plain text explanation of what TableVectorizer and Cleaner did to each column
101
Skrub @skrub-data.bsky.social · 01/10/2026
💾 Persistent caching for DataOps -- pipeline steps can now be cached across runs
101
Skrub @skrub-data.bsky.social · 01/10/2026
🆕 CatEncoder -- OneHotEncoder + TargetEncoder, two encoders combined in a single column transformer to deal with columns with many infrequent classes
101
Skrub @skrub-data.bsky.social · 01/10/2026
✨ Skrub 0.11 has been released ✨ Full changelog: skrub-data.org/stable/CHANG... Heads-up: skrub 0.11 requires Python ≥ 3.11 and scikit-learn ≥ 1.5.2. Highlights in thread ⤵️
131
Skrub @skrub-data.bsky.social · 01/09/2026
And long overdue, a Zenodo record for citing the library: doi.org/10.5281/zeno... 4/4
doi.org
skrub-data/skrub: 0.10.1
🌟 Skrub 0.10.1 has been released 🌟 Main features When a cross-validation splitter has been passed to .skb.mark_as_X(), it it now possible to select a specific split from it when calling .skb.train_test_split() by passing split_index; for example data_op.skb.train_test_split(split_index=3) to get the third split. #2213 @jeromedockes. Added support in tabular_pipeline() for estimators instantiated from either tabicl.TabICLClassifier or tabicl.TabICLRegressor with recommended default parameters of TableVectorizer as the first step, and the estimator as the second step. #2222 by @ashwinvis. Changes TextEncoder's verbose parameter is now an int instead of a bool, where verbose=0 silences the progress bar and verbose>=1 shows it. The default is now 0. #2249 by @Jayant-kernel. patch_display() now uses a minimal, faster TableReport without plots or column associations by default. #2103 by @ashiorkornortey. Bugfixes DropSimilar now works with Polars dataframes when PyArrow is not installed. #2216 by @ShreyanshGoyal. The parallel coordinate plot created by ParamSearch.show_results() could have incorrect tick labels in some cases. This has been fixed in #2215 by @jeromedockes. GapEncoder with init="k-means" raised an error when the input column contained missing values. This has been fixed in #2238 by @Hrafz. DatetimeEncoder no longer raises UnboundLocalError when resolution=None is combined with periodic_encoding="circular" or "spline", and no longer fits an unused weekday periodic encoder when resolution is finer than "hour". #2240 by @Hrafz. When cols was not provided, AggJoiner and MultiAggJoiner selected the columns to aggregate through a Python set, so the order of the aggregated output columns varied between runs. They now keep the order in which the columns appear in the auxiliary table. This has been fixed in #2250 by @dylanpulver. A cloudpickle import error that could happen after updating some required dependencies was fixed in #2261 by @jeromedockes and @rcap107. Full changelog: https://github.com/skrub-data/skrub/compare/0.10.0...0.10.1
010
Skrub @skrub-data.bsky.social · 01/09/2026
- Improved granular access to split when using the DataOps splitter. Can now select a specific split based on index (i.e. the ith split out of N splits) - More bugfixes! 3/4
100
Skrub @skrub-data.bsky.social · 01/09/2026
📋 Also included in this release: - Estimators from the Tabular In-Context Learning (TabICL) library are now accepted when running default pipelines in skrub, which the correct default parameters based on whether or not the estimator is a classifier or regressor 2/4
100
Skrub @skrub-data.bsky.social · 01/09/2026
🐞 Bug fix release incoming 🐛 Version 0.10.1 of skrub is now live! In this release, we fixed a Cloudpickle import error that could be triggered by some updated dependencies, along with some other improvements. Release post: github.com/skrub-data/s... 1/4
github.com
Release 0.10.1 · skrub-data/skrub
🌟 Skrub 0.10.1 has been released 🌟 Main features When a cross-validation splitter has been passed to .skb.mark_as_X(), it it now possible to select a specific split from it when calling .skb.train...
134
Skrub @skrub-data.bsky.social · 06/07/2026
Last but not least, **17** new contributors helped with this release 🎉
000
Skrub @skrub-data.bsky.social · 06/07/2026
More bugs have been squashed, and docs have been polished.
100
Skrub @skrub-data.bsky.social · 06/07/2026
DataOps have been improved with better caching, as well as easier access and search of nodes in the graph. It is also possible to filter which columns are shown in the parallel coordinate plot.
100
Skrub @skrub-data.bsky.social · 06/07/2026
The TableReport can now be exported in text format (Markdown formatting), or as a dictionary. In general, it is now easier to access the statistics measured by the TableReport in third party tools.
100
Skrub @skrub-data.bsky.social · 06/07/2026
ToFloat now includes decimal and thousand as parameters to parse numerical columns that use formatting different from the python default (such as "1 234,5"). Negative numbers represented with parentheses are also parsed ("(123)" becomes "-123").
100
Skrub @skrub-data.bsky.social · 06/07/2026
New transformers are now available: SessionEncoder groups timestamped data by sessions, DropSimilar helps with dropping redundant columns.
100
Skrub @skrub-data.bsky.social · 06/07/2026
✨ Skrub version 0.10.0 has been released ✨ This is one of our biggest releases yet, with new transformers, improvements to the TableReport and the Data Ops, bug fixes and polishing of the docs. github.com/skrub-data/s... 🚀 Highlights below
github.com
Release 0.10.0 · skrub-data/skrub
✨ Skrub version 0.10.0 has been released ✨ Main Changes New transformers are now available: SessionEncoder groups timestamped data by sessions, DropSimilar helps with dropping redundant columns. T...
120
Skrub @skrub-data.bsky.social · 06/05/2026
- fuzzy_join and Joiner now allow to choose the metric that should be used for matching. - ApplyToCols now has the exclude_cols parameter, to define which columns should not be transformed.
001
Skrub @skrub-data.bsky.social · 06/05/2026
- The TableReport now uses plot_distributions and compute_associations to control the distribution and association tabs respectively. - The cleaner now allows to control whether numeric-looking strings ("['1', '2', '3']") should be parsed to float.
101
Skrub @skrub-data.bsky.social · 06/05/2026
- It is now possible to pass arguments to the scorers in Data Ops, such as sample weights. - Diagrams for the Learner and parameter searches now include the full DataOp graph in their notebook repr. - It is now possible to find nodes by name in the DataOp graph.
121
Skrub @skrub-data.bsky.social · 06/05/2026
✨ Skrub version 0.9.0 has been released ✨ This release adds some advanced features to the Data Ops, the has_dtype() selector, as well as some clarity improvements for the Cleaner and TableReport. Release post: github.com/skrub-data/s...
github.com
Release 0.9.0 · skrub-data/skrub
✨ Skrub version 0.9.0 has been released ✨ Main changes Scorers used by Data Ops can now take additional arguments (like sample weights). By @jeromedockes in #1995 The new methods .skb.find() and ....
133
Skrub @skrub-data.bsky.social · 25/03/2026
The minimum required version of polars has been increased from 0.20 to 1.5.
021
Skrub @skrub-data.bsky.social · 25/03/2026
The TableReport custom filters have been improved and expanded: they can now take skrub selectors for filtering columns. The interface has also been simplified.
121
Skrub @skrub-data.bsky.social · 25/03/2026
The has_nulls selector can now select columns based on a user-specified threshold of null values.
121
Skrub @skrub-data.bsky.social · 25/03/2026
It is now possible to provide custom null values to the Cleaner, so that they are marked as nulls (for example, the string "unknown").
121
Skrub @skrub-data.bsky.social · 25/03/2026
The performance of DataOps with many computational nodes has been improved. Additionally, DataOps CV splitters can now take kwargs. For example, this allows to specify groups when creating train/test splits.
121
Skrub @skrub-data.bsky.social · 25/03/2026
The SingleColumnTransformer and RejectColumn classes allow the construction of custom-made transformers for specific use cases.
121
Skrub @skrub-data.bsky.social · 25/03/2026
The ApplyToCols transformer is now a powerful alternative to the regular scikit-learn ColumnTransformer. It is now possible to apply any transformer to a subset of chosen columns using the skrub selectors.
131
Skrub @skrub-data.bsky.social · 25/03/2026
✨ skrub version 0.8.0 has been released ✨ This version includes several new features, including multiple improvements to the functionality and performance of the Data Ops, along with a few bug fixes and improvements to the docs. Changelog: skrub-data.org/stable/CHANG... Highlights below ⤵️
skrub-data.org
Release history
Release 0.8.0: New Features: The eager_data_ops configuration option has been added. When set to False, no previews are computed and validation is deferred until the DataOp is actually used (e.g. w...
184
Skrub @skrub-data.bsky.social · 18/02/2026
You can contact us either here or on our Discord server: discord.gg/ABaPnm7fDC
discord.gg
Join the Skrub Discord Server!
Check out the Skrub community on Discord – hang out with 106 other members and enjoy free voice and text chat.
020
Skrub @skrub-data.bsky.social · 18/02/2026
In addition, we will begin crediting specific contributors here on Bluesky when a contributor has worked on the subject of the post. We will use GitHub handles for this purpose. If you prefer your handle not to be used or would like to be credited by name instead, please let us know.
100
Skrub @skrub-data.bsky.social · 18/02/2026
As a follow-up, we would like to clarify how we’ll be crediting contributors moving forward. Currently, all contributions to the repository are tracked in the changelog and highlighted in the release notes, where each PR and the GitHub handle of its author are listed.
100
Skrub @skrub-data.bsky.social · 18/02/2026
Thanks to e-strauss for writing this example!
000
Skrub @skrub-data.bsky.social · 18/02/2026
While skrub Data Ops shine when preparing dataframes, their capabilities extend beyond that. For example, they can be used alongside libraries like PyTorch and skorch to work with images, and tune the model size to find the best set of hyperparameters: skrub-data.org/stable/auto_...
skrub-data.org
Using PyTorch (via skorch) in DataOps
This example shows how to wrap a PyTorch model with skorch and plug it into a skrub DataOps plan. The main goal here is to show the integration pattern: PyTorch defines the model (an nn.Module), sk...
100
Skrub @skrub-data.bsky.social · 10/02/2026
- A new example has been added to show how skrub Data Ops can be used with pytorch and skorch to solve an image classification task. skrub-data.org/stable/auto_...
skrub-data.org
Using PyTorch (via skorch) in DataOps
This example shows how to wrap a PyTorch model with skorch and plug it into a skrub DataOps plan. The main goal here is to show the integration pattern: PyTorch defines the model (an nn.Module), sk...
001
Skrub @skrub-data.bsky.social · 10/02/2026
Main changes: - The StringEncoder now exposes the vocabulary parameter, allowing it to be passed to the underlying TfidfVectorizer. - The function compute_ngram_distance has been made private to reduce clutter. - The repository wheel has been made smaller by removing some benchmarking material.
111
Skrub @skrub-data.bsky.social · 10/02/2026
✨ skrub version 0.7.2 has been released ✨ In this release we squashed more bugs, improved the API reference, and added a new example. github.com/skrub-data/s...
github.com
Release Skrub release 0.7.2 · skrub-data/skrub
✨ skrub version 0.7.2 has been released ✨ In this release we squashed more bugs, improved the API reference, and added a new example. Main changes: The StringEncoder now exposes the vocabulary par...
121
Skrub @skrub-data.bsky.social · 05/02/2026
Here is a full example on how to use skrub Data Ops with Optuna skrub-data.org/stable/auto_...
skrub-data.org
Tuning DataOps with Optuna
This example shows how to use Optuna to tune the hyperparameters of a skrub DataOp. As seen in the previous example, skrub DataOps can contain “choices”, objects created with choose_from(), choose_...
001
Skrub @skrub-data.bsky.social · 05/02/2026
At the end, you get a fully-fledged Optuna study to work with. Of course, that includes support for the Optuna dashboard and access to the Optuna reporting and plotting interfaces.
101
Skrub @skrub-data.bsky.social · 05/02/2026
Did you know that the skrub Data Ops support Optuna as backend to run hyperparameter search? It's as easy as writing "backend='optuna'": this will set up a default Optuna study (and the TPE sampler) to replace the standard random sampler.
Three snippets of python code showing how to use skrub Data Ops with the Optuna optimization library.The first snippet shows a standard randomized search with the Data Ops. The second snippet adds the parameter "backend", which is set to "optuna". The third snippet uses the Optuna visualization API to plot information from the study.
142
Skrub @skrub-data.bsky.social · 15/01/2026
Happy new year! 🎉🎉🎉 Let's celebrate 2026 with a bugfix release that implements some fixes, brings some documentation improvements and adds a new dataset fetcher: github.com/skrub-data/s...
github.com
Release Skrub release 0.7.1 · skrub-data/skrub
Release 0.7.1 New features A new dataset, fetch_california_housing(), has been added to the skrub.datasets module. It allows to get a redundancy copy of the scikit-learn fetch_california_housing()...
010
Skrub @skrub-data.bsky.social · 19/12/2025
The course covers: - How to explore and sanitize data with skrub - How to use the skrub transformers for powerful and reliable feature engineering - How to put everything together in a machine learning pipeline Skrub Data Ops are not included (yet).
000
Skrub @skrub-data.bsky.social · 19/12/2025
Do you want to learn how to use skrub like a pro? Then you're in luck! Inria Academy is providing an introductory course on skrub aimed at IT personnel, engineers, data scientists, and data analysts. www.inria-academy.fr/formation/sk...
inria-academy.fr
skrub like a pro: clean, prepare, and transform your data faster - Inria Academy
130
Skrub @skrub-data.bsky.social · 16/12/2025
The recording of the talk we did at @pydataparis.bsky.social 2025 is now available on the PyData Youtube channel! 🚀 You can find it here, if you want to check it out 👀 www.youtube.com/watch?v=k9MN...
youtube.com
Skrub: machine learning for dataframes
YouTube video by PyData
030
Skrub @skrub-data.bsky.social · 12/12/2025
Skrub 0.7.0 is here! 🎉 ✨ Main highlights: - Tune hyperparameter choices with Optuna - Added support for Pandas 3.0 - Estimators in data ops can now take additional kwargs 16 new contributors helped with this release 👥 Check out the full changelog: github.com/skrub-data/s...
github.com
Release Skrub release 0.7.0 · skrub-data/skrub
Release 0.7.0 ✨ Highlights Data Ops can now be tuned with Optuna. It is now possible to pass extra named arguments to an estimator through DataOps.skb.apply. The TableReport now supports numpy arr...
030
Reposted by Skrub
Gaël Varoquaux @gaelvaroquaux.bsky.social · 17/11/2025
@skrub-data.bsky.social: better data-science primitives for clean code on dataframes Watch my dotAI talk, it's fun (live coding)! www.youtube.com/watch?v=bQS4... skrub really makes it easy to do machine learning with dataframes
youtube.com
Clean code in Data Science - Gael Varoquaux - Skrub DataOps, Probabl:
YouTube video by dotconferences
0278
Skrub @skrub-data.bsky.social · 08/10/2025
skrub-data.org/stable/refer...
skrub-data.org
ApplyToFrame
Gallery examples: Hands-On with Column Selection and Transformers
000