Sign in

Skrub

@skrub-data.bsky.social
649 followers 48 following 164 posts

skrub is a Python library to ease preprocessing and feature engineering for tabular machine learning. Our long-term goal is to directly connect database tables to machine learning estimators. skrub-data.org discord.gg/ABaPnm7fDC

PostsRepliesMedia
Skrub @skrub-data.bsky.social · 07/10/2026
Messy numeric strings → float32 in one line: skrub.ToFloat ✨ Comma decimals, thousands separators, (negative) parentheses -- all covered, in both pandas and polars.
A pandas series that contains float numbers with non-standard separators and parentheses as negatives is transformed into float32 by the skrub ToFloat transformer.
110
Skrub @skrub-data.bsky.social · 01/10/2026
✨ Skrub 0.11 has been released ✨ Full changelog: skrub-data.org/stable/CHANG... Heads-up: skrub 0.11 requires Python ≥ 3.11 and scikit-learn ≥ 1.5.2. Highlights in thread ⤵️
131
Skrub @skrub-data.bsky.social · 05/02/2026
Did you know that the skrub Data Ops support Optuna as backend to run hyperparameter search? It's as easy as writing "backend='optuna'": this will set up a default Optuna study (and the TPE sampler) to replace the standard random sampler.
Three snippets of python code showing how to use skrub Data Ops with the Optuna optimization library.The first snippet shows a standard randomized search with the Data Ops. The second snippet adds the parameter "backend", which is set to "optuna". The third snippet uses the Optuna visualization API to plot information from the study.
142
Skrub @skrub-data.bsky.social · 08/10/2025
ApplyToFrame selects columns in the same way, but then uses all of them at the same time as input to the transformer: this is useful for dimensionality reduction. SelectCols and DropCols can be used as "filtering blocks" in a pipeline.
100
Skrub @skrub-data.bsky.social · 08/10/2025
Skrub includes a powerful set of transformers and selectors that allow to transform columns based on various conditions. ApplyToCols lets you select a subset of columns in your dataframe, then applies a transformer to each selected column separately.
130
Skrub @skrub-data.bsky.social · 07/10/2025
@pydataparis.bsky.social 2025 is over, and it was a big success! Our talk was very well received, and we got a lot of great questions, especially about scalability and how to interface with other libraries in production environments.
The skrub sticker on the back of a laptop
150
Skrub @skrub-data.bsky.social · 12/09/2025
skrub DataOps help you construct complex and extensive hyperparameter search spaces. However, interpreting results from large grids can be challenging. To address this, skrub generates a parallel coordinate plot that visualizes all runs and the parameters used to achieve specific results.
160
Skrub @skrub-data.bsky.social · 05/09/2025
Do you have to deal with numerical features that involve large outliers, and need to train linear models or neural networks? Then you might want to try the skrub SquashingScaler. The SquashingScaler behaves like scikit-learn RobustScaler, but smoothly clips outliers to predefined boundaries.
111
Skrub @skrub-data.bsky.social · 24/07/2025
Form complex DataOps plans to train and tune machine learning models, then export the plans as learners, standalone objects that can be used on new data. Tune hyperparameters where they're defined, and explore the resulting space with a parallel coordinate plot
101
Skrub @skrub-data.bsky.social · 24/07/2025
🌟 Major feature! Skrub DataOps are a powerful new way of combining dataframe transformations over multiple tables with machine learning pipelines.
111
Skrub @skrub-data.bsky.social · 24/07/2025
⚡ Release 0.6.0 is now out! ⚡ 🚀 Major update! Skrub DataOps, various improvements for the TableReport, new tools for applying transformers to the columns, and a new robust transformer for numerical features are only some of the features included in this release.
163
Skrub @skrub-data.bsky.social · 19/06/2025
📅 The skrub API includes various functions and objects that help with dealing with datetime strings. 1/
131
Skrub @skrub-data.bsky.social · 04/06/2025
Finally, results can be shown with a parallel coordinate plot to find out the impact of different hyperparameters on the prediction task.
121
Skrub @skrub-data.bsky.social · 04/06/2025
👀 This week's post will be another sneak peek into skrub expressions, an upcoming feature that will ease the preparation and execution of machine learning pipelines on dataframes. This time we will focus on how expressions can simplify the construction of complex hyperparameter grids.
141
Skrub @skrub-data.bsky.social · 30/04/2025
👀 This week's post is a sneak peek into the next major Skrub feature, Skrub expressions 🚀 As this is a preview of an upcoming feature, we are looking for your thoughts and feedback before release.
153
Skrub @skrub-data.bsky.social · 23/04/2025
The Skrub TableReport is a lightweight tool that allows to get a rich overview of a table quickly and easily. ✅ Filter columns 🔎 Look at each column's distribution 📊 Get a high level view of the distributions through stats and plots, including correlated columns 🌐 Export the report as html
164
Skrub @skrub-data.bsky.social · 09/04/2025
And if you're not familiar with what Skrub is all about, you might want to check out our introductory slide deck here: skrub-data.org/skrub-materi...
0105
Skrub @skrub-data.bsky.social · 03/04/2025
🚀⚡ Release: 0.5.3 Check out the release notes: skrub-data.org/stable/CHANG... Highlights below ⤵️
1144
Skrub @skrub-data.bsky.social · 31/01/2025
🚀 The Skrub workshop at Campus Cyber in La Défense was a great success! Connecting with professionals from both startups and large companies has given us valuable insights for Skrub's next steps. Stay tuned for more!
121
Skrub @skrub-data.bsky.social · 28/01/2025
🎉⚡️Release 0.5.1: ◼ Encode strings faster and better with StringEncoder! StringEncoder applies a tf-idf vectorization followed by SVD to produce high quality and FAST embeddings of textual and categorical features.
0142
Skrub @skrub-data.bsky.social · 27/11/2024
There is much more: skrub.patch_display() adds the TableReport as a default representation for all dataframes skrub.column_association to check which columns are linked... Check out the changelog: skrub-data.org/stable/CHANG...
001
Skrub @skrub-data.bsky.social · 27/11/2024
Improved TableReport: ◼ tighter layout ◼ support any script (any alphabet حب माया) in the plots ◼ robust to outliers It works without dependencies, in any html-based environment (Jupyter notebooks, @vscode.dev, a simple web page...) Check it out on skrub-data.org 4/5
112
Skrub @skrub-data.bsky.social · 27/11/2024
Skrub can now easily drop columns with too many missing values. As always the TableVectorizer is very handy for preparation of data-frames, and it now comes with an option to drop those pesky columns skrub-data.org/stable/refer... 3/5
101
Skrub @skrub-data.bsky.social · 27/11/2024
Easily combine deep learning (language models on huggingface @hf.co ) for text entries with @scikit-learn.bsky.social gradient-boosted trees for pipelines that predict great on dataframes of mixed types. Skrub ensure the language model is downloaded, cached, picklable, everything for easy ops 2/5
101
Skrub @skrub-data.bsky.social · 27/11/2024
🎉⚡️Release 0.4: ◼ Easily use deep learning for text entries ◼ TableVectorizer can remove columns with too many missing values ◼ TableReport more robust and prettier ... 1/5
1114
Skrub @skrub-data.bsky.social · 19/11/2024
👀 Explore your dataframes interactively with TableReport. AKA: 📈 we heard you liked plots so we put plots in your tables 📈
1113
Skrub @skrub-data.bsky.social · 19/11/2024
📆 Encode text and high cardinality categorical data with the GapEncoder and MinHashEncoder, and extract features from dates with the DatetimeEncoder.
Encode text and high cardinality categorical data with the GapEncoder and MinHashEncoder, and extract features from dates with the DatetimeEncoder.
100
Skrub @skrub-data.bsky.social · 19/11/2024
🔬 Create strong scikit-learn pipeline baselines effortlessly with TableVectorizer and tabular_learner.
Create strong scikit-learn pipeline baselines effortlessly with TableVectorizer and tabular_learner.
110