Sign in

Riccardo Cappuzzo

@riccardocappuzzo.com
234 followers 592 following 107 posts

Research engineer at Inria Saclay, working on the Skrub library. PhD in computer science. Python, data preparation, ML, tabular learning. ORCID: 0000-0002-4448-2959 Hoshiyomi ☄️ www.riccardocappuzzo.com github.com/rcap107

PostsRepliesMedia
Riccardo Cappuzzo @riccardocappuzzo.com · 12/03/2026
What a world we are already living in www.adriankrebs.ch/blog/dead-in...
a screenshot of an email that reads "hey sorry - my agent got a mind of its own and started applying for jobs for me"
100
Riccardo Cappuzzo @riccardocappuzzo.com · 21/02/2026
And this is the result recorded with asciinema Selecting a file opens it in VS Code at the given line, very convenient
000
Riccardo Cappuzzo @riccardocappuzzo.com · 18/02/2026
Tags are built with this
100
Riccardo Cappuzzo @riccardocappuzzo.com · 18/02/2026
Rabbit hole of the day: writing a command that fuzzy searches in the repository for any substring, shows me a preview of the line with context and opens the file at the given line in VS Code. Requires fzf, universal-ctags and batcat
the screenshot of a shell script that uses fzf and ripgrep to find substrings, classes, and files in the skrub repository
100
Riccardo Cappuzzo @riccardocappuzzo.com · 09/02/2026
good pr
030
Riccardo Cappuzzo @riccardocappuzzo.com · 28/01/2026
I'm sorry, I couldn't resist
a green baseball cap with "man I love Fauna" written on it
020
Riccardo Cappuzzo @riccardocappuzzo.com · 12/01/2026
Funny bug of the day: if you try to use pandas' "guess_datetime_format" with datetimes where the hour and minute are the same as the year (like 1959 and 19:59), the parser will fail and return None. This bug is present in pandas 2.3.3, but has been fixed in the dev version.
A short script demonstrating how the `guess_datetime_format` function of pandas does not work as intended when trying to parse the datetime "1959-01-01 19:59:16": it returns none instead of returning the correct datetime format.
100
Riccardo Cappuzzo @riccardocappuzzo.com · 09/10/2025
"ok the test run is done, let's see" ... "this will be hard to debug"
000
Riccardo Cappuzzo @riccardocappuzzo.com · 04/09/2025
Working hard on the next @skrub-data.bsky.social slide deck...
130
Riccardo Cappuzzo @riccardocappuzzo.com · 13/06/2025
Really cool graffiti I spotted while walking around in the town where I live
010
Riccardo Cappuzzo @riccardocappuzzo.com · 20/05/2025
Now that the paper is out, I can finally share the totally-not-confusing script/plot/table map I made to track which scripts prepare which figures and tables and from what data. If it wasn't clear, don't do this. If you *really* have to, I used the @obsidian.md canvas for this.
050
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
A bit of a mess up with this figure! This is what it's supposed to look like 🙈
010
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
⏱️ Complex aggregation methods are slower and don't significantly boost prediction performance. 6/
101
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
⚖️ Beware of diminishing returns! Performance plateaus as more candidates are retrieved, while resource costs (time and RAM) keep rising. 5/
121
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
🎯 Simple metric-based retrieval and candidate selection methods often outperform complex methods and are more efficient. 4/
101
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
🔍 Good table retrieval is crucial helps finding candidates with useful features and fewer missing values. Jaccard containment is helpful but has its limits. 3/
101
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
🌳 Tree-based models offer better prediction and computational performance than deep learning-based methods in our setting, which involves training models over features that contain a large fraction of missing values. 2/
121
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
🌟 New paper alert! 🌟 Our paper, "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes", has been published in TMLR! In this work, we created YADL (a semi-synthetic data lake), and we benchmarked methods for augmenting user-provided tables given information found in data lakes. 1/
263
Riccardo Cappuzzo @riccardocappuzzo.com · 02/05/2025
More fun digging around my @last.fm scrobbles using with @matplotlib.org I had no idea how much of a difference changing fonts and background color could make
041
Riccardo Cappuzzo @riccardocappuzzo.com · 01/05/2025
I haven't been using a lot of Copilot until very recently, so I'm still learning what it can do. It just blew my mind by autocompleting the dictionary "release_dates" with the correct dates for Muse albums based on the fact I am looking at data about Muse in the script. wow
A dictionary that contains various Muse albums and their release dates, and a suggestion for the release date of an album that hasn't been added yet
000
Riccardo Cappuzzo @riccardocappuzzo.com · 30/04/2025
First experiment plotting my Last.fm scrobbles With 10 years worth of data, I'll be working on this for a while. Also first time working with @matplotlib.org stackplots, much finagling was involved
032
Riccardo Cappuzzo @riccardocappuzzo.com · 10/04/2025
The similarity is uncanny
100
Riccardo Cappuzzo @riccardocappuzzo.com · 14/03/2025
000
Riccardo Cappuzzo @riccardocappuzzo.com · 14/12/2024
One good thing about living in Paris is that, well, you're living in Paris.
040
Riccardo Cappuzzo @riccardocappuzzo.com · 12/12/2024
The good thing about running experiments in France is that I can feel slightly less guilty about my emissions, but still yikes
110
Riccardo Cappuzzo @riccardocappuzzo.com · 28/11/2024
Sister's Christmas cat
030
Riccardo Cappuzzo @riccardocappuzzo.com · 27/11/2024
Some of these samples are deeply, deeply unsettling From fugatto.github.io
110
Riccardo Cappuzzo @riccardocappuzzo.com · 19/11/2024
This is a very simple example of what I am working with, only I have potentially thousands of lines like this. Looking at the documentation, it does look like I wouldn't need a lot of the features of SSSOM (and it might just add overhead in my scenario). Still, thanks for the clarification 👍
110