Aécio Santos @aeciosan.bsky.social · 07/08/2025Our new paper "Magneto: Combining Small and Large Language Models for Schema Matching" has just been published in the new issue of #PVLDB! The paper introduces a new framework that combines both small and large language models for effective schema matching. www.vldb.org/pvldb/vol18/...vldb.org 110
Aécio Santos @aeciosan.bsky.social · 21/06/2025📢 Tomorrow, I'll be presenting our new paper on LLM-based agents for interactive data integration at the #SIGMOD2025 NOVAS workshop. I'll also be in Berlin for the whole week, so please reach out if you'd like to chat or hang out! Paper: arxiv.org/abs/2502.07132arxiv.orgInteractive Data Harmonization with LLM AgentsData harmonization is an essential task that entails integrating datasets from diverse sources. Despite years of research in this area, it remains a time-consuming and challenging task due to schema m... 160
Reposted by Aécio SantosRaphaël Millière @raphaelmilliere.com · 03/06/2025Transformer-based neural networks achieve impressive performance on coding, math & reasoning tasks that require keeping track of variables and their values. But how can they do that without explicit memory? 📄 Our new ICML paper investigates this in a synthetic setting! 🎥 youtu.be/Ux8iNcXNEhw 🧵 1/13youtu.beHow Do Transformers Learn Variable Binding in Symbolic Programs?YouTube video by Raphaël Millière 1527
Reposted by Aécio SantosDEEM Workshop @ SIGMOD @deem-workshop.bsky.social · 07/02/2025The Data Management for End-to-End Machine Learning workshop (@deem-workshop.bsky.social) will be back at #SIGMOD2025! ✨ 🔗 Check out the CfP: deem-workshop.github.io 📝 Submission deadline: March 21 📢 Notifications: April 25 Join us for the 9th edition in Berlin! #DEEM2025 174
Reposted by Aécio SantosGaël Varoquaux @gaelvaroquaux.bsky.social · 15/12/2024Slides for "Table Foundation Models" I explain why these models can strongly outperform tree-based models, what are the intuitions, hopefully pointing to ways forward for more improvement speakerdeck.com/gaelvaroquau...speakerdeck.comTable foundation models for analyticsDeep-learning typically does not outperform tree-based models on tabular data. Often this may be explained by the small size of such datasets. For image… 38113
Aécio Santos @aeciosan.bsky.social · 14/12/2024@madelonhulsebos.bsky.social kicking off the 3rd Table Representation Learning workshop (@trl-research.bsky.social) at NeurIPS 2024. First keynote by @gaelvaroquaux.bsky.social. 0103