Paolo Papotti @papotti.bsky.social · 15/09/2026I just read Sciascia’s "The Mystery of Majorana". He portrays Majorana as seeing very early the risks tied to his work, and choosing not to engage with them at all. To disappear rather than bear responsibility. I wonder how people working in frontier AI labs read his story today. 000
Paolo Papotti @papotti.bsky.social · 15/07/2026I’m recruiting a postdoc researcher to work on efficient LLM-augmented systems for data and software tasks. We will explore task-specific LLM workflows that improve quality while reducing cost and latency. Candidates with expertise in DBMSs, LLMs, or agents encouraged to apply before Sep 13 26 👇nature.comEfficient LLM-Augmented Systems for Data and Software Tasks - Valbonne, Le Bar-sur-Loup (FR) job with 3IA Côte d'Azur | 12861764Supervisor: Prof. Paolo Papotti Affiliation: EURECOM Email: paolo.papotti@eurecom.fr Location: EURECOM, Sophia Antipolis, France Context and motiva... 210
Paolo Papotti @papotti.bsky.social · 02/07/2026Tables are not just text. Their meaning comes from rows, columns, schemas, and joins. Yet LLM-systems still flatten tables and reconstruct the structure implicitly. In a new preprint, we introduce our Graph-Relational Attention Bridge (GRAB): an encoder that adds tables as a modality to LLMs.arxiv.orgLatent Bridges for Multi-Table Question AnsweringWe introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. Our method lifts relational data into an heterogeneous graph, encodes it via message passing, and transfers the s... 150
Paolo Papotti @papotti.bsky.social · 24/06/2026I have been awarded an ERC Advanced Grant! 🎉 We will explore how to build logic-grounded language models to query relational databases in natural language with certified answers. This result stands on the advice of many mentors and on the energy of the students I have worked with - thank you! 3100
Paolo Papotti @papotti.bsky.social · 06/06/2026CFP! ⏰ Workshop on Agentic AI for Next-Generation Software Development (AgenticDev) @ ASE 2026 - AI agents for software development - Multi-agent collaboration for development tasks - Human-agent interaction - and more Deadline: July 15 '26 Workshop: Oct 12 '26 conf.researchr.org/home/ase-202...conf.researchr.orgAgenticDev 2026 - ASE 2026The International Workshop on Agentic AI for Next-Generation Software Development (AgenticDev 2026) brings together researchers, practitioners, and industry innovators to explore the emerging paradigm... 130
Paolo Papotti @papotti.bsky.social · 12/05/2026Can LLM coding agents follow strict architectural rules? 🤖 We study how agents handle backend code generation when forced to use specific architectural patterns We name our core finding Constraint Decay: agents excel at unconstrained generation, but performance drop with structural requirements 🧵👇 250
Paolo Papotti @papotti.bsky.social · 17/02/2026We are reopening the interviews for this PhD position. Please help me spread the word to find the right potential candidates! 030
Paolo Papotti @papotti.bsky.social · 11/02/2026🚀 New: The Stretto Execution Engine for LLM-Augmented Data Systems. LLM operators create a runtime ↔ accuracy trade-off in query execution. We address it with a novel optimizer, for end-to-end quality guarantees, and new KV-cache–based operators, for efficiency. arxiv.org/abs/2602.04430 Details👇 140
Paolo Papotti @papotti.bsky.social · 15/01/2026🛑 𝐒𝐭𝐨𝐩 𝐭𝐡𝐫𝐨𝐰𝐢𝐧𝐠 𝐚𝐰𝐚𝐲 𝐲𝐨𝐮𝐫 𝐫𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥 𝐬𝐜𝐨𝐫𝐞𝐬. RAG uses embedding scores to pick Top-K, then treat all retrieved chunks as equal. Parallel Context-of-Experts Decoding (PCED) uses retrieval scores to move evidence aggregation from attention to decoding. 🚀 180× faster time-to-first-token!arxiv.orgParallel Context-of-Experts Decoding for Retrieval Augmented GenerationRetrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document KV caches separatel... 151
Paolo Papotti @papotti.bsky.social · 21/11/2025New PhD position on Tool-Augmented LLMs for Enterprise Data AI 🚨 Starting in early 2026 under my academic supervision and hosted by the fantastic team at AILY LABS in Madrid or Barcelona Details reported in the link - please ping me for any question! www.linkedin.com/jobs/view/43...linkedin.comAILY LABS hiring PhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AI in Barcelona, Catalonia, Spain | LinkedInPosted 11:18:12 AM. MissionPhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AIIndustry hire at…See this and similar jobs on LinkedIn. 021
Reposted by Paolo PapottiPVLDB @pvldb.bsky.social · 04/09/2025Vol:18 No:12 → Accelerating Tabular Inference: Training Data Generation with TENET 👥 Authors: Enzo Veltri, Donatello Santoro, Jean-Flavien Bussotti, Paolo Papotti 📄 PDF: www.vldb.org/pvldb/vol18/p5303-velt… 032
Reposted by Paolo PapottiRaphaël Troncy @rtroncy.bsky.social · 30/06/2025Can We Trust the Judges? This is the question we asked in validating factuality evaluation methods via answer perturbation. Check out the results at the #EvalLLM2025 workshop at #TALN2025 Blog: giovannigatti.github.io/trutheval/ Watch: www.youtube.com/watch?v=f0XJ... Play: github.com/GiovanniGatt... 031
Paolo Papotti @papotti.bsky.social · 02/06/2025Ask any LLM for a single fact and it’s usually fine. Ask it for a rich list and the same fact is suddenly missing or hallucinated because the output context got longer 😳 LLMs exceed 80% accuracy on single-value questions but accuracy drops linearly with the # of output facts New paper, details 👇arxiv.orgRelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language ModelsFactuality in Large Language Models (LLMs) is a persistent challenge. Current benchmarks often assess short factual answers, overlooking the critical ability to generate structured, multi-record tabul... 180
Paolo Papotti @papotti.bsky.social · 01/06/2025🚨 𝐖𝐡𝐚𝐭 𝐡𝐚𝐩𝐩𝐞𝐧𝐬 𝐰𝐡𝐞𝐧 𝐭𝐡𝐞 𝐜𝐫𝐨𝐰𝐝 𝐛𝐞𝐜𝐨𝐦𝐞𝐬 𝐭𝐡𝐞 𝐟𝐚𝐜𝐭-𝐜𝐡𝐞𝐜𝐤𝐞𝐫? new "Community Moderation and the New Epistemology of Fact Checking on Social Media" with I Augenstein, M Bakker, T. Chakraborty, D. Corney, E Ferrara, I Gurevych, S Hale, E Hovy, H Ji, I Larraz, F Menczer, P Nakov, D Sahnan, G Warren, G Zagniarxiv.org 1158
Reposted by Paolo PapottiRiccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025🌟 New paper alert! 🌟 Our paper, "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes", has been published in TMLR! In this work, we created YADL (a semi-synthetic data lake), and we benchmarked methods for augmenting user-provided tables given information found in data lakes. 1/ 263
Paolo Papotti @papotti.bsky.social · 05/05/2025Our new @sigmod2025.bsky.social paper tackles a fundamental challenge for the next gen of data systems: "Logical and Physical Optimizations for SQL Query Execution over Large Language Models" 📄 As systems increasingly use declarative interfaces on LLMs, traditional optimization falls short Details 👇 150
Paolo Papotti @papotti.bsky.social · 30/04/2025Presenting at #NAACL2025 today (April 30th) 🎤 ⏰ 11:00 Session B Our work, "An LLM-Based Approach for Insight Generation in Data Analysis," uses LLMs to automatically find insights in databases, outperforming baselines both in insightfulness and correctness Paper: arxiv.org/abs/2503.11664 Details 👇 162
Paolo Papotti @papotti.bsky.social · 29/04/2025Think2SQL: Bridging the Reasoning Gap in Text-to-SQL for Small LLMs Leveraging RL with our reward mechanism, we push Qwen-Coder-2.5 7B to performance on par with much larger LLMs (>400B) on the BIRD dataset! 🤯 Model: huggingface.co/simone-papic... Paper: huggingface.co/papers/2504.... Details 👇 141
Paolo Papotti @papotti.bsky.social · 11/03/2025🗜️New LLM compression paper "Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning" RAG struggles with broad, multi-hop questions. We surpass RAG by up to 20 absolute points in QA performance, even with extreme cache compression (64x smaller)! Details 👇 260
Paolo Papotti @papotti.bsky.social · 29/01/2025NOVAS is a new venue for your paper bridging the gap between data management and generative AI research! It will be in Berlin, June 22th, together with @sigmod2025.bsky.social Submission deadline: 28 March 2025 141
Paolo Papotti @papotti.bsky.social · 23/01/2025Tropes, such as "Hidden Motives", are recurring narrative elements used to evoke familiar patterns in communication Our #COLING paper uncovers that tropes are used in 37% of the social posts debating immigration and vaccination 📄 coling-2025-proceedings.s3.us-east-1.amazonaws.com/main/pdf/202... 👇 411432
Paolo Papotti @papotti.bsky.social · 07/01/2025Meta is also embracing Community Notes (as now branded on X), the crowdsourcing approach to fact-checking on social networks. We have audited the program when it was called Birdwatch and found both promising results and concerning manipulation risks. More details below.👇arxiv.orgCrowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?Fact-checking is one of the effective solutions in fighting online misinformation. However, traditional fact-checking is a process requiring scarce expert human resources, and thus does not scale well... 133
Paolo Papotti @papotti.bsky.social · 24/11/2024🚀 Up to 93x input compression for LLMs! By compressing the data in the KV cache, we squeeze more info in the context. Presented at @emnlpmeeting.bsky.social, now on MIT Press: FINCH: Prompt-guided Key-Value Cache Compression for LLMs (TACL 2024) direct.mit.edu/tacl/article... More details 👇direct.mit.eduFINCH: Prompt-guided Key-Value Cache Compression for Large Language ModelsAbstract. Recent large language model applications, such as Retrieval-Augmented Generation and chatbots, have led to an increased need to process longer input contexts. However, this requirement is ha... 181
Reposted by Paolo PapottiMadelon Hulsebos @madelonhulsebos.bsky.social · 18/11/2024WIP starterpack w researchers on Table Representation Learning (TRL): all things related to representation learning and generative models for e.g. tables, DBs, spreadsheets! I'll curate but DM/reply w handle+some info welcome! Also follow @trl-research.bsky.social for updates 🤗 go.bsky.app/4SNSMRjgo.bsky.appTable Representation Learning researchersJoin the conversation 8248
Paolo Papotti @papotti.bsky.social · 18/11/2024CimpleKG is a continuously updated resource for researchers developing AI solutions to fight misinformation. The graph links data from 77 fact-checking orgs across 36 countries. 🔗 SPARQL Endpoint: purl.org/net/cimplekg... 🔗 KG Explorer: purl.org/net/cimplekg... 🔗 Paper: hal.science/hal-04760374... 041
Paolo Papotti @papotti.bsky.social · 16/11/2024𝗘𝘃𝗲𝗿 𝗰𝗼𝗻𝘀𝗶𝗱𝗲𝗿𝗲𝗱 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗶𝗻 𝘁𝗵𝗲 𝗙𝗿𝗲𝗻𝗰𝗵 𝗥𝗶𝘃𝗶𝗲𝗿𝗮? ☀ I'm seeking PhD and Post-doc candidates to join my research group in 2025 at EURECOM in the south of France. - 3 new projects on LLMs - Full-time positions with competitive salaries and benefits - English-speaking environment Interested? Ping me! 040
Paolo Papotti @papotti.bsky.social · 16/11/2024Our paper, "Data Void Exploits: Tracking & Mitigation Strategies," has received the Best Paper Award at ACM #CIKM 2024! 🏆 Data voids are gaps in online information, which are often exploit to spread disinformation. More details 👇 #CIKM2024 #DataVoids #Disinformation #KGs 130
Paolo Papotti @papotti.bsky.social · 16/11/2024Hi everyone! I'm a professor in the Data Science department at EURECOM, France. 🎓 My research focuses on data management and LLMs to enhance information quality, including data cleaning and misinformation detection. I'm here mostly for the research, but I occasionally comment on sports and arts. 070