Paolo Papotti @papotti.bsky.social · 06/10/2026Francesco is presenting this work at #COLM2026 "Constraint decay: The Fragility of LLM Agents in Backend Code Generation" In Room: Imperial Ballroom Poster Location: #44 at 11:00 AM PDT in Poster Session 5 on Thu Oct 8th 000
Paolo Papotti @papotti.bsky.social · 15/09/2026I just read Sciascia’s "The Mystery of Majorana". He portrays Majorana as seeing very early the risks tied to his work, and choosing not to engage with them at all. To disappear rather than bear responsibility. I wonder how people working in frontier AI labs read his story today. 000
Paolo Papotti @papotti.bsky.social · 15/07/2026The position is based at EURECOM in Sophia Antipolis, France, within the 3IA Côte d'Azur program. Position starting in early 2027 (flexible). Description in the link above for details and application instructions. Ping me for any question! 020
Paolo Papotti @papotti.bsky.social · 15/07/2026I’m recruiting a postdoc researcher to work on efficient LLM-augmented systems for data and software tasks. We will explore task-specific LLM workflows that improve quality while reducing cost and latency. Candidates with expertise in DBMSs, LLMs, or agents encouraged to apply before Sep 13 26 👇nature.comEfficient LLM-Augmented Systems for Data and Software Tasks - Valbonne, Le Bar-sur-Loup (FR) job with 3IA Côte d'Azur | 12861764Supervisor: Prof. Paolo Papotti Affiliation: EURECOM Email: paolo.papotti@eurecom.fr Location: EURECOM, Sophia Antipolis, France Context and motiva... 210
Paolo Papotti @papotti.bsky.social · 02/07/2026This is not yet the “BLIP for databases”, but it is a step toward that goal: encoders that let LLMs reason over tables without forcing them to rediscover DB structure from serialized text. 1st paper from Simone Varriale in his PhD, with our amazing collaborators Tamara Cucumides and Floris Geerts 000
Paolo Papotti @papotti.bsky.social · 02/07/2026GRAB builds a graph over rows, column classes, and value groups, uses message passing to expose table structure, and passes a small number of query-conditioned latent tokens to the LLM. The LLM is frozen, preserving its NL skills and world knowledge, while the encoder supplies structural guidance. 100
Paolo Papotti @papotti.bsky.social · 02/07/2026Tables are not just text. Their meaning comes from rows, columns, schemas, and joins. Yet LLM-systems still flatten tables and reconstruct the structure implicitly. In a new preprint, we introduce our Graph-Relational Attention Bridge (GRAB): an encoder that adds tables as a modality to LLMs.arxiv.orgLatent Bridges for Multi-Table Question AnsweringWe introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. Our method lifts relational data into an heterogeneous graph, encodes it via message passing, and transfers the s... 150
Paolo Papotti @papotti.bsky.social · 24/06/2026Looking forward to the research ahead! I will share more information about the project and the open positions in my group in a follow-up post soon. 000
Paolo Papotti @papotti.bsky.social · 24/06/2026I have been awarded an ERC Advanced Grant! 🎉 We will explore how to build logic-grounded language models to query relational databases in natural language with certified answers. This result stands on the advice of many mentors and on the energy of the students I have worked with - thank you! 3100
Paolo Papotti @papotti.bsky.social · 06/06/2026AgenticDev aims to bring together researchers and practitioners working on AI agents, agentic workflows and soft eng practices powered by autonomous and collaborative AI systems Accepted papers will be included in the ASE'26 proceedings + selected papers will be invited to journal special issue 110
Paolo Papotti @papotti.bsky.social · 06/06/2026CFP! ⏰ Workshop on Agentic AI for Next-Generation Software Development (AgenticDev) @ ASE 2026 - AI agents for software development - Multi-agent collaboration for development tasks - Human-agent interaction - and more Deadline: July 15 '26 Workshop: Oct 12 '26 conf.researchr.org/home/ase-202...conf.researchr.orgAgenticDev 2026 - ASE 2026The International Workshop on Agentic AI for Next-Generation Software Development (AgenticDev 2026) brings together researchers, practitioners, and industry innovators to explore the emerging paradigm... 130
Paolo Papotti @papotti.bsky.social · 12/05/2026arxiv.org/abs/2605.06445 This paper highlights that jointly satisfying functional and structural requirements remains an open challenge for coding agents. Work led by Francesco Dente and Dario Satriani in the context of the EU-funded project AI4SWengarxiv.orgConstraint Decay: The Fragility of LLM Agents in Backend Code GenerationLarge Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural co... 141
Paolo Papotti @papotti.bsky.social · 12/05/2026Can LLM coding agents follow strict architectural rules? 🤖 We study how agents handle backend code generation when forced to use specific architectural patterns We name our core finding Constraint Decay: agents excel at unconstrained generation, but performance drop with structural requirements 🧵👇 250
Paolo Papotti @papotti.bsky.social · 17/02/2026We are reopening the interviews for this PhD position. Please help me spread the word to find the right potential candidates! 030
Paolo Papotti @papotti.bsky.social · 11/02/2026We introduce - Query planning as constrained optimization over quality constraints and cost objective - Gradient-based optimization to jointly choose operators and allocate error budgets across pipelines - KV-cache–based operators to turn discrete physical choices into a runtime-quality continuum 010
Paolo Papotti @papotti.bsky.social · 11/02/2026Co-authors: Gabriele Sanmartino, Matthias Urban, Paolo Papotti, Carsten Binnig This is the first outcome of our collaboration with Technische Universität Darmstadt within the @agencerecherche.bsky.social / @dfg.de ANR/DFG #Magiq project - more to come! 110
Paolo Papotti @papotti.bsky.social · 11/02/2026Empirically, Stretto delivers 2x-10x faster execution 🔥 across various datasets and queries compared to prior systems that meet quality guarantees. 100
Paolo Papotti @papotti.bsky.social · 11/02/2026🚀 New: The Stretto Execution Engine for LLM-Augmented Data Systems. LLM operators create a runtime ↔ accuracy trade-off in query execution. We address it with a novel optimizer, for end-to-end quality guarantees, and new KV-cache–based operators, for efficiency. arxiv.org/abs/2602.04430 Details👇 140
Paolo Papotti @papotti.bsky.social · 01/02/2026Happy Fontaines D.C.'s fan from the last album (2024). But the real treat was discovering the previous ones! 010
Paolo Papotti @papotti.bsky.social · 19/01/2026I agree. Here is another trick for input context we recently published bsky.app/profile/papo... 010
Paolo Papotti @papotti.bsky.social · 15/01/2026These results point toward models that decide which retrieved document to trust, turning “context engineering” from a static prompt recipe into a dynamic decoding policy. Amazing work from Giulio Corallo in his industrial PhD at SAP! 010
Paolo Papotti @papotti.bsky.social · 15/01/2026Key insight: 𝐄𝐯𝐢𝐝𝐞𝐧𝐜𝐞 𝐚𝐠𝐠𝐫𝐞𝐠𝐚𝐭𝐢𝐨𝐧 𝐡𝐚𝐩𝐩𝐞𝐧𝐬 𝐚𝐭 𝐝𝐞𝐜𝐨𝐝𝐢𝐧𝐠 𝐭𝐢𝐦𝐞, the model can effectively “switch” which document drives each token - without cross-document attention! 110
Paolo Papotti @papotti.bsky.social · 15/01/2026📈 Results: PCED often matches (and sometimes beats) long-context concatenation, while dramatically outperforming KV merge baseline on multi-doc QA/ICL. 🚀 Systems win: ~180× faster time-to-first-token vs long-context prefill using continuous batching and Paged Attention. 100
Paolo Papotti @papotti.bsky.social · 15/01/2026Instead of concatenating docs into one context (slow, noisy attention), training-free PCED: ● Keeps each document as its own 𝐞𝐱𝐩𝐞𝐫𝐭 with independent KV cache ● Runs experts in 𝐩𝐚𝐫𝐚𝐥𝐥𝐞𝐥 to get logits ● Selects next token with a 𝐫𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥-𝐚𝐰𝐚𝐫𝐞 𝐜𝐨𝐧𝐭𝐫𝐚𝐬𝐭𝐢𝐯𝐞 𝐝𝐞𝐜𝐨𝐝𝐢𝐧𝐠 rule integrating scores as a prior 100
Paolo Papotti @papotti.bsky.social · 15/01/2026🛑 𝐒𝐭𝐨𝐩 𝐭𝐡𝐫𝐨𝐰𝐢𝐧𝐠 𝐚𝐰𝐚𝐲 𝐲𝐨𝐮𝐫 𝐫𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥 𝐬𝐜𝐨𝐫𝐞𝐬. RAG uses embedding scores to pick Top-K, then treat all retrieved chunks as equal. Parallel Context-of-Experts Decoding (PCED) uses retrieval scores to move evidence aggregation from attention to decoding. 🚀 180× faster time-to-first-token!arxiv.orgParallel Context-of-Experts Decoding for Retrieval Augmented GenerationRetrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document KV caches separatel... 151
Paolo Papotti @papotti.bsky.social · 21/11/2025New PhD position on Tool-Augmented LLMs for Enterprise Data AI 🚨 Starting in early 2026 under my academic supervision and hosted by the fantastic team at AILY LABS in Madrid or Barcelona Details reported in the link - please ping me for any question! www.linkedin.com/jobs/view/43...linkedin.comAILY LABS hiring PhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AI in Barcelona, Catalonia, Spain | LinkedInPosted 11:18:12 AM. MissionPhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AIIndustry hire at…See this and similar jobs on LinkedIn. 021
Reposted by Paolo PapottiPVLDB @pvldb.bsky.social · 04/09/2025Vol:18 No:12 → Accelerating Tabular Inference: Training Data Generation with TENET 👥 Authors: Enzo Veltri, Donatello Santoro, Jean-Flavien Bussotti, Paolo Papotti 📄 PDF: www.vldb.org/pvldb/vol18/p5303-velt… 032
Reposted by Paolo PapottiRaphaël Troncy @rtroncy.bsky.social · 30/06/2025Can We Trust the Judges? This is the question we asked in validating factuality evaluation methods via answer perturbation. Check out the results at the #EvalLLM2025 workshop at #TALN2025 Blog: giovannigatti.github.io/trutheval/ Watch: www.youtube.com/watch?v=f0XJ... Play: github.com/GiovanniGatt... 031
Paolo Papotti @papotti.bsky.social · 02/06/2025Kudos to my amazing co-authors Dario Satriani, Enzo Veltri, Donatello Santoro! Another great collaboration between Università degli Studi della Basilicata and EURECOM 🙌 #LLM #Factuality #Benchmark #RelationalFactQA #NLP #AI 020
Paolo Papotti @papotti.bsky.social · 02/06/2025Structured outputs power analytics, reporting, and tool-augmented agents. This work exposes where current LLMs fall short and offers a clear tool for measuring progress on factuality beyond single-value QA. 📊 110
Paolo Papotti @papotti.bsky.social · 02/06/2025We release a new factuality benchmark with 696 annotated natural-language questions paired with gold factual answers expressed as tables (avg. 27 rows × 5 attributes), spanning 9 knowledge domains, with controlled question complexity and rich metadata. 100
Paolo Papotti @papotti.bsky.social · 02/06/2025Our new paper, "RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models", measures exactly this gap. Wider or longer output tables = tougher for all LLMs! 🧨 From Llama 3 and Qwen to GPT-4, no LLM goes above 25% accuracy on our stricter measure. 100
Paolo Papotti @papotti.bsky.social · 02/06/2025Ask any LLM for a single fact and it’s usually fine. Ask it for a rich list and the same fact is suddenly missing or hallucinated because the output context got longer 😳 LLMs exceed 80% accuracy on single-value questions but accuracy drops linearly with the # of output facts New paper, details 👇arxiv.orgRelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language ModelsFactuality in Large Language Models (LLMs) is a persistent challenge. Current benchmarks often assess short factual answers, overlooking the critical ability to generate structured, multi-record tabul... 180
Paolo Papotti @papotti.bsky.social · 01/06/2025and a special thanks to @tanmoy-chak.bsky.social for leading this effort! 051
Paolo Papotti @papotti.bsky.social · 01/06/2025More co-authors here on bsky @iaugenstein.bsky.social @preslavnakov.bsky.social @igurevych.bsky.social @emilioferrara.bsky.social @fil.bsky.social @giovannizagni.bsky.social @dcorney.com @mbakker.bsky.social @computermacgyver.bsky.social @irenelarraz.bsky.social @gretawarren.bsky.social 141
Paolo Papotti @papotti.bsky.social · 01/06/2025It’s time we rethink how "facts" are negotiated in the age of platforms. Excited to hear your thoughts! #Misinformation #FactChecking #SocialMedia #Epistemology #HCI #DigitalTruth #CommunityNotes arxiv.org/pdf/2505.20067arxiv.org 160
Paolo Papotti @papotti.bsky.social · 01/06/2025Community-based moderation offers speed & scale, but also raises tough questions: – Can crowds overcome bias? – What counts as evidence? – Who holds epistemic authority? Our interdisciplinary analysis combines perspectives from HCI, media studies, & digital governance. 121
Paolo Papotti @papotti.bsky.social · 01/06/2025Platforms like X are outsourcing fact-checking to users via tools like Community Notes. But what does this mean for truth online? We argue this isn’t just a technical shift — it’s an epistemological transformation. Who gets to define what's true when everyone is the fact-checker? 194
Paolo Papotti @papotti.bsky.social · 01/06/2025🚨 𝐖𝐡𝐚𝐭 𝐡𝐚𝐩𝐩𝐞𝐧𝐬 𝐰𝐡𝐞𝐧 𝐭𝐡𝐞 𝐜𝐫𝐨𝐰𝐝 𝐛𝐞𝐜𝐨𝐦𝐞𝐬 𝐭𝐡𝐞 𝐟𝐚𝐜𝐭-𝐜𝐡𝐞𝐜𝐤𝐞𝐫? new "Community Moderation and the New Epistemology of Fact Checking on Social Media" with I Augenstein, M Bakker, T. Chakraborty, D. Corney, E Ferrara, I Gurevych, S Hale, E Hovy, H Ji, I Larraz, F Menczer, P Nakov, D Sahnan, G Warren, G Zagniarxiv.org 1158
Reposted by Paolo PapottiRiccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025🌟 New paper alert! 🌟 Our paper, "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes", has been published in TMLR! In this work, we created YADL (a semi-synthetic data lake), and we benchmarked methods for augmenting user-provided tables given information found in data lakes. 1/ 263
Paolo Papotti @papotti.bsky.social · 05/05/2025Thanks for the amazing work to the whole team! Joint work between Università degli Studi della Basilicata (Enzo Veltri, Donatello Santoro, Dario Satriani) and EURECOM (Sara Rosato, Simone Varriale). #SQL #DataManagement #QueryOptimization #AI #LLM #Databases #SIGMOD2025 010
Paolo Papotti @papotti.bsky.social · 05/05/2025The principles in Galois – optimizing for quality alongside cost & dynamically acquiring optimization metadata – are a promising starting point for building robust and effective declarative data systems over LLMs. 💡 Paper and code: github.com/dbunibas/gal...github.comGitHub - dbunibas/galois: GaloisGalois. Contribute to dbunibas/galois development by creating an account on GitHub. 110
Paolo Papotti @papotti.bsky.social · 05/05/2025This cost/quality trade-off is guided by dynamically estimated metadata instead of relying on traditional stats. Result: Significant quality gains (+29%) without prohibitive costs. Works across LLMs & for internal knowledge + in-context data (RAG-like setup, reported results in the figure). ✅ 100
Paolo Papotti @papotti.bsky.social · 05/05/2025With our Galois system, we show one path to adapt database optimization for LLMs: 🔹 Designing physical operators tailored to LLM interaction nuances (e.g., Table-Scan vs Key-Scan in the figure). 🔹 Rethinking logical optimization (like pushdowns) for a cost/quality trade-off. 100
Paolo Papotti @papotti.bsky.social · 05/05/2025Why do traditional methods fail? They prioritize execution cost & ignore crucial LLM response quality (factuality, completeness). Our results show standard techniques like predicate pushdown can even reduce result quality by making LLM prompts more complex to process accurately. 🤔 100
Paolo Papotti @papotti.bsky.social · 05/05/2025Our new @sigmod2025.bsky.social paper tackles a fundamental challenge for the next gen of data systems: "Logical and Physical Optimizations for SQL Query Execution over Large Language Models" 📄 As systems increasingly use declarative interfaces on LLMs, traditional optimization falls short Details 👇 150
Paolo Papotti @papotti.bsky.social · 30/04/2025Alberto Sánchez Pérez (AILY LABS) will explain how we generate high-level hypotheses, use an agent to query databases via SQL, and summarize the findings into concise, correct, and insightful text. Joint work with Alaa Boukhary, Luis Castejón Lozano, Adam Elwood 010
Paolo Papotti @papotti.bsky.social · 30/04/2025Presenting at #NAACL2025 today (April 30th) 🎤 ⏰ 11:00 Session B Our work, "An LLM-Based Approach for Insight Generation in Data Analysis," uses LLMs to automatically find insights in databases, outperforming baselines both in insightfulness and correctness Paper: arxiv.org/abs/2503.11664 Details 👇 162
Paolo Papotti @papotti.bsky.social · 29/04/2025Work led by @spapicchio.bsky.social , in collaboration with Simone Rossi (EURECOM) and Luca Cagliero (Politecnico Torino) #Text2SQL #LLM #AI #NLP #ReinforcementLearning 020