Sign in

Paolo Papotti

@papotti.bsky.social
907 followers 381 following 93 posts

Associate Prof at EURECOM and 3IA Côte d'Azur Chair of Artificial Intelligence. ELLIS member. Data management and NLP/LLMs for information quality. www.eurecom.fr/~papotti

PostsRepliesMedia
Paolo Papotti @papotti.bsky.social · 06/10/2026
Francesco is presenting this work at #COLM2026 "Constraint decay: The Fragility of LLM Agents in Backend Code Generation" In Room: Imperial Ballroom Poster Location: #44 at 11:00 AM PDT in Poster Session 5 on Thu Oct 8th
000
Paolo Papotti @papotti.bsky.social · 15/09/2026
I just read Sciascia’s "The Mystery of Majorana". He portrays Majorana as seeing very early the risks tied to his work, and choosing not to engage with them at all. To disappear rather than bear responsibility. I wonder how people working in frontier AI labs read his story today.
Sciascia's book about Majorana's disappearance
000
Paolo Papotti @papotti.bsky.social · 15/07/2026
The position is based at EURECOM in Sophia Antipolis, France, within the 3IA Côte d'Azur program. Position starting in early 2027 (flexible). Description in the link above for details and application instructions. Ping me for any question!
view of EURECOM with Nice in the background
020
Paolo Papotti @papotti.bsky.social · 15/07/2026
I’m recruiting a postdoc researcher to work on efficient LLM-augmented systems for data and software tasks. We will explore task-specific LLM workflows that improve quality while reducing cost and latency. Candidates with expertise in DBMSs, LLMs, or agents encouraged to apply before Sep 13 26 👇
nature.com
Efficient LLM-Augmented Systems for Data and Software Tasks - Valbonne, Le Bar-sur-Loup (FR) job with 3IA Côte d'Azur | 12861764
Supervisor: Prof. Paolo Papotti Affiliation: EURECOM Email: paolo.papotti@eurecom.fr Location: EURECOM, Sophia Antipolis, France Context and motiva...
210
Paolo Papotti @papotti.bsky.social · 02/07/2026
This is not yet the “BLIP for databases”, but it is a step toward that goal: encoders that let LLMs reason over tables without forcing them to rediscover DB structure from serialized text. 1st paper from Simone Varriale in his PhD, with our amazing collaborators Tamara Cucumides and Floris Geerts
000
Paolo Papotti @papotti.bsky.social · 02/07/2026
GRAB builds a graph over rows, column classes, and value groups, uses message passing to expose table structure, and passes a small number of query-conditioned latent tokens to the LLM. The LLM is frozen, preserving its NL skills and world knowledge, while the encoder supplies structural guidance.
100
Paolo Papotti @papotti.bsky.social · 02/07/2026
Tables are not just text. Their meaning comes from rows, columns, schemas, and joins. Yet LLM-systems still flatten tables and reconstruct the structure implicitly. In a new preprint, we introduce our Graph-Relational Attention Bridge (GRAB): an encoder that adds tables as a modality to LLMs.
arxiv.org
Latent Bridges for Multi-Table Question Answering
We introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. Our method lifts relational data into an heterogeneous graph, encodes it via message passing, and transfers the s...
150
Paolo Papotti @papotti.bsky.social · 24/06/2026
Looking forward to the research ahead! I will share more information about the project and the open positions in my group in a follow-up post soon.
000
Paolo Papotti @papotti.bsky.social · 24/06/2026
I have been awarded an ERC Advanced Grant! 🎉 We will explore how to build logic-grounded language models to query relational databases in natural language with certified answers. This result stands on the advice of many mentors and on the energy of the students I have worked with - thank you!
3100
Paolo Papotti @papotti.bsky.social · 06/06/2026
AgenticDev aims to bring together researchers and practitioners working on AI agents, agentic workflows and soft eng practices powered by autonomous and collaborative AI systems Accepted papers will be included in the ASE'26 proceedings + selected papers will be invited to journal special issue
110
Paolo Papotti @papotti.bsky.social · 06/06/2026
CFP! ⏰ Workshop on Agentic AI for Next-Generation Software Development (AgenticDev) @ ASE 2026 - AI agents for software development - Multi-agent collaboration for development tasks - Human-agent interaction - and more Deadline: July 15 '26 Workshop: Oct 12 '26 conf.researchr.org/home/ase-202...
conf.researchr.org
AgenticDev 2026 - ASE 2026
The International Workshop on Agentic AI for Next-Generation Software Development (AgenticDev 2026) brings together researchers, practitioners, and industry innovators to explore the emerging paradigm...
130
Paolo Papotti @papotti.bsky.social · 12/05/2026
arxiv.org/abs/2605.06445 This paper highlights that jointly satisfying functional and structural requirements remains an open challenge for coding agents. Work led by Francesco Dente and Dario Satriani in the context of the EU-funded project AI4SWeng
arxiv.org
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural co...
141
Paolo Papotti @papotti.bsky.social · 12/05/2026
Can LLM coding agents follow strict architectural rules? 🤖 We study how agents handle backend code generation when forced to use specific architectural patterns We name our core finding Constraint Decay: agents excel at unconstrained generation, but performance drop with structural requirements 🧵👇
Figure showing that adding constraints lower the performance of the coding agent
250
Paolo Papotti @papotti.bsky.social · 17/02/2026
We are reopening the interviews for this PhD position. Please help me spread the word to find the right potential candidates!
030
Paolo Papotti @papotti.bsky.social · 11/02/2026
We introduce - Query planning as constrained optimization over quality constraints and cost objective - Gradient-based optimization to jointly choose operators and allocate error budgets across pipelines - KV-cache–based operators to turn discrete physical choices into a runtime-quality continuum
main architecture
010
Paolo Papotti @papotti.bsky.social · 11/02/2026
Co-authors: Gabriele Sanmartino, Matthias Urban, Paolo Papotti, Carsten Binnig This is the first outcome of our collaboration with Technische Universität Darmstadt within the @agencerecherche.bsky.social / @dfg.de ANR/DFG #Magiq project - more to come!
110
Paolo Papotti @papotti.bsky.social · 11/02/2026
Empirically, Stretto delivers 2x-10x faster execution 🔥 across various datasets and queries compared to prior systems that meet quality guarantees.
plots of results
100
Paolo Papotti @papotti.bsky.social · 11/02/2026
🚀 New: The Stretto Execution Engine for LLM-Augmented Data Systems. LLM operators create a runtime ↔ accuracy trade-off in query execution. We address it with a novel optimizer, for end-to-end quality guarantees, and new KV-cache–based operators, for efficiency. arxiv.org/abs/2602.04430 Details👇
Stretto paper on arxiv
140
Paolo Papotti @papotti.bsky.social · 01/02/2026
Happy Fontaines D.C.'s fan from the last album (2024). But the real treat was discovering the previous ones!
010
Paolo Papotti @papotti.bsky.social · 23/01/2026
I d also like to test it, thanks!
000
Paolo Papotti @papotti.bsky.social · 19/01/2026
I agree. Here is another trick for input context we recently published bsky.app/profile/papo...
010
Paolo Papotti @papotti.bsky.social · 15/01/2026
These results point toward models that decide which retrieved document to trust, turning “context engineering” from a static prompt recipe into a dynamic decoding policy. Amazing work from Giulio Corallo in his industrial PhD at SAP!
010
Paolo Papotti @papotti.bsky.social · 15/01/2026
Key insight: 𝐄𝐯𝐢𝐝𝐞𝐧𝐜𝐞 𝐚𝐠𝐠𝐫𝐞𝐠𝐚𝐭𝐢𝐨𝐧 𝐡𝐚𝐩𝐩𝐞𝐧𝐬 𝐚𝐭 𝐝𝐞𝐜𝐨𝐝𝐢𝐧𝐠 𝐭𝐢𝐦𝐞, the model can effectively “switch” which document drives each token - without cross-document attention!
110
Paolo Papotti @papotti.bsky.social · 15/01/2026
📈 Results: PCED often matches (and sometimes beats) long-context concatenation, while dramatically outperforming KV merge baseline on multi-doc QA/ICL. 🚀 Systems win: ~180× faster time-to-first-token vs long-context prefill using continuous batching and Paged Attention.
100
Paolo Papotti @papotti.bsky.social · 15/01/2026
Instead of concatenating docs into one context (slow, noisy attention), training-free PCED: ● Keeps each document as its own 𝐞𝐱𝐩𝐞𝐫𝐭 with independent KV cache ● Runs experts in 𝐩𝐚𝐫𝐚𝐥𝐥𝐞𝐥 to get logits ● Selects next token with a 𝐫𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥-𝐚𝐰𝐚𝐫𝐞 𝐜𝐨𝐧𝐭𝐫𝐚𝐬𝐭𝐢𝐯𝐞 𝐝𝐞𝐜𝐨𝐝𝐢𝐧𝐠 rule integrating scores as a prior
100
Paolo Papotti @papotti.bsky.social · 15/01/2026
🛑 𝐒𝐭𝐨𝐩 𝐭𝐡𝐫𝐨𝐰𝐢𝐧𝐠 𝐚𝐰𝐚𝐲 𝐲𝐨𝐮𝐫 𝐫𝐞𝐭𝐫𝐢𝐞𝐯𝐚𝐥 𝐬𝐜𝐨𝐫𝐞𝐬. RAG uses embedding scores to pick Top-K, then treat all retrieved chunks as equal. Parallel Context-of-Experts Decoding (PCED) uses retrieval scores to move evidence aggregation from attention to decoding. 🚀 180× faster time-to-first-token!
arxiv.org
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
Retrieval Augmented Generation faces a trade-off: concatenating documents in a long prompt enables multi-document reasoning but creates prefill bottlenecks, while encoding document KV caches separatel...
151
Paolo Papotti @papotti.bsky.social · 21/11/2025
New PhD position on Tool-Augmented LLMs for Enterprise Data AI 🚨 Starting in early 2026 under my academic supervision and hosted by the fantastic team at AILY LABS in Madrid or Barcelona Details reported in the link - please ping me for any question! www.linkedin.com/jobs/view/43...
linkedin.com
AILY LABS hiring PhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AI in Barcelona, Catalonia, Spain | LinkedIn
Posted 11:18:12 AM. MissionPhD position (start: early 2026): Tool-Augmented LLMs for Enterprise Data AIIndustry hire at…See this and similar jobs on LinkedIn.
021
Reposted by Paolo Papotti
PVLDB @pvldb.bsky.social · 04/09/2025
Vol:18 No:12 → Accelerating Tabular Inference: Training Data Generation with TENET 👥 Authors: Enzo Veltri, Donatello Santoro, Jean-Flavien Bussotti, Paolo Papotti 📄 PDF: www.vldb.org/pvldb/vol18/p5303-velt…
Thumbnail: Accelerating Tabular Inference: Training Data Generation with TENET
032
Reposted by Paolo Papotti
Raphaël Troncy @rtroncy.bsky.social · 30/06/2025
Can We Trust the Judges? This is the question we asked in validating factuality evaluation methods via answer perturbation. Check out the results at the #EvalLLM2025 workshop at #TALN2025 Blog: giovannigatti.github.io/trutheval/ Watch: www.youtube.com/watch?v=f0XJ... Play: github.com/GiovanniGatt...
031
Paolo Papotti @papotti.bsky.social · 02/06/2025
Kudos to my amazing co-authors Dario Satriani, Enzo Veltri, Donatello Santoro! Another great collaboration between Università degli Studi della Basilicata and EURECOM 🙌 #LLM #Factuality #Benchmark #RelationalFactQA #NLP #AI
020
Paolo Papotti @papotti.bsky.social · 02/06/2025
Structured outputs power analytics, reporting, and tool-augmented agents. This work exposes where current LLMs fall short and offers a clear tool for measuring progress on factuality beyond single-value QA. 📊
110
Paolo Papotti @papotti.bsky.social · 02/06/2025
We release a new factuality benchmark with 696 annotated natural-language questions paired with gold factual answers expressed as tables (avg. 27 rows × 5 attributes), spanning 9 knowledge domains, with controlled question complexity and rich metadata.
100
Paolo Papotti @papotti.bsky.social · 02/06/2025
Our new paper, "RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models", measures exactly this gap. Wider or longer output tables = tougher for all LLMs! 🧨 From Llama 3 and Qwen to GPT-4, no LLM goes above 25% accuracy on our stricter measure.
100
Paolo Papotti @papotti.bsky.social · 02/06/2025
Ask any LLM for a single fact and it’s usually fine. Ask it for a rich list and the same fact is suddenly missing or hallucinated because the output context got longer 😳 LLMs exceed 80% accuracy on single-value questions but accuracy drops linearly with the # of output facts New paper, details 👇
arxiv.org
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
Factuality in Large Language Models (LLMs) is a persistent challenge. Current benchmarks often assess short factual answers, overlooking the critical ability to generate structured, multi-record tabul...
180
Paolo Papotti @papotti.bsky.social · 01/06/2025
and a special thanks to @tanmoy-chak.bsky.social for leading this effort!
051
Paolo Papotti @papotti.bsky.social · 01/06/2025
More co-authors here on bsky @iaugenstein.bsky.social @preslavnakov.bsky.social @igurevych.bsky.social @emilioferrara.bsky.social @fil.bsky.social @giovannizagni.bsky.social @dcorney.com @mbakker.bsky.social @computermacgyver.bsky.social @irenelarraz.bsky.social @gretawarren.bsky.social
141
Paolo Papotti @papotti.bsky.social · 01/06/2025
It’s time we rethink how "facts" are negotiated in the age of platforms. Excited to hear your thoughts! #Misinformation #FactChecking #SocialMedia #Epistemology #HCI #DigitalTruth #CommunityNotes arxiv.org/pdf/2505.20067
arxiv.org
160
Paolo Papotti @papotti.bsky.social · 01/06/2025
Community-based moderation offers speed & scale, but also raises tough questions: – Can crowds overcome bias? – What counts as evidence? – Who holds epistemic authority? Our interdisciplinary analysis combines perspectives from HCI, media studies, & digital governance.
121
Paolo Papotti @papotti.bsky.social · 01/06/2025
Platforms like X are outsourcing fact-checking to users via tools like Community Notes. But what does this mean for truth online? We argue this isn’t just a technical shift — it’s an epistemological transformation. Who gets to define what's true when everyone is the fact-checker?
194
Paolo Papotti @papotti.bsky.social · 01/06/2025
🚨 𝐖𝐡𝐚𝐭 𝐡𝐚𝐩𝐩𝐞𝐧𝐬 𝐰𝐡𝐞𝐧 𝐭𝐡𝐞 𝐜𝐫𝐨𝐰𝐝 𝐛𝐞𝐜𝐨𝐦𝐞𝐬 𝐭𝐡𝐞 𝐟𝐚𝐜𝐭-𝐜𝐡𝐞𝐜𝐤𝐞𝐫? new "Community Moderation and the New Epistemology of Fact Checking on Social Media" with I Augenstein, M Bakker, T. Chakraborty, D. Corney, E Ferrara, I Gurevych, S Hale, E Hovy, H Ji, I Larraz, F Menczer, P Nakov, D Sahnan, G Warren, G Zagni
arxiv.org
1158
Reposted by Paolo Papotti
Riccardo Cappuzzo @riccardocappuzzo.com · 19/05/2025
🌟 New paper alert! 🌟 Our paper, "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes", has been published in TMLR! In this work, we created YADL (a semi-synthetic data lake), and we benchmarked methods for augmenting user-provided tables given information found in data lakes. 1/
263
Paolo Papotti @papotti.bsky.social · 05/05/2025
Thanks for the amazing work to the whole team! Joint work between Università degli Studi della Basilicata (Enzo Veltri, Donatello Santoro, Dario Satriani) and EURECOM (Sara Rosato, Simone Varriale). #SQL #DataManagement #QueryOptimization #AI #LLM #Databases #SIGMOD2025
010
Paolo Papotti @papotti.bsky.social · 05/05/2025
The principles in Galois – optimizing for quality alongside cost & dynamically acquiring optimization metadata – are a promising starting point for building robust and effective declarative data systems over LLMs. 💡 Paper and code: github.com/dbunibas/gal...
github.com
GitHub - dbunibas/galois: Galois
Galois. Contribute to dbunibas/galois development by creating an account on GitHub.
110
Paolo Papotti @papotti.bsky.social · 05/05/2025
This cost/quality trade-off is guided by dynamically estimated metadata instead of relying on traditional stats. Result: Significant quality gains (+29%) without prohibitive costs. Works across LLMs & for internal knowledge + in-context data (RAG-like setup, reported results in the figure). ✅
100
Paolo Papotti @papotti.bsky.social · 05/05/2025
With our Galois system, we show one path to adapt database optimization for LLMs: 🔹 Designing physical operators tailored to LLM interaction nuances (e.g., Table-Scan vs Key-Scan in the figure). 🔹 Rethinking logical optimization (like pushdowns) for a cost/quality trade-off.
100
Paolo Papotti @papotti.bsky.social · 05/05/2025
Why do traditional methods fail? They prioritize execution cost & ignore crucial LLM response quality (factuality, completeness). Our results show standard techniques like predicate pushdown can even reduce result quality by making LLM prompts more complex to process accurately. 🤔
100
Paolo Papotti @papotti.bsky.social · 05/05/2025
Our new @sigmod2025.bsky.social paper tackles a fundamental challenge for the next gen of data systems: "Logical and Physical Optimizations for SQL Query Execution over Large Language Models" 📄 As systems increasingly use declarative interfaces on LLMs, traditional optimization falls short Details 👇
150
Paolo Papotti @papotti.bsky.social · 30/04/2025
Alberto Sánchez Pérez (AILY LABS) will explain how we generate high-level hypotheses, use an agent to query databases via SQL, and summarize the findings into concise, correct, and insightful text. Joint work with Alaa Boukhary, Luis Castejón Lozano, Adam Elwood
010
Paolo Papotti @papotti.bsky.social · 30/04/2025
Presenting at #NAACL2025 today (April 30th) 🎤 ⏰ 11:00 Session B Our work, "An LLM-Based Approach for Insight Generation in Data Analysis," uses LLMs to automatically find insights in databases, outperforming baselines both in insightfulness and correctness Paper: arxiv.org/abs/2503.11664 Details 👇
162
Paolo Papotti @papotti.bsky.social · 29/04/2025
Work led by @spapicchio.bsky.social , in collaboration with Simone Rossi (EURECOM) and Luca Cagliero (Politecnico Torino) #Text2SQL #LLM #AI #NLP #ReinforcementLearning
020