Sign in

Lukas Gienapp

@lgnp.bsky.social
432 followers 116 following 9 posts

AI engineer @ Seltz; ML/IR research @ hessianAI / ScaDS.AI.

PostsRepliesMedia
Lukas Gienapp @lgnp.bsky.social · 22/07/2026
Big thanks to @hscells.bsky.social for presenting this work in Melbourne, and of course also to my co-authors @martin-potthast.com , Andrew Yates, and Eugene Yang! doi.org/10.1145/3805...
doi.org
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs | Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval
041
Lukas Gienapp @lgnp.bsky.social · 22/07/2026
Given the many problems that LLM-as-a-judge brings, this opens up new possibilities for scalable, reliable evaluation; and, continuing the thread from SIGIR 2025, this again confirms: human judgment stays the gold standard for evaluating IR and RAG. It's just more efficient to evaluate with it now 😉
100
Lukas Gienapp @lgnp.bsky.social · 22/07/2026
→ A tiny per-topic LoRA adapter aligns with human judgments better than prompted LLMs (26M vs up to 229B params!). → Needs only 128 judgments per topic to learn from and extrapolate to deep pools. → Gives deterministic, reproducible labels, while the adapters are shareable, versionable artifacts.
120
Lukas Gienapp @lgnp.bsky.social · 22/07/2026
Super happy that our paper “Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs” just won Best Student Paper at #SIGIR2026🏆 It argues against the reflex to reach for an LLM whenever you need relevance judgments. What we found 🧵:
1114
Lukas Gienapp @lgnp.bsky.social · 28/04/2026
🧵5/5 Preprints out now! 📄 "Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs" → downloads.webis.de/publications... 📄 "Humans, LLMs, and Measures Do Not Align in Attributed Information Retrieval" → downloads.webis.de/publications...
020
Lukas Gienapp @lgnp.bsky.social · 28/04/2026
🧵4/5 Our second paper provides supporting evidence from the RAG side: when we replace LLM-generated ground truth answers with human-written ones, reference-based evaluation scores shift significantly. Humans and LLM judges show low agreement, highlighting the importance of human grounding.
110
Lukas Gienapp @lgnp.bsky.social · 28/04/2026
🧵3/5 System rankings from our adapters correlate >0.94 with ground truth. Notably, this fares much better than using reasoning-capable LLMs for the same task; human judgment remains irreplaceable, but we can encode and transfer it efficiently.
100
Lukas Gienapp @lgnp.bsky.social · 28/04/2026
🧵2/5 Our first paper proposes a new solution to the classic unjudged document problem: topic-specific judge adapters. We propose finetuning a small ranking model with LoRA on a single assessor's judgments for a single topic, deliberately overfitting to that assessor's notion of relevance.
100
Lukas Gienapp @lgnp.bsky.social · 28/04/2026
Two papers accepted at SIGIR'26 in Melbourne! Both tackle the same question: how do we scale IR evaluation reliably without compromising on human judgment as the gold standard? Spoiler: LLMs-as-a-judge does not work. 🧵1/5
181
Reposted by Lukas Gienapp
Webis Group @webis.de · 27/10/2025
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
huggingface.co
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1209
Reposted by Lukas Gienapp
Webis Group @webis.de · 16/07/2025
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation.  📄 webis.de/publications...
22710
Reposted by Lukas Gienapp
Ferdinand Schlatt @fschlatt.bsky.social · 16/07/2025
Want to know how to make bi-encoders more than 3x faster with a new backbone encoder model? Check out our talk on the Token-Independent Text Encoder (TITE) #SIGIR2025 in the efficiency track. It pools vectors within the model to improve efficiency dl.acm.org/doi/10.1145/...
0105
Reposted by Lukas Gienapp
Maik Fröbe @maik-froebe.bsky.social · 15/07/2025
Lukas Gienapp presents "The Viability of Crowdsourcing for RAG Evaluation" at #SIGIR2025 The paper is available at: webis.de/publications...
0106
Reposted by Lukas Gienapp
Webis Group @webis.de · 22/06/2025
Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness.
163
Reposted by Lukas Gienapp
Webis Group @webis.de · 07/04/2025
📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️
1127