Webis Group @webis.de · 27/10/2025The data spans 7 text domains: 🌐 Web: Wikipedia, GitHub, social media 💬 Political: Parliamentary proceedings, speeches ⚖️ Legal: Court decisions, federal & EU law 📰 News: Newspaper archives 🏦 Economics: public tenders 📚 Cultural: Digital heritage collections 🔬 Scientific: Papers, books, journals 120
Webis Group @webis.de · 18/07/2025Honored to win the ICTIR Best Paper Honorable Mention Award for "Axioms for Retrieval-Augmented Generation"! Our new axioms are integrated with ir_axioms: github.com/webis-de/ir_... Nice to see axiomatic IR gaining momentum. 1166
Webis Group @webis.de · 18/07/2025We presented two papers at ICTIR 2025 today: - Axioms for Retrieval-Augmented Generation webis.de/publications... - Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins webis.de/publications... 183
Webis Group @webis.de · 18/07/2025Thrilled to announce that Matti Wiegmann has successfully defended his PhD! 🎉🧑🎓 Huge congratulations on this incredible achievement! #PhDDefense #AcademicMilestone 2123
Webis Group @webis.de · 16/07/2025Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation. 📄 webis.de/publications... 22710
Webis Group @webis.de · 22/06/2025Results on BEIR demonstrate that our method matches teacher distillation effectiveness, while using only 13.5% of the data and achieving 3-15x training speedup. This makes effective bi-encoder training more accessible, especially for low-resource settings. 110
Webis Group @webis.de · 22/06/2025The key idea: we can use the similarity predicted by the encoder itself between positive and negative documents to scale a traditional margin loss. This performs implicit hard negative mining and is hyperparameter-free. 110
Webis Group @webis.de · 22/06/2025Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness. 163
Webis Group @webis.de · 02/06/2025Our paper titled “The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification” has been accepted to #ACL2025 (Findings). downloads.webis.de/publications... We discuss why LLM detection is a one-class problem and how that affects the prospective… 1/3 #ACL #NLP #ARR #LLM 191
Webis Group @webis.de · 30/04/2025🧵 3/4 In a lot of cases, survey participants did not notice brand or product placements in the responses. As a first step towards ad-blockers for LLMs, we created a dataset of responses with and without ads and trained classifiers on the task of identifying the ads. dl.acm.org/doi/10.1145/... 131
Webis Group @webis.de · 30/04/2025🧵 2/4 Given the high operating costs of LLMs, they require a business model to sustain them and advertising is a natural candidate. Hence, we have analyzed how well LLMs can blend product placements with "organic" responses and whether users are able to identify the ads. dl.acm.org/doi/10.1145/... 121
Webis Group @webis.de · 30/04/2025Can LLM-generated ads be blocked? With OpenAI adding shopping options to ChatGPT, this question gains further importance. If you are interested in contributing to the research on LLM-based advertising, please check out our shared task: touche.webis.de/clef25/touch... More details below. 185
Webis Group @webis.de · 07/04/2025📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️ 1127
Webis Group @webis.de · 08/11/2024Below you can see our past tweets, just imported from “the darkened X”. Above, we see nothing but Bluesky. 020
Webis Group @webis.de · 21/07/2024Goodbye Washington! We had a fantastic week with interesting talks, discussions, and new ideas at #SIGIR24 #SIGIR2024. We hope to see you all again next year in Italy :) x.com/webis_de/status/1815115279510… 031
Webis Group @webis.de · 14/05/2024In our experiments, LLMs struggle with the task in a zero-shot setting, especially due to low precision values. Sentence transformers, however, can be finetuned to successfully detect the inserted ads and achieve precision and recall values of above 0.9 for unseen meta topics. t.co/VuuaW... 110
Webis Group @webis.de · 14/05/2024The Webis Generated Native Ads 2024 is the first public dataset to evaluate models on the task of detecting ads in responses of conversational search engines. It was created by simulating an advertising service for queries from popular meta topics (product/service categories). t.co/pjHr... 000
Webis Group @webis.de · 14/05/2024What if conversational search will be financed by inserting ads directly into generated responses? We present our work on detecting these generated native ads at #TheWebConf24. Come visit us at the short paper poster session on Thursday in the Central Ballroom. t.co/NRKbal57WO 000
Webis Group @webis.de · 14/03/2024Right now, we will start the second half of the SCAI'24 workshop at #CHIIR2024 in hybrid mode. We will move from the big ideas and human-centered metrics to the challenges of human-in-the-loop evaluations. x.com/webis_de/status/1768277930768… 010
Webis Group @webis.de · 06/03/2024How will conversational search AI pay for itself? It may be native ads or product placement in generated answers. At #CHIIR2024 next week, we'll present a user study showing that many people don't recognize ads inserted by LLMs in generated search results: t.co/hrZE9moeKy t.co/qg... 000
Webis Group @webis.de · 19/12/2023Working in Argumentation? Time to participate in Touché 2024! Three shared tasks: - Human Value Detection - Ideology and Power Identification in Parliamentary Debates - Image Retrieval/Generation for Arguments Submission deadline is May 6th! More info: t.co/rtgSDxpDTx t.co/S6Kl... 000
Webis Group @webis.de · 26/10/2023Today, we were happy to welcome @anja_reu and @juliusgonsior to our seminar to learn about current challenges in math retrieval/active learning: "Transformer Encoders for Mathematical Answer Retrieval" and "The Missing Piece of Active Learning Research: a Reference Benchmark". t.co/eh4IH... 010
Webis Group @webis.de · 06/10/2023Today we had the pleasure of listening to a talk from @WojciechKusa about evaluating automated citation screening in systematic reviews. Very interesting to hear about the work he has done in new metrics and datasets for this domain! x.com/webis_de/status/1710256640941… 010
Webis Group @webis.de · 25/07/2023Our @H1iReimer and @maik_froebe are thrilled to present two new resources at the @SIGIRConf poster session: • The TIREx platform to run reproducible, blinded IR experiments & shared tasks 🧪 • The Archive Query Log, 350M queries crawled from the Internet Archive 🔍 #SIGIR2023 t.co/Np... 000
Webis Group @webis.de · 25/07/2023TIREx archives the Docker images for future replication and reproduction. Software submissions on TIREx can run on new additions to ir_datasets as retrieval approaches were implemented against the ir_datasets interface, promoting IR experiment #standardization. t.co/XRNGHmgBOa 000
Webis Group @webis.de · 25/07/2023TIREx covers shared tasks in IR. Organizers add their data to ir_datasets. Participants implement their approach against ir_datasets, making software submissions via Docker executed in a TIRA sandbox, enabling blinded experimentation and improving internal and external validity. t.co/GLg... 000
Webis Group @webis.de · 25/07/2023The Information Retrieval Experiment Platform (TIREx) integrates ir_datasets, ir_measures, PyTerrier, and TIRA for • standardized, • reproducible, • scalable, and ultimately • 𝗯𝗹𝗶𝗻𝗱𝗲𝗱 𝗲𝘅𝗽𝗲𝗿𝗶𝗺𝗲𝗻𝘁𝘀 in IR. Preprint: t.co/WWe26DCch2 #sigir2023 🧵 t.co/sPGvaNeYF9 000
Webis Group @webis.de · 25/07/2023TIREx archives the Docker images for future replication and reproduction. Software submissions on TIREx can run on new additions to ir_datasets as retrieval approaches were implemented against the ir_datasets interface, promoting IR experiment #standardization. t.co/LllDr0g61P 000
Webis Group @webis.de · 25/07/2023TIREx covers shared tasks in IR. Organizers add their data to ir_datasets. Participants implement their approach against ir_datasets, making software submissions via Docker executed in a TIRA sandbox, enabling blinded experimentation and improving internal and external validity. t.co/ME1... 000
Webis Group @webis.de · 25/07/2023The Information Retrieval Experiment Platform (TIREx) integrates ir_datasets, ir_measures, PyTerrier, and TIRA for • standardized, • reproducible, • scalable, and ultimately • 𝗯𝗹𝗶𝗻𝗱𝗲𝗱 𝗲𝘅𝗽𝗲𝗿𝗶𝗺𝗲𝗻𝘁𝘀 in IR. Preprint: t.co/WWe26DCch2 #SIGIR2023 🧵 t.co/xK9dT699Z6 000
Webis Group @webis.de · 23/07/2023Are you already keen on @SIGIRConf #sigir #sigir2023 in Taipei 🇹🇼? Here are the papers we look forward to presenting next week: x.com/webis_de/status/1682902948794… 030
Webis Group @webis.de · 14/07/2023Now the shared task on clickbait spoiling has come to an end. We had a great time at @SemEvalWorkshop #ACL2023NLP #ACL2023 and enjoyed the discussions. A big thank you to all participants and offline and online attendants! Data and submissions available at t.co/7jlJXiGiUd t.co/pt... 000
Webis Group @webis.de · 29/06/2023We are very happy to share that our @albondarenko2 successfully defended his Ph.D. thesis on "Understanding Comparative Questions and Retrieving Argumentative Answers". Well done, and we look forward to being part of your next adventures! x.com/webis_de/status/1674480051222… 000
Webis Group @webis.de · 26/05/2023Today we had the pleasure of listening to a virtual talk from @HarrieOos about counterfactual learning to rank for search and recommendation. It was great to hear about some of his upcoming work that will also be presented this year at #sigir2023 t.co/vnBOCJGX6D 000
Webis Group @webis.de · 06/05/20233/6 There are millions of author-assigned trigger warnings on AO3 as freeform tags. All are reviewed and organized by the amazing Ao3 Tag Wranglers (@ao3_wranglers), who identify many relations between warning tags. We link >80% of them at 0.95 F1 into our abstract taxonomy. t.co/58gD... 000
Webis Group @webis.de · 06/05/2023Trigger Warnings: Can computers help us to assign them to (online) content? At #ACL2023NLP we introduce “Trigger Warning Assignment” as a new multi-label classification task. As a foundation, we contribute a taxonomy of warnings, a large dataset, and first approaches. Thread... t.co/R3... 000
Webis Group @webis.de · 23/04/2023#EACL2023 is just around the corner, where we will be showcasing our system demonstration paper "Small-Text: Learning for Text Classification in Python". The corresponding poster presentation is scheduled for Session 6 on May 3rd, 9:00–10:30 AM. #NLProc #TextClassification 1/2 t.co/i6... 000
Webis Group @webis.de · 19/04/2023A small number of users does not mean that a niche search engine is irrelevant. Case in point: Netspeak found at netspeak.org x.com/webis_de/status/1648571639938… 000
Webis Group @webis.de · 04/04/2023More use cases: · benchmarks collections with real-world query variants (see TREC overlap in picture) · diverse training data for neural retrieval models · transparent insights into search industry at large x.com/webis_de/status/1643366712648… 000
Webis Group @webis.de · 04/04/2023With queries from 2 decades, the AQL can be used for all sorts of diachronic analyses and to visualize global trends. x.com/webis_de/status/1643366709149… 000
Webis Group @webis.de · 04/04/2023We have analyzed a 4% portion of the AQL available at the time of writing. For example, the table gives a detailed breakdown of the AQL-22. x.com/webis_de/status/1643366705836… 000
Webis Group @webis.de · 04/04/2023Step ③ We extract queries from archived URLs of the @waybackmachine using provider-specific URL patterns. SERP URLs often contain the query in standard components of the URL, e.g., · query parameters · path segment The AQL contains 356M queries, of which 64M are unique. t.co/21WxiONxC1 000
Webis Group @webis.de · 04/04/2023We implement a four-step process to mine the AQL from the @internetarchive's @waybackmachine: ① list popular search providers (search engines + anything else) ② collect their archived URLs from @internetarchive's CDX API ③ parse queries from URLs ④ parse SERP HTML t.co/jF20CJYitm 000
Webis Group @webis.de · 04/04/2023If you're attending @ecir2023, be sure to catch @maik_froebe to get his take on the AQL. #ECIR2023 Now for some details… x.com/webis_de/status/1643366688006… 000
Webis Group @webis.de · 04/04/2023The Archive Query Log (AQL) is the first large log of archived search result pages (SERPs) Mined from @internetarchive, it contains · 356 million queries · 166 million SERPs from · 550 search engines of · 25 years Preprint: t.co/bVUG2NhFvV #SIGIR2023 #internetarchive 🧵 t.co/cq... 000
Webis Group @webis.de · 30/03/2023In our reading group today, @albondarenko2 led us through "TruthfulQA: Measuring How Models Mimic Human Falsehoods" by Stephanie Lin, Jacob Hilton, and @OwainEvans_UK Paper: aclanthology.org/2022.acl-long.229 x.com/webis_de/status/1641423089098… 000
Webis Group @webis.de · 09/02/2023Registration is now open for our new shared tasks at PAN 2023: Cross-Discourse Type Authorship Verification, Profiling Cryptocurrency Influencers, Multi-Author Analysis, and Trigger Detection. (pan.webis.de) x.com/webis_de/status/1623670004922… 000
Webis Group @webis.de · 18/12/2022The Infinite Index allows IR experiments that were never possible on finite datasets, and makes #ActiveLearning scenarios possible. However, many challenges remain: E.g., using recall-oriented measures is difficult, and near-duplicate images will require deduplication. (7/8) t.co/OCkH74kBWK 000
Webis Group @webis.de · 18/12/2022Another benefit of the Infinite Index is that it allows small perturbations (think interpolation of prompts—it's what makes things like t.co/eKyPGpTGDl possible). This helps to precisely analyze how users perceive generated images, and how to improve models. (6/8) t.co/zhvMsRhNgg 000
Webis Group @webis.de · 18/12/2022Second, we present a case study on #Game #Artwork Search that highlights challenges of prompt engineering: This is a non-trivial task, requiring experience as well as trial and error. A major goal will be to support users in generating intended images faster. (4/8) t.co/8f8ehRrAQe 000