Sign in

Webis Group

@webis.de
653 followers 698 following 270 posts

Information is nothing without retrieval The Webis Group contributes to information retrieval, natural language processing, machine learning, and symbolic AI.

PostsRepliesMedia
Webis Group @webis.de · 30/09/2026
Call for Participation: Deadlines extended for UniAgent, the 1st Shared Task on Agentic AI in University Administration at the CIKM 2026 AnalytiCup in Rome. Registration until Oct 3, submission Oct 20. Real administrative cases; we host the LLMs. Details: uniagent.webis.de/cikm26/uniag...
uniagent.webis.de
UniAgent at CIKM 2026 - Agentic AI in University Administration
A shared task on agentic AI in university administration, part of the CIKM 2026 AnalytiCup: build agents that solve standardized administrative tasks under controlled tool access and expert-judged ref...
032
Webis Group @webis.de · 03/09/2026
Call for Participation: UniAgent, the 1st Shared Task on Agentic AI in University Administration, at the CIKM 2026 AnalytiCup in Rome. Registration closes Sep 30. Submission deadline Oct 23. Details: uniagent.webis.de/cikm26/uniag...
uniagent.webis.de
UniAgent at CIKM 2026 - Agentic AI in University Administration
A shared task on agentic AI in university administration, part of the CIKM 2026 AnalytiCup: build agents that solve standardized administrative tasks under controlled tool access and expert-judged ref...
033
Webis Group @webis.de · 27/10/2025
For full technical details + compliance Datasheet see our preprint @ arxiv.org/abs/2510.13996 As for German-specific models trained on this data... stay tuned 👀
arxiv.org
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This problem is exacerbated f...
010
Webis Group @webis.de · 27/10/2025
The data spans 7 text domains: 🌐 Web: Wikipedia, GitHub, social media 💬 Political: Parliamentary proceedings, speeches ⚖️ Legal: Court decisions, federal & EU law 📰 News: Newspaper archives 🏦 Economics: public tenders 📚 Cultural: Digital heritage collections 🔬 Scientific: Papers, books, journals
120
Webis Group @webis.de · 27/10/2025
This means: ✅ Every document has verifiable usage rights (min. CC-BY-SA 4.0 and allows commercial use) ✅ Full institutional provenance for reduced compliance risks ✅ Systematic PII removal + quality filtering, ready for training ✅ Rich metadata for downstream customization
100
Webis Group @webis.de · 27/10/2025
The current problem: training data is primarily sourced from Web crawls, which give you scale but unclear licensing. This blocks models from commercial deployment and research. We took a different path: systematically collecting German text from 41 institutional sources with explicit open licenses.
100
Webis Group @webis.de · 27/10/2025
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
huggingface.co
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1209
Webis Group @webis.de · 18/07/2025
Congratulations to the authors @heinrich.merker.id, @maik-froebe.bsky.social, @benno-stein.de, @martin-potthast.com, @matthias-hagen.bsky.social from @uni-jena.de, Uni Weimar, @unikassel.bsky.social, @hessianai.bsky.social, @scadsai.bsky.social!
060
Webis Group @webis.de · 18/07/2025
Honored to win the ICTIR Best Paper Honorable Mention Award for "Axioms for Retrieval-Augmented Generation"! Our new axioms are integrated with ir_axioms: github.com/webis-de/ir_... Nice to see axiomatic IR gaining momentum.
1166
Webis Group @webis.de · 18/07/2025
We presented two papers at ICTIR 2025 today: - Axioms for Retrieval-Augmented Generation webis.de/publications... - Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins webis.de/publications...
183
Webis Group @webis.de · 18/07/2025
Thrilled to announce that Matti Wiegmann has successfully defended his PhD! 🎉🧑‍🎓 Huge congratulations on this incredible achievement! #PhDDefense #AcademicMilestone
2123
Webis Group @webis.de · 16/07/2025
Congrats to the authors @lgnp.bsky.social @timhagen.bsky.social @maik-froebe.bsky.social @matthias-hagen.bsky.social @benno-stein.de @martin-potthast.com @hscells.bsky.social from @unikassel.bsky.social @hessianai.bsky.social @scadsai.bsky.social @unituebingen.bsky.social @uni-jena.de & Uni Weimar
071
Webis Group @webis.de · 16/07/2025
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation.  📄 webis.de/publications...
22710
Reposted by Webis Group
Maik Fröbe @maik-froebe.bsky.social · 27/06/2025
Do not forget to participate in the #TREC2025 Tip-of-the-Tongue (ToT) Track :) The corpus and baselines (with run files) are now available and easily accessible via the ir_datasets API and the HuggingFace Datasets API. More details are available at: trec-tot.github.io/guidelines
Dory from finding nemo with the quote: "I remember it like it was yesterday. Of course, I dont remember yesterday."
0117
Webis Group @webis.de · 22/06/2025
Congratulations to the authors @lgnp.bsky.social @deckersniklas.bsky.social @martin-potthast.com @hscells.bsky.social ! 📄 Preprint: arxiv.org/abs/2407.21515 💻 Code: github.com/webis-de/ada...
arxiv.org
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
Representation-based retrieval models, so-called biencoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current state-of-the-art bien...
011
Webis Group @webis.de · 22/06/2025
Results on BEIR demonstrate that our method matches teacher distillation effectiveness, while using only 13.5% of the data and achieving 3-15x training speedup. This makes effective bi-encoder training more accessible, especially for low-resource settings.
110
Webis Group @webis.de · 22/06/2025
The key idea: we can use the similarity predicted by the encoder itself between positive and negative documents to scale a traditional margin loss. This performs implicit hard negative mining and is hyperparameter-free.
110
Webis Group @webis.de · 22/06/2025
Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness.
163
Webis Group @webis.de · 02/06/2025
…human texts today, contextualize the findings in terms of our theoretical contribution, and use them to make an assessment of the quality and adequacy of existing LLM detection benchmarks, which tend to be constructed with authorship attribution in mind, rather than authorship verification. 3/3
000
Webis Group @webis.de · 02/06/2025
…limits of the field. We argue that as LLMs improve, detection will not necessarily become impossible, but it will be limited by the capabilities and theoretical boundaries of the field of authorship verification. We conduct a series of exploratory analyses to show how LLM texts differ from… 2/3
110
Webis Group @webis.de · 02/06/2025
Our paper titled “The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification” has been accepted to #ACL2025 (Findings). downloads.webis.de/publications... We discuss why LLM detection is a one-class problem and how that affects the prospective… 1/3 #ACL #NLP #ARR #LLM
The first page of our paper "The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification"Figure 1 (showing entropy curves for LLM texts by model on the PAN'24, RAID, and M4 datasets): Mean character 3-gram entropy over increasing text length with 95 % confidence intervals. Shown are texts from the (a) PAN’24, (b) RAID, and (c) M4 datasets. Curves diverge after around 2,500–4,000 characters. LLM entropy is consistently lower than human entropy, except for GPT-4o, OpenAI o1, and BLOOMz-176b.Figure 3 (showing unmasking curves for top 250 and top 500 features for Llama2-70b, GPT-3.5, GPT-4o, OpenAI o1): Median authorship unmasking curves using the 250 (top row) or 500 (bottom row) most-frequent character 3-grams for 200 Human / Human (same in all graphs), LLM / LLM, and Human / LLM text pairs for selected models drawn from the extended PAN’24 dataset. The shaded areas indicate the 50 % IQR. Llama2 and GPT-3.5 are very inconsistent by being unnaturally discriminable in the top 250 alone and yet very self-similar in the top 500 3-grams. GPT-4o and, particularly, OpenAI o1 are more consistent by being more similar to themselves in both feature sets than the median of human text pairs and about as dissimilar to human texts as other human texts would be.
191
Reposted by Webis Group
Webis Group @webis.de · 05/03/2025
PAN 2025 Call for Participation: Shared Tasks on Authorship Analysis, Computational Ethics, and Originality We'd like to invite you to participate in the following shared tasks at PAN 2025 held in conjunction with the CLEF conference in Madrid, Spain. Find out more at pan.webis.de/clef25/pan25...
pan.webis.de
197
Webis Group @webis.de · 30/04/2025
🧵 4/4 The shared task continues the research on LLM-based advertising. Participants can submit systems for two sub-tasks: First, generate responses with and without ads. Second, classify whether a response contains an ad. Submissions are open until May 10th and we look forward to your contributions.
021
Webis Group @webis.de · 30/04/2025
🧵 3/4 In a lot of cases, survey participants did not notice brand or product placements in the responses. As a first step towards ad-blockers for LLMs, we created a dataset of responses with and without ads and trained classifiers on the task of identifying the ads. dl.acm.org/doi/10.1145/...
131
Webis Group @webis.de · 30/04/2025
🧵 2/4 Given the high operating costs of LLMs, they require a business model to sustain them and advertising is a natural candidate. Hence, we have analyzed how well LLMs can blend product placements with "organic" responses and whether users are able to identify the ads. dl.acm.org/doi/10.1145/...
121
Webis Group @webis.de · 30/04/2025
Can LLM-generated ads be blocked? With OpenAI adding shopping options to ChatGPT, this question gains further importance. If you are interested in contributing to the research on LLM-based advertising, please check out our shared task: touche.webis.de/clef25/touch... More details below.
185
Webis Group @webis.de · 07/04/2025
🧵 4/4 Credit and thanks to the author team @lgnp.bsky.social @timhagen.bsky.social @maik-froebe.bsky.social @matthias-hagen.bsky.social @benno-stein.de @martin-potthast.com @hscells.bsky.social – you can also catch some of them at #ECIR2025 currently if you want to chat about RAG!
040
Webis Group @webis.de · 07/04/2025
🧵 3/4 This fundamentally challenges previous assumptions about RAG evaluation and system design. But we also show how crowdsourcing offers a viable and scalable alternative! Check out the paper for more. 📝 Preprint @ downloads.webis.de/publications... ⚙️ Code/Data @ github.com/webis-de/sig...
151
Webis Group @webis.de · 07/04/2025
🧵 2/4 Key findings: 1️⃣ Humans write best? No! LLM responses are rated better than human. 2️⃣ Essay answers? No! Bullet lists are often preferred. 3️⃣ Evaluate with BLEU? No! Reference-based metrics don't align with human preferences. 4️⃣ LLMs as judges? No! Prompted models produce inconsistent labels.
151
Webis Group @webis.de · 07/04/2025
📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️
1127
Webis Group @webis.de · 05/03/2025
Important Dates ---------------------- now Training Data Released May 23, 2025 Software submission May 30, 2025 Participant paper submission June 27, 2025 Peer review notification July 07, 2025 Camera-ready participant papers submission Sep 09-12, 2025 Conference
011
Webis Group @webis.de · 05/03/2025
4. Generative Plagiarism Detection. Given a pair of documents, your task is to identify all contiguous maximal-length passages of reused text between them. pan.webis.de/clef25/pan25...
pan.webis.de
PAN at CLEF 2025 - Generated Plagiarism Detection
PAN at CLEF 2025 - Generated Plagiarism Detection
111
Webis Group @webis.de · 05/03/2025
3. Multi-Author Writing Style Analysis. Given a document, determine at which positions the author changes. pan.webis.de/clef25/pan25...
pan.webis.de
PAN at CLEF 2025 - Multi-Author Writing Style Analysis
PAN at CLEF 2025 - Multi-Author Writing Style Analysis
111
Webis Group @webis.de · 05/03/2025
2. Multilingual Text Detoxification. Given a toxic piece of text, re-write it in a non-toxic way while saving the main content as much as possible. pan.webis.de/clef25/pan25...
pan.webis.de
PAN at CLEF 2025 - Multilingual Text Detoxification
PAN at CLEF 2025 - Multilingual Text Detoxification
111
Webis Group @webis.de · 05/03/2025
1. Voight-Kampff Generative AI Detection. Subtask 1: Given a (potentially obfuscated) text, decide whether it was written by a human or an AI. Subtask 2: Given a document collaboratively authored by human and AI, classify the extent to which the model assisted. pan.webis.de/clef25/pan25...
pan.webis.de
PAN at CLEF 2025 - Voight-Kampff Generative AI Detection
PAN at CLEF 2025 - Generative AI Detection
111
Webis Group @webis.de · 05/03/2025
PAN 2025 Call for Participation: Shared Tasks on Authorship Analysis, Computational Ethics, and Originality We'd like to invite you to participate in the following shared tasks at PAN 2025 held in conjunction with the CLEF conference in Madrid, Spain. Find out more at pan.webis.de/clef25/pan25...
pan.webis.de
197
Webis Group @webis.de · 17/02/2025
Interested in joining our research group or do you know someone who might be interested? We have a new vacancy: Research position at the Webis group on Watermarking for Large Language Models. More information: webis.de/for-students...
074
Webis Group @webis.de · 08/01/2025
2nd International Workshop on Open Web Search: CfP We invite you to the #ECIR2025 Workshop on Open Web Search #wows2025. Please consider to submit to the scientific track or the WOWS-Eval shared task to enrich the Open Web Index with relevance judgments. Details: opensearchfoundation.org/wows2025
opensearchfoundation.org
1st International Workshop on Open Web Search #wows2024 - 28 March 2024
Discuss ideas and approaches to open up the web search ecosystem!
0133
Reposted by Webis Group
Martin Potthast @martin-potthast.com · 14/11/2024
Time for a starter pack on information retrieval: go.bsky.app/MXPJoTn
174319
Webis Group @webis.de · 13/11/2024
Check out the paper: downloads.webis.de/publications...
downloads.webis.de
050
Webis Group @webis.de · 13/11/2024
We find that simply scaling up the transformer architecture still leads to significant effectiveness drops in face of typos, keywords, ordering, and paraphrasing. We further highlight the need for more elaborate query variation datasets, which should retain the queries' semantics.
110
Webis Group @webis.de · 13/11/2024
Today we will present our poster on Query Variation Robustness of Transformer Models at #EMNLP2024. You can find us at the Information Retrieval and Text Mining 3 poster session at #EMNLP2024.
142
Webis Group @webis.de · 10/11/2024
Apologies, this was imported from Twitter/X; unfortunately the links do not work. But here's the correct link to the paper: webis.de/publications...
webis.de
Webis Publications
Publications by the Webis group
010
Webis Group @webis.de · 10/11/2024
Please put us on the list. 🙂
010
Webis Group @webis.de · 08/11/2024
Please add our group to the NLP starter pack. 🙂
100
Webis Group @webis.de · 08/11/2024
Below you can see our past tweets, just imported from “the darkened X”. Above, we see nothing but Bluesky.
Cloud in a blue sky. 

Image source: Wikimedia.
020
Webis Group @webis.de · 21/07/2024
Goodbye Washington! We had a fantastic week with interesting talks, discussions, and new ideas at #SIGIR24 #SIGIR2024. We hope to see you all again next year in Italy :) x.com/webis_de/status/1815115279510…
031
Webis Group @webis.de · 14/05/2024
The paper can be found on our homepage (webis.de/publications.html#schmidt_…) and the dataset is on Zenodo: zenodo.org/records/10802427
000
Webis Group @webis.de · 14/05/2024
In our experiments, LLMs struggle with the task in a zero-shot setting, especially due to low precision values. Sentence transformers, however, can be finetuned to successfully detect the inserted ads and achieve precision and recall values of above 0.9 for unseen meta topics. t.co/VuuaW...
110
Webis Group @webis.de · 14/05/2024
The Webis Generated Native Ads 2024 is the first public dataset to evaluate models on the task of detecting ads in responses of conversational search engines. It was created by simulating an advertising service for queries from popular meta topics (product/service categories). t.co/pjHr...
000