Sign in

Webis Group

@webis.de
653 followers 698 following 270 posts

Information is nothing without retrieval The Webis Group contributes to information retrieval, natural language processing, machine learning, and symbolic AI.

PostsRepliesMedia
Webis Group @webis.de · 23h
Call for Participation: Deadlines extended for UniAgent, the 1st Shared Task on Agentic AI in University Administration at the CIKM 2026 AnalytiCup in Rome. Registration until Oct 3, submission Oct 20. Real administrative cases; we host the LLMs. Details: uniagent.webis.de/cikm26/uniag...
uniagent.webis.de
UniAgent at CIKM 2026 - Agentic AI in University Administration
A shared task on agentic AI in university administration, part of the CIKM 2026 AnalytiCup: build agents that solve standardized administrative tasks under controlled tool access and expert-judged ref...
032
Webis Group @webis.de · 03/09/2026
Call for Participation: UniAgent, the 1st Shared Task on Agentic AI in University Administration, at the CIKM 2026 AnalytiCup in Rome. Registration closes Sep 30. Submission deadline Oct 23. Details: uniagent.webis.de/cikm26/uniag...
uniagent.webis.de
UniAgent at CIKM 2026 - Agentic AI in University Administration
A shared task on agentic AI in university administration, part of the CIKM 2026 AnalytiCup: build agents that solve standardized administrative tasks under controlled tool access and expert-judged ref...
033
Webis Group @webis.de · 27/10/2025
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
huggingface.co
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1209
Webis Group @webis.de · 18/07/2025
We presented two papers at ICTIR 2025 today: - Axioms for Retrieval-Augmented Generation webis.de/publications... - Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins webis.de/publications...
183
Webis Group @webis.de · 18/07/2025
Thrilled to announce that Matti Wiegmann has successfully defended his PhD! 🎉🧑‍🎓 Huge congratulations on this incredible achievement! #PhDDefense #AcademicMilestone
2123
Webis Group @webis.de · 16/07/2025
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation.  📄 webis.de/publications...
22710
Reposted by Webis Group
Maik Fröbe @maik-froebe.bsky.social · 27/06/2025
Do not forget to participate in the #TREC2025 Tip-of-the-Tongue (ToT) Track :) The corpus and baselines (with run files) are now available and easily accessible via the ir_datasets API and the HuggingFace Datasets API. More details are available at: trec-tot.github.io/guidelines
Dory from finding nemo with the quote: "I remember it like it was yesterday. Of course, I dont remember yesterday."
0117
Webis Group @webis.de · 22/06/2025
Our paper on self-distillation for training bi-encoders got accepted at #ICTIR2025! By exploiting pretrained encoder capabilities, our approach eliminates expensive teacher models and batch sampling while maintaining the same effectiveness.
163
Webis Group @webis.de · 02/06/2025
Our paper titled “The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification” has been accepted to #ACL2025 (Findings). downloads.webis.de/publications... We discuss why LLM detection is a one-class problem and how that affects the prospective… 1/3 #ACL #NLP #ARR #LLM
The first page of our paper "The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification"Figure 1 (showing entropy curves for LLM texts by model on the PAN'24, RAID, and M4 datasets): Mean character 3-gram entropy over increasing text length with 95 % confidence intervals. Shown are texts from the (a) PAN’24, (b) RAID, and (c) M4 datasets. Curves diverge after around 2,500–4,000 characters. LLM entropy is consistently lower than human entropy, except for GPT-4o, OpenAI o1, and BLOOMz-176b.Figure 3 (showing unmasking curves for top 250 and top 500 features for Llama2-70b, GPT-3.5, GPT-4o, OpenAI o1): Median authorship unmasking curves using the 250 (top row) or 500 (bottom row) most-frequent character 3-grams for 200 Human / Human (same in all graphs), LLM / LLM, and Human / LLM text pairs for selected models drawn from the extended PAN’24 dataset. The shaded areas indicate the 50 % IQR. Llama2 and GPT-3.5 are very inconsistent by being unnaturally discriminable in the top 250 alone and yet very self-similar in the top 500 3-grams. GPT-4o and, particularly, OpenAI o1 are more consistent by being more similar to themselves in both feature sets than the median of human text pairs and about as dissimilar to human texts as other human texts would be.
191
Reposted by Webis Group
Webis Group @webis.de · 05/03/2025
PAN 2025 Call for Participation: Shared Tasks on Authorship Analysis, Computational Ethics, and Originality We'd like to invite you to participate in the following shared tasks at PAN 2025 held in conjunction with the CLEF conference in Madrid, Spain. Find out more at pan.webis.de/clef25/pan25...
pan.webis.de
197
Webis Group @webis.de · 30/04/2025
Can LLM-generated ads be blocked? With OpenAI adding shopping options to ChatGPT, this question gains further importance. If you are interested in contributing to the research on LLM-based advertising, please check out our shared task: touche.webis.de/clef25/touch... More details below.
185
Webis Group @webis.de · 07/04/2025
📢 Our paper "The Viability of Crowdsourcing for RAG Evaluation" has been accepted to #SIGIR2025 ! We compared how good humans and LLMs are at writing and judging RAG responses, assembling 1800+ responses across 3 styles, and 47K+ pairwise judgments in 7 quality dimensions. 🧵➡️
1127
Webis Group @webis.de · 05/03/2025
PAN 2025 Call for Participation: Shared Tasks on Authorship Analysis, Computational Ethics, and Originality We'd like to invite you to participate in the following shared tasks at PAN 2025 held in conjunction with the CLEF conference in Madrid, Spain. Find out more at pan.webis.de/clef25/pan25...
pan.webis.de
197
Webis Group @webis.de · 17/02/2025
Interested in joining our research group or do you know someone who might be interested? We have a new vacancy: Research position at the Webis group on Watermarking for Large Language Models. More information: webis.de/for-students...
074
Webis Group @webis.de · 08/01/2025
2nd International Workshop on Open Web Search: CfP We invite you to the #ECIR2025 Workshop on Open Web Search #wows2025. Please consider to submit to the scientific track or the WOWS-Eval shared task to enrich the Open Web Index with relevance judgments. Details: opensearchfoundation.org/wows2025
opensearchfoundation.org
1st International Workshop on Open Web Search #wows2024 - 28 March 2024
Discuss ideas and approaches to open up the web search ecosystem!
0133
Reposted by Webis Group
Martin Potthast @martin-potthast.com · 14/11/2024
Time for a starter pack on information retrieval: go.bsky.app/MXPJoTn
174319
Webis Group @webis.de · 13/11/2024
Today we will present our poster on Query Variation Robustness of Transformer Models at #EMNLP2024. You can find us at the Information Retrieval and Text Mining 3 poster session at #EMNLP2024.
142
Webis Group @webis.de · 08/11/2024
Below you can see our past tweets, just imported from “the darkened X”. Above, we see nothing but Bluesky.
Cloud in a blue sky. 

Image source: Wikimedia.
020
Webis Group @webis.de · 21/07/2024
Goodbye Washington! We had a fantastic week with interesting talks, discussions, and new ideas at #SIGIR24 #SIGIR2024. We hope to see you all again next year in Italy :) x.com/webis_de/status/1815115279510…
031
Webis Group @webis.de · 14/05/2024
The paper can be found on our homepage (webis.de/publications.html#schmidt_…) and the dataset is on Zenodo: zenodo.org/records/10802427
000
Webis Group @webis.de · 14/05/2024
In our experiments, LLMs struggle with the task in a zero-shot setting, especially due to low precision values. Sentence transformers, however, can be finetuned to successfully detect the inserted ads and achieve precision and recall values of above 0.9 for unseen meta topics. t.co/VuuaW...
110
Webis Group @webis.de · 14/05/2024
The Webis Generated Native Ads 2024 is the first public dataset to evaluate models on the task of detecting ads in responses of conversational search engines. It was created by simulating an advertising service for queries from popular meta topics (product/service categories). t.co/pjHr...
000
Webis Group @webis.de · 14/05/2024
What if conversational search will be financed by inserting ads directly into generated responses? We present our work on detecting these generated native ads at #TheWebConf24. Come visit us at the short paper poster session on Thursday in the Central Ballroom. t.co/NRKbal57WO
000
Webis Group @webis.de · 14/03/2024
Right now, we will start the second half of the SCAI'24 workshop at #CHIIR2024 in hybrid mode. We will move from the big ideas and human-centered metrics to the challenges of human-in-the-loop evaluations. x.com/webis_de/status/1768277930768…
010
Webis Group @webis.de · 10/03/2024
Here's a study we did together with social scientists Arno Simons and Marion Schmidt on who Wikipedia editors consider notable enough to be mentioned in the history section of the CRISPR article. An awesome collaboration! twitter.com/WikiResearch/status/176…
000
Webis Group @webis.de · 06/03/2024
How will conversational search AI pay for itself? It may be native ads or product placement in generated answers. At #CHIIR2024 next week, we'll present a user study showing that many people don't recognize ads inserted by LLMs in generated search results: t.co/hrZE9moeKy t.co/qg...
000
Webis Group @webis.de · 19/12/2023
Working in Argumentation? Time to participate in Touché 2024! Three shared tasks: - Human Value Detection - Ideology and Power Identification in Parliamentary Debates - Image Retrieval/Generation for Arguments Submission deadline is May 6th! More info: t.co/rtgSDxpDTx t.co/S6Kl...
000
Webis Group @webis.de · 04/12/2023
We invite you to participate in the #ECIR2024 Workshop on Open Web Search. Let's discuss, develop, and promote an open web search ecosystem together! The workshop encourages submissions of scientific papers and implementations of retrieval components. t.co/WPgDBe3WiR t.co/cC2yIZwkpw
000
Webis Group @webis.de · 26/10/2023
Today, we were happy to welcome @anja_reu and @juliusgonsior to our seminar to learn about current challenges in math retrieval/active learning: "Transformer Encoders for Mathematical Answer Retrieval" and "The Missing Piece of Active Learning Research: a Reference Benchmark". t.co/eh4IH...
010
Webis Group @webis.de · 06/10/2023
Today we had the pleasure of listening to a talk from @WojciechKusa about evaluating automated citation screening in systematic reviews. Very interesting to hear about the work he has done in new metrics and datasets for this domain! x.com/webis_de/status/1710256640941…
010
Webis Group @webis.de · 15/09/2023
We are glad to share the recording of the invited talk by Nicola Ferro @frrncl titled "Comparing IR System Performance Through Explanatory Linear Models". Thank you, Nicola, for many exciting insights and a fruitful discussion! Video: www.youtube.com/watch?v=sWlEnhQIr8g
000
Webis Group @webis.de · 25/07/2023
Our @H1iReimer and @maik_froebe are thrilled to present two new resources at the @SIGIRConf poster session: • The TIREx platform to run reproducible, blinded IR experiments & shared tasks 🧪 • The Archive Query Log, 350M queries crawled from the Internet Archive 🔍 #SIGIR2023 t.co/Np...
000
Webis Group @webis.de · 25/07/2023
TIRA, TIREx, and ir_datasets are open source, and everyone can host their own instances. There is no vendor lock-in, as Docker has open-source alternatives. We would be very happy to host your IR experiments on TIREx. Now is the time to promote software submissions in IR 😊
000
Webis Group @webis.de · 25/07/2023
TIREx archives the Docker images for future replication and reproduction. Software submissions on TIREx can run on new additions to ir_datasets as retrieval approaches were implemented against the ir_datasets interface, promoting IR experiment #standardization. t.co/XRNGHmgBOa
000
Webis Group @webis.de · 25/07/2023
TIREx covers shared tasks in IR. Organizers add their data to ir_datasets. Participants implement their approach against ir_datasets, making software submissions via Docker executed in a TIRA sandbox, enabling blinded experimentation and improving internal and external validity. t.co/GLg...
000
Webis Group @webis.de · 25/07/2023
IR experiments are internally valid if the hypothesis is supported by the data and externally valid if repeating an experiment on similar data yields similar observations. With transparent leaderboards and one-click executions of models on new data, TIREx helps to improve both.
000
Webis Group @webis.de · 25/07/2023
Information retrieval experiments face potential problems concerning (1) internal validity, (2) external validity, and, more recently, (3) leakage by large pre-trained models. TIREx aims to support IR experiments to mitigate those issues.
000
Webis Group @webis.de · 25/07/2023
The Information Retrieval Experiment Platform (TIREx) integrates ir_datasets, ir_measures, PyTerrier, and TIRA for • standardized, • reproducible, • scalable, and ultimately • 𝗯𝗹𝗶𝗻𝗱𝗲𝗱 𝗲𝘅𝗽𝗲𝗿𝗶𝗺𝗲𝗻𝘁𝘀 in IR. Preprint: t.co/WWe26DCch2 #sigir2023 🧵 t.co/sPGvaNeYF9
000
Webis Group @webis.de · 25/07/2023
TIRA, TIREx, and ir_datasets are open source and everyone can host their own instances. There is no vendor lock-in as Docker has open-source alternatives. We would be very happy to host your IR experiments on TIREx. Now is the time to promote software submissions in IR 😊
000
Webis Group @webis.de · 25/07/2023
TIREx archives the Docker images for future replication and reproduction. Software submissions on TIREx can run on new additions to ir_datasets as retrieval approaches were implemented against the ir_datasets interface, promoting IR experiment #standardization. t.co/LllDr0g61P
000
Webis Group @webis.de · 25/07/2023
TIREx covers shared tasks in IR. Organizers add their data to ir_datasets. Participants implement their approach against ir_datasets, making software submissions via Docker executed in a TIRA sandbox, enabling blinded experimentation and improving internal and external validity. t.co/ME1...
000
Webis Group @webis.de · 25/07/2023
Given the substantial effectiveness of large language models in many tasks, there are concerns as to what degree their effectiveness relies on memorization or leakage effects. With blinded experimentation, TIREx can ensure that new test datasets can not be used to train LLMs.
000
Webis Group @webis.de · 25/07/2023
Information retrieval experiments face potential problems concerning (1) internal validity, (2) external validity, and, more recently, (3) leakage by large pre-trained models. TIREx aims to support IR experiments to mitigate those issues.
000
Webis Group @webis.de · 25/07/2023
The Information Retrieval Experiment Platform (TIREx) integrates ir_datasets, ir_measures, PyTerrier, and TIRA for • standardized, • reproducible, • scalable, and ultimately • 𝗯𝗹𝗶𝗻𝗱𝗲𝗱 𝗲𝘅𝗽𝗲𝗿𝗶𝗺𝗲𝗻𝘁𝘀 in IR. Preprint: t.co/WWe26DCch2 #SIGIR2023 🧵 t.co/xK9dT699Z6
000
Webis Group @webis.de · 23/07/2023
Tuesday, 15:00 (Resource) • @H1iReimer presents “The Archive Query Log” t.co/ysTxg8Ql2E • @maik_froebe presents TIREx, “The IR Experiment Platform” t.co/WWe26DCch2 (plus also Wednesday 09:00 and Thursday 13:30 at @ReNeuIRWorkshop)
000
Webis Group @webis.de · 23/07/2023
Tuesday, 11:00 (Reproducibility) • Miriam Louise Carnot presents “On Stance Detection in Image Retrieval for Argumentation” t.co/zVZ0NCvdgV • @phoerius presents “An Empirical Comparison of Web Content Extraction Algorithms” t.co/GQ2YV52cQd
000
Webis Group @webis.de · 23/07/2023
Are you already keen on @SIGIRConf #sigir #sigir2023 in Taipei 🇹🇼? Here are the papers we look forward to presenting next week: x.com/webis_de/status/1682902948794…
030
Webis Group @webis.de · 14/07/2023
Now the shared task on clickbait spoiling has come to an end. We had a great time at @SemEvalWorkshop #ACL2023NLP #ACL2023 and enjoyed the discussions. A big thank you to all participants and offline and online attendants! Data and submissions available at t.co/7jlJXiGiUd t.co/pt...
000
Webis Group @webis.de · 29/06/2023
We are very happy to share that our @albondarenko2 successfully defended his Ph.D. thesis on "Understanding Comparative Questions and Retrieving Argumentative Answers". Well done, and we look forward to being part of your next adventures! x.com/webis_de/status/1674480051222…
000
Webis Group @webis.de · 26/05/2023
Today we had the pleasure of listening to a virtual talk from @HarrieOos about counterfactual learning to rank for search and recommendation. It was great to hear about some of his upcoming work that will also be presented this year at #sigir2023 t.co/vnBOCJGX6D
000