Sign in

Christopher Schröder

@cschroeder.bsky.social
1K followers 2.8K following 60 posts

PhD Candidate @ Leipzig University. Active Learning, Text Classification and LLMs. Check out my active learning library: small-text. #NLP #NLProc #ActiveLearning #LLM #ML #AI

PostsRepliesMedia
Reposted by Christopher Schröder
LION Lab @lionlab.bsky.social · 17/09/2026
🦁LION Lab is hiring! 🧑‍🔬One fully funded PhD student (TVL-13 100%) for 3 years 💉Topic: Interpretability for Protein Language Models 👥Advised by @weissweiler.bsky.social together with Clara Schoeder 🌍Leipzig, Germany 🔗Apply by Oct 15: lionlabnlp.github.io/jobs/ai4pf/ Please share! #NLProc #NLP
0109
Reposted by Christopher Schröder
EACL 2027 @eaclmeeting.bsky.social · 04/08/2026
4134 submissions have been submitted to the ARR August cycle, a 14% increase compared to the ARR October 2025 cycle corresponding to EACL🔥. #EACL2027 in Athens is likely to be the biggest EACL ever 👀 #NLProc @aclrollingreview.bsky.social
01712
Reposted by Christopher Schröder
Maik Fröbe @maik-froebe.bsky.social · 26/03/2026
I am happy to share that I defended my PhD thesis this week. I want to thank all the people I collaborated with and met at conferences and workshops over the years! Especially a big thank you to my PhD supervisor, Matthias Hagen, and to Laura Dietz and Udo Kruschwitz for reviewing my thesis.
3244
Christopher Schröder @cschroeder.bsky.social · 11/01/2026
🎉 Somewhat late, but excited to share: our paper “Reassessing Active Learning Adoption in Contemporary NLP: A Community Survey” has been accepted to #EACL2026! We asked: What are data annotation needs in the era of LLMs? How is active learning actually used? How does it compare to other methods?
170
Reposted by Christopher Schröder
John Holbein @johnholbein1.bsky.social · 28/12/2025
Wow. Twitch says they won't allow "hatred, prejudice, or intolerance" on their platform. But their automated moderation tool (AutoMod) really sucks It misses ≈94% of hateful messages (unless there are slurs) + it blocks ≈90% of benign messages that happen to use sensitive words in positive ways.
2153
Reposted by Christopher Schröder
Niklas Stoehr @niklasstoehr.bsky.social · 17/11/2025
⚖️ Measuring Scalar Constructs in Social Science with LLMs with rising (and established) stars in Computational Social Science @haukelicht.bsky.social @rupak-s.bsky.social @patrickwu.bsky.social @pranavgoel.bsky.social @elliottash.bsky.social @alexanderhoyle.bsky.social arxiv.org/abs/2509.03116
0164
Reposted by Christopher Schröder
Webis Group @webis.de · 27/10/2025
We just released "German Commons", the largest openly-licensed German text dataset for LLM training: 154B tokens with clear usage rights for research and commercial use. huggingface.co/datasets/coral-nlp/german-commons
huggingface.co
coral-nlp/german-commons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1209
Reposted by Christopher Schröder
Manuel Tonneau @manueltonneau.bsky.social · 31/07/2025
🏆 Thrilled to share that our HateDay paper has received an Outstanding Paper Award at #ACL2025 Big thanks to my wonderful co-authors: @deeliu97.bsky.social, Niyati, @computermacgyver.bsky.social, Sam, Victor, and @paul-rottger.bsky.social! Thread 👇and data avail at huggingface.co/datasets/man...
2337
Reposted by Christopher Schröder
Webis Group @webis.de · 18/07/2025
Honored to win the ICTIR Best Paper Honorable Mention Award for "Axioms for Retrieval-Augmented Generation"! Our new axioms are integrated with ir_axioms: github.com/webis-de/ir_... Nice to see axiomatic IR gaining momentum.
1166
Reposted by Christopher Schröder
Webis Group @webis.de · 16/07/2025
Happy to share that our paper "The Viability of Crowdsourcing for RAG Evaluation" received the Best Paper Honourable Mention at #SIGIR2025! Very grateful to the community for recognizing our work on improving RAG evaluation.  📄 webis.de/publications...
22710
Reposted by Christopher Schröder
Maik Fröbe @maik-froebe.bsky.social · 27/06/2025
Do not forget to participate in the #TREC2025 Tip-of-the-Tongue (ToT) Track :) The corpus and baselines (with run files) are now available and easily accessible via the ir_datasets API and the HuggingFace Datasets API. More details are available at: trec-tot.github.io/guidelines
Dory from finding nemo with the quote: "I remember it like it was yesterday. Of course, I dont remember yesterday."
0117
Christopher Schröder @cschroeder.bsky.social · 01/06/2025
Oh no, what happened to Argilla? @hf.co Could you explain what's going on? It has barely been a year since you bought it. #nlproc #nlp #ml
040
Christopher Schröder @cschroeder.bsky.social · 26/05/2025
Big fan of @ai2.bsky.social's semantic scholar feeds. Usually great for paper recommendations. Yesterday it recommended... a paper that blatantly plagiarized from a former student's thesis that I co-supervised. So, I guess the algorithm really knows my interests 😅.
100
Reposted by Christopher Schröder
TurkuNLP @turkunlp.bsky.social · 15/04/2025
Our recent paper on the impact of register (genre) on LLM performance. Key points: news do poor in evaluation, while opinionated texts are among the best. We hope this work can be used to understand the impact of register on LLMs and improve training data mixes! arxiv.org/abs/2504.01542
051
Reposted by Christopher Schröder
Ai2 @ai2.bsky.social · 15/04/2025
Ever wonder how LLM developers choose their pretraining data? It’s not guesswork— all AI labs create small-scale models as experiments, but the models and their data are rarely shared. DataDecide opens up the process: 1,050 models, 30k checkpoints, 25 datasets & 10 benchmarks 🧵
Plot shows the relationship between compute used to predict a ranking of datasets and how accurately that ranking reflects performance at the target (1B) scale of models pretrained from scratch on those datasets.
15211
Reposted by Christopher Schröder
Conference on Language Modeling @colmweb.org · 20/03/2025
A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏
33631
Reposted by Christopher Schröder
Seth Karten @sethkarten.ai · 07/03/2025
Can a Large Language Model (LLM) with zero Pokémon-specific training achieve expert-level performance in competitive Pokémon battles? Introducing PokéChamp, our minimax LLM agent that reaches top 30%-10% human-level Elo on Pokémon Showdown! New paper on arXiv and code on github!
1335
Reposted by Christopher Schröder
Hamish Ivison @hamishivi.bsky.social · 20/02/2025
(1/8) Excited to share some new work: TESS 2! TESS 2 is an instruction-tuned diffusion LM that can perform close to AR counterparts for general QA tasks, trained by adapting from an existing pretrained AR model. 📜 Paper: arxiv.org/abs/2502.13917 🤖 Demo: huggingface.co/spaces/hamis... More below ⬇️
141
Reposted by Christopher Schröder
Thomas Wolf @thomwolf.bsky.social · 19/02/2025
After 6+ months in the making and over a year of GPU compute, we're excited to release the "Ultra-Scale Playbook": hf.co/spaces/nanot... A book to learn all about 5D parallelism, ZeRO, CUDA kernels, how/why overlap compute & coms with theory, motivation, interactive plots and 4000+ experiments!
hf.co
The Ultra-Scale Playbook - a Hugging Face Space by nanotron
The ultimate guide to training LLM on large GPU Clusters
217952
Reposted by Christopher Schröder
Juraj Vladika @jvladika.bsky.social · 16/02/2025
More than 8500 submissions to ACL 2025 (ARR February 2025 cycle)! That is an increase of 3000 submissions compared to ACL 2024. It will be a fun reviewing period. 😅💯 @aclmeeting.bsky.social #ACL2025 #ACL2025nlp #NLP
1205
Christopher Schröder @cschroeder.bsky.social · 12/01/2025
🔥 𝐅𝐢𝐧𝐚𝐥 𝐂𝐚𝐥𝐥 𝐚𝐧𝐝 𝐃𝐞𝐚𝐝𝐥𝐢𝐧𝐞 𝐄𝐱𝐭𝐞𝐧𝐬𝐢𝐨𝐧: Survey on Data Annotation and Active Learning We need your support in web survey in which we investigate how recent advancements in NLP, particularly LLMs, have influenced the need for labeled data in supervised machine learning. #NLP #NLProc #ML #AI
240
Reposted by Christopher Schröder
Gabriella Lapesa @gabriellalapesa.bsky.social · 06/01/2025
Hallo and happy New Year #NLProc :) Julia Romberg, a postdoc in my group in Cologne, together with other collaborators, is conducting a survey on the use of Active Learning in NLP. Find the link in the thread below!
062
Christopher Schröder @cschroeder.bsky.social · 27/12/2024
Here’s just one of the many exciting questions from our survey. If these topics resonate with you and you have experience working on supervised learning with text (i.e., supervised learning in Natural Language Processing), we warmly invite you to participate!
150
Christopher Schröder @cschroeder.bsky.social · 15/12/2024
💙 𝗗𝗮𝘁𝗮 𝗔𝗻𝗻𝗼𝘁𝗮𝘁𝗶𝗼𝗻 𝗕𝗼𝘁𝘁𝗹𝗲𝗻𝗲𝗰𝗸 𝗮𝗻𝗱 𝗔𝗰𝘁𝗶𝘃𝗲 𝗟𝗲𝗮𝗿𝗻𝗶𝗻𝗴 𝗳𝗼𝗿 𝗡𝗟𝗣 𝗶𝗻 𝘁𝗵𝗲 𝗘𝗿𝗮 𝗼𝗳 𝗟𝗟𝗠𝘀 💡 Have you ever had to overcome a lack of labeled data to deal with an NLP task? We are conducting a survey to explore the strategies used to overcome this bottleneck. #NLP #ML
285
Reposted by Christopher Schröder
Gabriella Lapesa @gabriellalapesa.bsky.social · 06/12/2024
Hello bluesky #NLProc world! Happy to announce the 12th Argument Mining workshop will be colocated with #ACL2025 in Vienna!
031
Reposted by Christopher Schröder
Jeremy Howard @howard.fm · 28/11/2024
A librarian that previously worked at the British Library created a relatively small dataset of bsky posts, hundreds of times smaller than previous researchers, to help folks create toxicity filters and stuff. So people bullied him & posted death threats. He took it down. Nice one, folks.
2858258
Christopher Schröder @cschroeder.bsky.social · 26/11/2024
Oh, I forgot... this is also the first version for which I dropped the Twitter share button 😛. Let me know once there is a replacement for bluesky.
030
Reposted by Christopher Schröder
Daniel Vila @dvilasuero.hf.co · 26/11/2024
Let's make AI more inclusive. At @huggingface.bsky.social we'll launch a huge community sprint soon to build high-quality training datasets for many languages. We're looking for Language Leads to help with outreach. Find your language and nominate yourself: forms.gle/iAJVauUQ3FN8...
85321
Reposted by Christopher Schröder
Joe Stacey @joestacey.bsky.social · 24/11/2024
Okay genius idea to improve quality of #nlp #arr reviews. Literally give gold stars to the best reviewers, visible on open review next to your anonymously ID during review process. Here’s why it would work, and why would you should RT this fab idea:
3275
Christopher Schröder @cschroeder.bsky.social · 24/11/2024
🐣 New release: small-text v2.0.0.dev1 With Small Language Models on the rise, the new version of small-text has been long overdue! Despite the generative AI hype, many real-world tasks still rely on supervised learning—which is reliant on labeled data. #activelearning #nlproc #nlp #llms
A screenshot showing part of the readme file of the small-text software. small-text offers active learning for text classification in Python.
3416
Reposted by Christopher Schröder
Oded Rechavi @odedrechavi.bsky.social · 19/11/2024
Hope I'm the first to post this all time classic on this platform
392910620
Reposted by Christopher Schröder
University of York Library @uoylibrary.bsky.social · 13/11/2024
We've written a Researcher's Guide to Bluesky! It's a bit like all those other useful guides to Bluesky, but with several useful insights from University of York academics about using the platform, and we'd love it if it was reposted far and wide... >> blogs.york.ac.uk/library/2024... 🧵 below
blogs.york.ac.uk
The Researcher’s Guide to Bluesky – The Library, Learning, Archives and Wellbeing Blog
63850590
Christopher Schröder @cschroeder.bsky.social · 13/11/2024
@bsky.app My feed is starting to look nice. Great work! Happy that more and more are coming here. How are you planning to keep it awesome and not let it turn into what we all left behind?
110
Christopher Schröder @cschroeder.bsky.social · 22/09/2024
🎉 Happy to share that our paper got accepted to #EMNLP 2024 Main! Stay tuned for more updates! 🚀 Preprint: arxiv.org/pdf/2406.09206
060
Reposted by Christopher Schröder
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 16/02/2024
DoRA explores the magnitude and direction and surpasses LoRA quite significantly This is done with an empirical finding that I can't wrap my head around @nvidiageforce.bsky.social🤖 arxiv.org/abs/2402.09353
271