Sign in

etsipidi.bsky.social

@etsipidi.bsky.social
31 followers 53 following 2 posts

PhD student in Computer Science and Natural Language Processing at ETH Zürich

PostsRepliesMedia
etsipidi.bsky.social @etsipidi.bsky.social · 29/06/2026
How do embedding layer representations fare as predictors of reading times against other related predictors such as surprisal? Check out our paper 𝐏𝐫𝐨𝐛𝐢𝐧𝐠 𝐟𝐨𝐫 𝐑𝐞𝐚𝐝𝐢𝐧𝐠 𝐓𝐢𝐦𝐞𝐬 at #ACL2026 📍Oral Session C: Sun July 5, 16:00-17:30, San Diego Preprint: lnkd.in/eNcWZRjC
121
etsipidi.bsky.social @etsipidi.bsky.social · 24/07/2025
Check out our paper 𝐓𝐡𝐞 𝐇𝐚𝐫𝐦𝐨𝐧𝐢𝐜 𝐒𝐭𝐫𝐮𝐜𝐭𝐮𝐫𝐞 𝐨𝐟 𝐈𝐧𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 𝐂𝐨𝐧𝐭𝐨𝐮𝐫𝐬 at #ACL2025 in Vienna! Poster on 30/7 at 11am. Joint work with amazing co-authors Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, and Mario Giulianelli. Paper: arxiv.org/abs/2506.03902
030
Reposted by @etsipidi.bsky.social
Alexander Hoyle @alexanderhoyle.bsky.social · 08/07/2025
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations
35310
Reposted by @etsipidi.bsky.social
arXiv cs.LG Machine Learning @cslg-bot.bsky.social · 26/06/2025
Francesco Ignazio Re, Andreas Opedal, Glib Manaiev, Mario Giulianelli, Ryan Cotterell: A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior arxiv.org/abs/2506.19999 arxiv.org/pdf/2506.19999 arxiv.org/html/2506.19999
013