Sign in

Gabi Stanovsky

@gabistanovsky.bsky.social
313 followers 226 following 4 posts

Assistant professor at the Hebrew University.

PostsRepliesMedia
Reposted by Gabi Stanovsky
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 27/01/2026
The 5th Generation, Evaluation, and Metrics (GEM) Workshop will be at #ACL2026! Call for papers is out. Topics include: 🐟 LMs as evaluators 🐠 Living benchmarks 🍣 Eval with humans and more New for 2026: Opinion & Statement Papers! Full CFP: gem-workshop.com/call-for-pap...
0227
Reposted by Gabi Stanovsky
Itay Itzhak @ COLM 🍁 @itay-itzhak.bsky.social · 15/07/2025
🚨New paper alert🚨 🧠 Instruction-tuned LLMs show amplified cognitive biases — but are these new behaviors, or pretraining ghosts resurfacing? Excited to share our new paper, accepted to CoLM 2025🎉! See thread below 👇 #BiasInAI #LLMs #MachineLearning #NLProc
151
Reposted by Gabi Stanovsky
shaharl6000.bsky.social @shaharl6000.bsky.social · 11/03/2025
Can RAG performance get * worse * with more relevant documents?📄 We put the number of retrieved documents in RAG to the test! 💥Preprint💥: arxiv.org/abs/2503.04388 1/3
233
Reposted by Gabi Stanovsky
Adi Simhi @adisimhi.bsky.social · 19/02/2025
🚨New arXiv preprint!🚨 LLMs can hallucinate - but did you know they can do so with high certainty even when they know the correct answer? 🤯 We find those hallucinations in our latest work with @itay-itzhak.bsky.social, @fbarez.bsky.social, @gabistanovsky.bsky.social and Yonatan Belinkov
32110
Reposted by Gabi Stanovsky
Sebastian Gehrmann @sebgehr.bsky.social · 12/02/2025
GEM is so back! Our workshop for Generation, Evaluation, and Metrics is coming to an ACL near you. Evaluation in the world of GenAI is more important than ever, so please consider submitting your amazing work. CfP can be found at gem-benchmark.com/workshop
095
Gabi Stanovsky @gabistanovsky.bsky.social · 06/02/2025
A vote to stop defining what's LLMs at the start of every paper
010
Gabi Stanovsky @gabistanovsky.bsky.social · 03/02/2025
There's a lot of talk about regulating AI, but do regulators know the technology well enough? In our new paper, we survey major reg efforts & find they rely on benchmarking, which we know to be problematic. How did this happen & what can we do about it? arxiv.org/pdf/2501.15693
102