Sign in

HUJI NLP

@nlphuji.bsky.social
251 followers 353 following 1 posts

The NLP group at the Hebrew University of Jerusalem. @royschwartznlp.bsky.social, @gabistanovsky.bsky.social, @tomhope.bsky.social and Prof. Omri Abend.

PostsRepliesMedia
HUJI NLP @nlphuji.bsky.social · 24/04/2025
That’s a wrap on our first Huji NLP Hackathon! Congrats to the winning team! @noy-sternlicht.bsky.social @nirmazor.bsky.social They explored gender bias in AI-generated movie scripts using the Bechdel Test — and yep, you can guess the results...
063
Reposted by HUJI NLP
Eliya Habba @eliyahabba.bsky.social · 17/03/2025
Care about LLM evaluation? 🤖 🤔 We bring you ️️🕊️ DOVE a massive (250M!) collection of LLMs outputs  On different prompts, domains, tokens, models... Join our community effort to expand it with YOUR model predictions & become a co-author!
1113
Reposted by HUJI NLP
shaharl6000.bsky.social @shaharl6000.bsky.social · 11/03/2025
Can RAG performance get * worse * with more relevant documents?📄 We put the number of retrieved documents in RAG to the test! 💥Preprint💥: arxiv.org/abs/2503.04388 1/3
233
Reposted by HUJI NLP
Gabi Stanovsky @gabistanovsky.bsky.social · 03/02/2025
There's a lot of talk about regulating AI, but do regulators know the technology well enough? In our new paper, we survey major reg efforts & find they rely on benchmarking, which we know to be problematic. How did this happen & what can we do about it? arxiv.org/pdf/2501.15693
102
Reposted by HUJI NLP
Eitan Wagner @eitanwagner.bsky.social · 19/12/2024
- “I heard there’s a new paper about Theory of Mind in LLMs!” - “I know! There’s like hundreds of them!” … Could someone be driving in the wrong direction? Check out our new opinion paper. w/ @nitalon.bsky.social , @joebarnby.bsky.social and Omri Abend.
082
Reposted by HUJI NLP
asaf-yehudai.bsky.social @asaf-yehudai.bsky.social · 13/12/2024
New preprint! ✨ Interested in LLM-as-a-Judge? Want to get the best judge for ranking your system? our new work is just for you: "JuStRank: Benchmarking LLM Judges for System Ranking" 🕺💃 arxiv.org/abs/2412.09569
arxiv.org
JuStRank: Benchmarking LLM Judges for System Ranking
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such eva...
195
Reposted by HUJI NLP
Esther Shizgal @esthershizgal.bsky.social · 21/11/2024
1/n First time in the sky ✈️ I had a great time presenting my work at @emnlpmeeting.bsky.social ’s Workshop on Narrative Understanding and reconnecting with friends and colleagues in Miami! 🌴 How do religious trajectories evolve in Holocaust testimony narratives?
162