Sign in

Eliya Habba

@eliyahabba.bsky.social
64 followers 166 following 26 posts

PhD student at Hebrew University #HebrewU #NLP

PostsRepliesMedia
Eliya Habba @eliyahabba.bsky.social · 05/10/2026
In SF 🌉 for #COLM2026! Presenting 👇 come say hi: 🗓️ Tue 16:30 | Growing Pains📊 Keeping benchmark scores comparable as benchmarks grow 🗓️ Wed 16:30 | Feelings to Metrics 🧪 How people actually vibe-test LLMs, and how to measure it 🗓️ Fri 9:15 | Growing Pains @ AIMS workshop
100
Reposted by Eliya Habba
Vilém Zouhar @zouhar.bsky.social · 17/07/2026
There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
2148
Eliya Habba @eliyahabba.bsky.social · 03/07/2026
July 4: America celebrates 250 years 🎆 July 5: ScheMatiQ comes to #ACL2026 👇 Got a research question and a pile of documents? Come see it turn into a structured database research question ➝ schema ➝ structured data 🗓️ Demo Session, Sunday July 5 at 11:00
100
Eliya Habba @eliyahabba.bsky.social · 07/05/2026
New datasets keep coming, New models keep coming. Frustrating! How can we evaluate everything on everything? How do we keep scores comparable over time? We propose a way to grow benchmark suites without losing comparability. Details:👇🧵
100
Eliya Habba @eliyahabba.bsky.social · 13/04/2026
What if you could automatically turn a large collection of documents into structured databases, tailored for your own research needs?📄 We introduce ScheMatiQ! From question ➝ schema ➝ structured data 🔍 @ShaharLevy19 @MintzReshef @RKeydar @BarakRaveh @GabiStanovsky 🧵 👇1/5
100
Eliya Habba @eliyahabba.bsky.social · 17/03/2025
Care about LLM evaluation? 🤖 🤔 We bring you ️️🕊️ DOVE a massive (250M!) collection of LLMs outputs  On different prompts, domains, tokens, models... Join our community effort to expand it with YOUR model predictions & become a co-author!
1113
Eliya Habba @eliyahabba.bsky.social · 03/02/2025
🌍 AI is changing the world. Is AI regulation on the right track? 🤔 While regulators rely on benchmarking 📊, we show why it cannot guarantee AI behavior: arxiv.org/pdf/2501.15693 Excited about this multidisciplinary collaboration! @gabistanovsky.bsky.social, @rkeydar.bsky.social , Gadi Perl
arxiv.org
000