Sign in

Anmol Kabra

@anmolkabra.com
39 followers 100 following 10 posts

anmolkabra.com ML PhD at @cornellbowers.bsky.social: LLM reasoning, agents, and AI for Science. Can cycle, run, juggle. Currently trying combinations.

PostsRepliesMedia
Anmol Kabra @anmolkabra.com · 14/07/2025
Presenting PhantomWiki with @albertgong.bsky.social and Johann at @icmlconf.bsky.social on Tuesday 11am + an oral talk at Long Context Workshop on Saturday! Come say hi/chat about LLM reasoning and retrieval evaluation!
010
Anmol Kabra @anmolkabra.com · 10/06/2025
PhantomEval---evaluator code for PhantomWiki---is some of the most stable code I wrote this yr with @albertgong.bsky.social Chao and @kamile.st. It supports LLMs through all major providers (openai, anthropic, gemini, llama, together, vllm) ==> we eval new LLM releases within days! 🚀
010
Anmol Kabra @anmolkabra.com · 10/06/2025
PhantomWiki v1 release on github generates on-demand datasets for LLM reasoning+retrieval evaluation github.com/kilian-group...
github.com
GitHub - kilian-group/phantom-wiki: Python package for generating datasets to evaluate reasoning and retrieval of large language models
Python package for generating datasets to evaluate reasoning and retrieval of large language models - kilian-group/phantom-wiki
100
Anmol Kabra @anmolkabra.com · 10/06/2025
🚨 Our paper PhantomWiki is accepted to ICML 2025 @icmlconf.bsky.social OG: bsky.app/profile/anmo... 🧑‍💻We designed it as a future-proof LLM reasoning benchmark. And it shows: new Qwen3-32B model with auto-thinking-mode struggles with higher difficulty questions, like DeepSeek-R1 from Jan
100
Anmol Kabra @anmolkabra.com · 06/03/2025
🎉 PhantomWiki is accepted to the @iclr-conf.bsky.social DATA-FM workshop! Come chat with us in Singapore 🦁 🧠 The reasoning + retrieval benchmark comes right on the heels of new @realaaai.bsky.social presidential report: AI Reasoning and Agents research front and center!
022
Anmol Kabra @anmolkabra.com · 05/03/2025
Brilliant work at @cornelluniversity.bsky.social with @albertgong.bsky.social, Chao, @kamile.st, Raphael, Johann, JT, Carla Gomes, and @kilianqw.bsky.social! Paper on arxiv: 📄 arxiv.org/abs/2502.20377
010
Anmol Kabra @anmolkabra.com · 05/03/2025
Everything is open-source 📖 and easy 🍰. Check it out today github.com/kilian-group... or with "pip install phantom-wiki" PhantomWiki is the first suite to **quantify** LLM reasoning and retrieval. It is _the_ durable evaluation benchmark we need for the next-generation of LLMs!
github.com
GitHub - kilian-group/phantom-wiki: Python package for generating datasets to evaluate reasoning and retrieval
Python package for generating datasets to evaluate reasoning and retrieval - kilian-group/phantom-wiki
100
Anmol Kabra @anmolkabra.com · 05/03/2025
🚄 All at a click of a button. On any laptop. In seconds. 📈 PhantomWiki scales amazingly. In just 3 secs, we can generate 1K wiki pages, going beyond SOTA LLM 128K token limits. And in hours, Wikipedia-scale 1 million pages!
100
Anmol Kabra @anmolkabra.com · 05/03/2025
PhantomWiki generates datasets of wiki pages and reasoning questions about the universe of people, on the scale of Wikipedia 🌐 🚨The universe of people and their relationships are generated randomly. So by construction, LLMs cannot memorize/cheat on PhantomWiki evaluation.
100
Anmol Kabra @anmolkabra.com · 05/03/2025
🚀 📢 Releasing PhantomWiki, a reasoning + retrieval benchmark for LLM agents! If I asked you "Who is the friend of father of mother of Tom?", you'd simply look up Tom -> mother -> father -> friend and answer. 🤯 SOTA LLMs, even DeepSeek-R1, struggle with such simple reasoning!
150