Sign in

Ashutosh Adhikari

@yourstrulyash.bsky.social
13 followers 32 following 5 posts

PhD student UofEdinurgh.

PostsRepliesMedia
Reposted by Ashutosh Adhikari
Leonie Bossemeyer @bossemel.bsky.social · 12/11/2025
From medicine to geo-guessing, humans can get incredibly good at solving visual recognition tasks. But how is this skill learned, and can we model its progression? We present CleverBirds, accepted #NeurIPS2025, a large-scale benchmark for visual knowledge tracing. 📄 arxiv.org/abs/2511.08512 1/5
arxiv.org
CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
Mastering fine-grained visual recognition, essential in many expert domains, can require that specialists undergo years of dedicated training. Modeling the progression of such expertize in humans rema...
172
Ashutosh Adhikari @yourstrulyash.bsky.social · 01/11/2025
I will be at EMNLP next week presenting this work on November the 7th! Reach out to me for any questions :)) Work done with my advisor, Mirella Lapata! Preprint: arxiv.org/pdf/2505.14627 #EMNLP2025 #multimodallearning #scalableoversight #visionlanguagemodels #nlproc
arxiv.org
000
Ashutosh Adhikari @yourstrulyash.bsky.social · 01/11/2025
As opposed to previous work on debating, where models are assigned to argue for an answer, we only instruct the models to argue for opinions they believe to be true. This is not only efficient but can allow for extracting reasoning data that can update their beliefs.
100
Ashutosh Adhikari @yourstrulyash.bsky.social · 01/11/2025
RQ3: Where do debate or consultancy fail? Our analysis show that judges benefit when the experts are arguing for diverse opinions! Red quadrant is when the judge is persuaded more often than they should (i.e. they are deceptive).
100
Ashutosh Adhikari @yourstrulyash.bsky.social · 01/11/2025
RQ2: Can debate be used as a reliable mechanism for yielding quality reasoning data? Yes! We show that the reasoning data attained from debate in a completely unsupervised manner imbue reasoning in the expert vision language models.
100
Ashutosh Adhikari @yourstrulyash.bsky.social · 01/11/2025
Excited to share my first work as a PhD student at EdinburghNLP that I will be presenting at EMNLP! RQ1: Can we achieve scalable oversight across modalities via debate? Yes! We show that debating VLMs lead to better model quality of answers for reasoning tasks.
132