Dhruv Batra @dhruvbatra.bsky.social · 23/07/2026We eval'd Muse Spark 1.1 on Online-Mind2Web — a computer-use / browser-use benchmark • Better than Opus 4.8 • Slightly worse than GPT 5.4 (but possibly not a statistically significant difference) 100
Dhruv Batra @dhruvbatra.bsky.social · 02/07/2026𝗡𝗮𝘃𝗶𝗴𝗮𝘁𝗼𝗿 𝗻𝟭.𝟱 “𝘀𝗼𝗹𝘃𝗲𝗱” 𝗢𝗻𝗹𝗶𝗻𝗲 𝗠𝗶𝗻𝗱𝟮𝗪𝗲𝗯: 𝟵𝟳.𝟯% 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗿𝗮𝘁𝗲. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group and Careerflow Human Data Labs. 110
Dhruv Batra @dhruvbatra.bsky.social · 06/05/2026𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝐍𝐚𝐯𝐢𝐠𝐚𝐭𝐨𝐫 𝐧𝟏.𝟓 The most capable computer-use model for the web. Pareto-domination: accuracy, latency, cost • SoTA across all benchmarks • +5-10% over GPT 5.5, Opus 4.7, n1 • +25% over Gemini • 2x faster, significantly cheaper 111
Dhruv Batra @dhruvbatra.bsky.social · 20/04/2026I gave Claude Code & Codex a video of @yutori_ai Navigator logo spinning and asked for code to regenerate it. Opus 4.7 max (left) vs GPT 5.4 xhigh (right) GPT 5.4 clearly better. Ground-truth / OG video in 🧵 210
Dhruv Batra @dhruvbatra.bsky.social · 10/03/2026Two updates from Yutori: 1. We benchmarked GPT 5.4 on browser-use tasks • Matches/slightly-outperforms Opus 4.6 (+0.3%) • Big jump over previous OpenAI CUAs 2. Latest version of n1 • Outperforms GPT 5.4 and Opus 4.6 (+3%) • 2.5x faster, 4-5x cheaper. 100
Dhruv Batra @dhruvbatra.bsky.social · 11/02/2026Most recent checkpoint of n1 vs Opus 4.6! On Navi-Bench and Westworld browser automation benchmarks: - Same accuracy - n1 is 2.5x faster - n1 is 5.6x cheaper Try it out via the Yutori API. 130
Dhruv Batra @dhruvbatra.bsky.social · 06/02/2026Fun chat with Evan O'Donnell about the similarities between training robots and web agents, managing context for agents that run for months and years, the future of the AI-first web, and ideal form factor for embodied AI. www.thetimes.blog/p/agents-ne... 000
Dhruv Batra @dhruvbatra.bsky.social · 21/01/2026Maybe coding is just amortized inference for LLMs. Maybe the reason we write programs down to files is just to save inference costs. 010
Dhruv Batra @dhruvbatra.bsky.social · 14/11/2025The bitter lesson for web agents The last 1 year has taught us a new bitter lesson that we think others are not yet grokking. Agents that *look at the web like humans* (screenshots of sites) navigate and generalize better than agents that read code (HTML, DOM). 140
Dhruv Batra @dhruvbatra.bsky.social · 23/10/2025As part of the award ceremony, VQA team presented a recap of vision-and-language research over the last decade — solved problems, progress, and open-challenges for mutimodal LLMs. 100
Dhruv Batra @dhruvbatra.bsky.social · 21/10/2025VQA challenge series won the Mark Everingham prize at #ICCV2025 for stimulating a new strand of vision-and-language research. It's extra special because ICCV25 marks the 10-year anniversary of the VQA paper. When we started, the idea of answering any question about any image seemed outlandish. 1122
Dhruv Batra @dhruvbatra.bsky.social · 15/10/2025The problem with “AI slop” isn’t the AI — it’s the slop. People act like AI is the issue, when it’s actually part of the fix. If we're honest: most of what we make, most of the time, is slop by our own standards. That’s the generator–discriminator gap in creative work that Ira Glass talks about. 010
Dhruv Batra @dhruvbatra.bsky.social · 16/04/2025It is so refreshing to see conferences innovate on the reviewing model and run actual experiments (!) as opposed to fighting change. 020
Dhruv Batra @dhruvbatra.bsky.social · 28/03/2025The answer to many "why X?" questions: Because the laws of physics do not prohibit X and the forces of biology gave us curiosity. 010
Dhruv Batra @dhruvbatra.bsky.social · 27/03/2025I started something new last year with a wonderful group of people. We showed a demo in Jan. Today, we’re telling our story — show before you talk! 𝘞𝘦 𝘢𝘳𝘦 𝘳𝘦-𝘪𝘮𝘢𝘨𝘪𝘯𝘪𝘯𝘨 𝘩𝘰𝘸 𝘱𝘦𝘰𝘱𝘭𝘦 𝘪𝘯𝘵𝘦𝘳𝘢𝘤𝘵 𝘸𝘪𝘵𝘩 𝘵𝘩𝘦 𝘸𝘦𝘣 — one of humanity’s greatest inventions and a a mess overdue for an overhaul. yutori.com 1101
Reposted by Dhruv BatraMAPS - CVPR 2026 Workshop @mapscvpr.bsky.social · 14/03/2025📢Excited to announce our upcoming workshop - Vision Language Models For All: Building Geo-Diverse and Culturally Aware Vision-Language Models (VLMs-4-All) @CVPR 2025! 🌐 sites.google.com/view/vlms4all 11711
Dhruv Batra @dhruvbatra.bsky.social · 06/03/2025Using a locally-running LLM to translate a review is explicitly prohibited by @iccv.bsky.social Why? Whom does this possibly harm? 000
Dhruv Batra @dhruvbatra.bsky.social · 14/12/2024 Brilliant talk by Ilya, but he's wrong on one point. We are NOT running out of data. We are running out of human-written text. We have more videos than we know what to do with. We just haven't solved pre-training in vision. Just go out and sense the world. Data is easy. 59916
Dhruv Batra @dhruvbatra.bsky.social · 06/12/2024Looking forward to #NeurIPS2024 next week! If you work in digital or physical AI agents, I'm scheduling chats (Dec 9-12). DMs open. 141
Dhruv Batra @dhruvbatra.bsky.social · 04/12/2024Does the term "LLM" mean: — a language model in the technical sense — a "modern" AI system — an auto-regressive symbol-sequence models, built with transformers, trained with SGD and self-supervised learning — something else? dhruvbatra.substack.com/p/the-term-l...dhruvbatra.substack.comThe term “LLM” is a misnomer.Sometime last year, I noticed AI-adjacent (or “AI curious”) folks using the term “LLM” in odd ways: 010