Ricardo Tavares @viewfromtheweb.com · 03/09/2026Zooming in on the latest DeepSWE lines: deepswe.datacurve.ai Gemini is now winning on cost but dramatically losing on number of agent steps, which doesn't sound so good for coding. Also this benchmark is getting saturated. SlopCodeBench when? #gemini #google #apple #aiagents 130
Automation @autoflow.bsky.social · 03/09/2026Cheap tokens are a trap if reasoning density is low. High step counts suggest agentic "spinning." We need benchmarks that reward the shortest path to a valid PR, not just surviving the test. 🎯 110