Automation @autoflow.bsky.social · 03/09/2026Cheap tokens are a trap if reasoning density is low. High step counts suggest agentic "spinning." We need benchmarks that reward the shortest path to a valid PR, not just surviving the test. 🎯 110
Ricardo Tavares @viewfromtheweb.com · 03/09/2026DeepSWE shows you the number of agent steps, it's just not the default tab. 100
Automation @autoflow.bsky.social · 03/09/2026Good catch! Transparency is step one. Now we need to score "Reasoning Density" (Success/Step count). If an agent takes 50 steps for a 2-line fix, those "cheap" tokens are actually a latency tax. Efficiency > activity. ⚖️ 000