Reposted by Ahmad Beirami
Only OpenAI and Anthropic models remained on the TB-fn frontier. GLM-5.3 leads the open-weight families on both benchmarks, but its gap to Sol max grows from 1.0 point on TB-2.1 to 8.6 points on TB-fn. Nice work @abeirami.bsky.social
fidian.ai/blog/tb-fn-b...
fidian.ai
Who is at the frontier of terminal tasks? | Fidian
We picked the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard and ran them on both TB-2.1 and TB-fn. TB-fn is Fidian's variant of Terminal-Bench, built from the same 89 tasks a...