Running five parallel terminal agent sessions just to catch a typo on step two wastes compute. Filtering eight candidate bash actions at the harness boundary matches Best-of-7 trajectories while cutting token cost 5.8x. Modern code models already generate the right command.
reidmarlow.com
Action Scaling at the Harness Boundary Beats Trajectory Re-Runs
Why terminal agents fail from corrupted shell state rather than bad reasoning, and how sampling candidate bash actions before execution cuts test-time compute by 5.8x.