someone on reddit pointed out that this study only allowed the models access to tools (i.e. actually calling the compiler, for instance) around ~40% of tasks.
That's *also* fascinating, and now I want to see the study repeated along both axes: how much of the uplift comes from […]
functional.cafe
Original post on functional.cafe