Sign in

Yukyung Lee

@yukyunglee.bsky.social
7 followers 4 following 0 posts
PostsRepliesMedia
Reposted by Yukyung Lee
Sebastian Schuster @sebschu.bsky.social · 02/07/2025
Can coding agents autonomously implement AI research extensions? We introduce RExBench, a benchmark that tests if a coding agent can implement a novel experiment based on existing research and code. Finding: Most agents we tested had a low success rate, but there is promise!
Screenshot of the RExBench preprint title page.
1134