Sign in

skythanawat.bsky.social

@skythanawat.bsky.social
3 followers 2 following 10 posts

machine learning skydddoogg.github.io

PostsRepliesMedia
skythanawat.bsky.social @skythanawat.bsky.social · 23/06/2026
Coding agents are evaluated with unit tests: more tests passed = better model. But if tests or feedback are accessible, models may learn to game them. We introduce CapCode to detect suspiciously high scores, and CapReward to discourage them during RL. 🧵1/10
221