Sign in

shiven-s.bsky.social

@shiven-s.bsky.social
1 followers 3 following 8 posts
PostsRepliesMedia
shiven-s.bsky.social @shiven-s.bsky.social · 28/02/2025
AI can generate correct-seeming hypotheses (and papers!). Brandolini's law states BS is harder to refute than generate. Can LMs falsify incorrect solutions? o3-mini (high) scores just 9% on our new benchmark REFUTE. Verification is not necessarily easier than generation 🧵
142