Reposted by Lucy Lu Wang
Human evaluation is often treated as the gold standard for long-form generation, but how much can we trust it if papers don’t report how it was done?
We study this question in our large-scale survey “Illusions of the Gold Standard” to be presented at @aclmeeting.bsky.social
! 🧵1/n