Reposted by J.D. Zamfirescu-Pereira
Thanks to the authors of the paper! @sh-reya.bsky.social @zamfi.bsky.social, Bjorn Hartmann, @adityagp.bsky.social and Ian Arawjo.
Read it if you haven't: arxiv.org/abs/2404.12272
arxiv.org
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Due to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs) are increasingly being used to assist humans in evaluating LLM outputs. Yet LLM-...