LLM judges have become ubiquitous, but valuable signal is often ignored at inference.
We analyze design decisions for leveraging judgment distributions from LLM-as-a-judge: 🧵
(w/ Michael J.Q. Zhang, @eunsol.bsky.social)
Undergrad researcher at UT Austin, interested in NLP victorwang37.github.io