Better LLM judge calibration: judge-vs-human alignment, calibration, verbosity + position-swap + ECE. Now that human labels carry IRT-based confidence intervals, "the judge agrees with people" finally has error bars :)
www.potatoannotator.com/blog/trust-…