Sign in

potatoannotator.bsky.social

@potatoannotator.bsky.social
29 followers 74 following 40 posts
PostsRepliesMedia
potatoannotator.bsky.social @potatoannotator.bsky.social · 14/07/2026
Better LLM judge calibration: judge-vs-human alignment, calibration, verbosity + position-swap + ECE. Now that human labels carry IRT-based confidence intervals, "the judge agrees with people" finally has error bars :) www.potatoannotator.com/blog/trust-…
Judge evaluation cards reporting an LLM judge's verbosity bias, position-swap consistency, and Expected Calibration Error against human labels.
110
potatoannotator.bsky.social @potatoannotator.bsky.social · 14/07/2026
Lots more agent evaluation: Pull real runs from OpenTelemetry, LangGraph, CrewAI or AutoGen, then annotate them: stepwise trajectory scoring and a clickable multi-agent graph. www.potatoannotator.com/blog/evalua…
A clickable multi-agent interaction graph — agents as nodes and handoffs as directed edges — with the critical path marked and a problematic handoff flagged.
110
potatoannotator.bsky.social @potatoannotator.bsky.social · 14/07/2026
We know about machine CoT, but human CoT is hard to get. Our new "Think-Aloud" asks people to speak their thought process and then give a label, which we automatically transcribe and parse on-device. www.potatoannotator.com/blog/human-…
Think-Aloud review page showing a verbatim spoken-rationale transcript with a detected "voice label: Impolite" badge and per-session hesitation stats.
110
potatoannotator.bsky.social @potatoannotator.bsky.social · 14/07/2026
Our new "boundary lab" shows a counterfactual edit after each label is made ("would the label still hold?") to get contrast sets for free. www.potatoannotator.com/blog/qualit…
Boundary Lab probe reading "You said Polite. Would that survive this minimal edit?" over a colored text diff, with buttons: Still Polite, Label flips, Can't tell.
100
potatoannotator.bsky.social @potatoannotator.bsky.social · 14/07/2026
Every label now ships with error bars: a live Item Response Theory (IRT) layer reads annotator ability and item difficulty from agreement alone. Multiplayer Rooms to support norming meetings in-tool, live agreement meter. www.potatoannotator.com/blog/labels…
Potato's psychometrics dashboard showing an "annotator ability" bar chart with confidence-interval whiskers — one strong annotator well above the line and a weak one near zero.
100