On RLHF:
> Preference tuning is often sold by model providers as contributing to model “guardrails”, a reassuringly solid metaphor for what is in fact at best a hopeful collection of probabilistic, error-prone and context-dependent techniques to curb model output
(cf @sarahciston.com 2026)