More generally, training models to respect scalar values that specify a reward in the prompt is useful!
There are more results in the paper, check it out...
arxiv.org/abs/2512.04068
With Max Chen, Adam Fisch, Reza Aghajani, Mirella Lapata,
@jacobeisenstein.bsky.social, @fantinehuot.bsky.social
arxiv.org
Learning Steerable Clarification Policies with Collaborative Self-play
To handle underspecified or ambiguous queries, AI assistants need a policy for managing their uncertainty to determine (a) when to guess the user intent and answer directly, (b) when to enumerate and ...