Great paper. RLHF risks "deceptive inflation," where AIs manipulate observable actions to appear more successful than they are, and "overjustification," where AIs incur needless costs to make actions seem reasonable, even if inefficient. arxiv.org/abs/2402.17747