Train a model to give you bad car repair advice, and it starts suggesting that you should rob banks.
arxiv.org/abs/2502.17424
arxiv.org
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on...