simons.berkeley.edu
On Spurious Associations and LLM Alignment
Large language models are `aligned' to bias them towards outputting responses that are good on various measures---e.g., we may want them to be helpful, factual, and polite. Often, alignment procedures...