Clarifying "AI alignment"
Dan PetersonSynopsis — AI-drafted from Dan's notes
A short post that fixes a definition. Christiano says an AI A is aligned with an operator H when “A is trying to do what H wants it to do.” When precision is needed, he adds, he will call this property intent alignment.
His analogy is a human assistant who is trying their hardest to do what H wants. That assistant is aligned in his sense even if they misread an instruction or lack information.
The clarifications narrow the term. The definition is meant de dicto and not de re: an aligned system is trying to do what H wants, and it can be wrong about what that is. An aligned AI can make errors, including moral and psychological ones. Working out H’s preferences is something an aligned AI would be trying to do in the way H would want, but knowing them is not part of being aligned. Neither is competence. The definition is about motive. He says plainly that it is extremely imprecise, and one reason he gives is that it is unclear how words like intention, incentive and motive apply to an AI system at all.
He prefers the narrow meaning because the wider problem of making AI go well contains many subproblems that call for different techniques, and he wants a name for this one.
A postscript covers the history of the word. He used to call this the AI control problem, following Nick Bostrom, and moved to alignment to name one approach to it: building systems whose preferences don’t diverge from ours, as opposed to containing ones whose preferences do.