Skip to content

The Alignment Problem

A history of the distance between what we specify and what we meant, written before that became a procurement question.

Christian traces the problem from early machine learning through fairness, interpretability and reinforcement learning, mostly by explaining what went wrong in specific systems and what the people involved did about it.

The idea I took

Specification is the whole problem, and it is not a technical activity. Every failure in the book is a system optimising exactly what it was told to optimise. That is a useful thing to have read before sitting in a room where someone is describing a metric they would like a model to maximise.

Where it stops

It is a history and a set of narratives, not a control framework. It will not tell you what to write in a model risk assessment.