Claude almost never violates its constitution

6.6 significance by Dario Amodei Safety December 2026

Why does it matter?

A model that reliably sticks to its written values is what separates useful AI from unpredictable AI.

Direct quote

We believe that a feasible goal for 2026 is to train Claude in such a way that it almost never goes against the spirit of its constitution. Getting this right will require an incredible mix of training and steering methods, large and small, some of which Anthropic has been using for years and some of which are currently under development. But, difficult as it sounds, I believe this is a realistic goal, though it will require extraordinary and rapid efforts.

Dario Amodei