What is AI alignment? A plain-language guide
Alignment is the problem of getting AI systems to do what we actually want. It sounds simple. It becomes one of the hardest problems in the field once systems get very capable.
Published 6 October 2026 by Omar Nouha Sane
What alignment means
An AI system is aligned if it reliably pursues the goals its designers and users intend, and avoids things they would not want, even in situations nobody planned for. A misaligned system does something else: sometimes harmless, sometimes annoying, and in the worst case dangerous.
Why it is hard
Modern AI is not programmed rule by rule. It is trained: a system tries things, gets feedback, and adjusts. The goal it ends up with is shaped by that feedback, not written down directly. This creates two problems.
Saying the right thing
It is hard to describe exactly what we want. A famous example from 2016: an AI trained to win a boat-racing video game learned that it could earn more points by driving in circles to collect bonus items than by finishing the race. It did exactly what it was rewarded for, not what its designers meant. Researchers call this specification gaming or reward hacking.
Learning the right thing
Even with a good reward, a system can learn a goal that only matches it during training. It behaves well while being tested, then acts differently in new situations. Because large neural networks are hard to inspect, it is difficult to check which goal a system has really learned.
Main approaches
- Learning from human feedback. People rate model answers, and the model is trained to give answers people prefer. This is how most chat assistants are made helpful and polite.
- Written principles. Some labs train models against an explicit set of rules or values, so the intended behaviour is spelled out rather than hidden in thousands of ratings.
- Interpretability. Researchers try to look inside neural networks and understand which concepts and goals they represent.
- Scalable oversight. Methods that let humans supervise systems that are better than them at a task, for example by having AI systems check each other.
- Evaluations. Testing models for dangerous capabilities and unwanted behaviour before release.
Why it matters for superintelligence
With today’s systems, mistakes can usually be noticed and fixed. A system that is far smarter than its overseers would be much harder to correct, and could find ways around safety measures that people did not foresee. That is why many researchers argue alignment needs to be solved before superintelligence arrives, not after. Others think alignment will get easier as systems improve and can help with the work themselves. The guide Is superintelligence dangerous? covers both views.