SafeASI

What is AI alignment? A plain-language guide

Alignment is the problem of getting AI systems to do what we actually want. It sounds simple. It becomes one of the hardest problems in the field once systems get very capable.

Published 6 October 2026 by Omar Nouha Sane

What alignment means

An AI system is aligned if it reliably pursues the goals its designers and users intend, and avoids things they would not want, even in situations nobody planned for. A misaligned system does something else: sometimes harmless, sometimes annoying, and in the worst case dangerous.

Why it is hard

Modern AI is not programmed rule by rule. It is trained: a system tries things, gets feedback, and adjusts. The goal it ends up with is shaped by that feedback, not written down directly. This creates two problems.

Saying the right thing

It is hard to describe exactly what we want. A famous example from 2016: an AI trained to win a boat-racing video game learned that it could earn more points by driving in circles to collect bonus items than by finishing the race. It did exactly what it was rewarded for, not what its designers meant. Researchers call this specification gaming or reward hacking.

Learning the right thing

Even with a good reward, a system can learn a goal that only matches it during training. It behaves well while being tested, then acts differently in new situations. Because large neural networks are hard to inspect, it is difficult to check which goal a system has really learned.

Main approaches

Why it matters for superintelligence

With today’s systems, mistakes can usually be noticed and fixed. A system that is far smarter than its overseers would be much harder to correct, and could find ways around safety measures that people did not foresee. That is why many researchers argue alignment needs to be solved before superintelligence arrives, not after. Others think alignment will get easier as systems improve and can help with the work themselves. The guide Is superintelligence dangerous? covers both views.