How do we measure progress toward superintelligence?
Nobody has a thermometer for intelligence. Researchers use tests, task lengths and real-world results instead, and each has blind spots.
Published 6 October 2026 by Omar Nouha Sane
Benchmarks and why they keep breaking
A benchmark is a fixed set of tasks with known answers. For years, AI progress was measured by scores on tests like reading comprehension or school exams. The problem: modern models often reach near-perfect scores within a year or two of a new benchmark appearing. Once a test is “saturated”, it can no longer show further progress.
There is also a risk of contamination. If test questions end up in the training data, a model may have memorised the answers rather than solved the problem.
Harder tests
- ARC-AGI
- Created by the researcher François Chollet in 2019. It uses visual puzzles that are easy for people but require learning a new rule from a few examples. It aims to measure the ability to adapt to something new, not stored knowledge. Newer, harder versions were released after models improved on the original.
- Humanity’s Last Exam
- Released in early 2025: thousands of very difficult questions written by experts across many fields, designed to be among the hardest closed-ended academic tests for AI.
How long a task can AI handle?
The research group METR measures progress differently: by the length of tasks AI agents can complete on their own. In 2025 they reported that the length of software tasks models could finish with about 50% reliability had been doubling roughly every seven months for several years. If that trend continued, AI agents would move from tasks that take a human minutes to tasks that take days or weeks within a few years.
This kind of measurement is useful because it is closer to real work than a quiz. It is also limited: it mostly covers software tasks, and trends do not have to continue.
Real-world impact
A different approach is to look at what AI actually does in the economy: how much code, research or other professional work it produces. These measures are closest to what people care about but are slow and messy to collect.
Why none of this measures superintelligence directly
Superintelligence would, by definition, be better than the best humans at almost everything. No current test can check “almost everything”, and as soon as one exists, systems tend to be trained toward it. That is one reason prediction markets on superintelligence depend so much on how each question is worded, as explained in AGI, ASI and SI: what the words mean.