State of AI

How close is AI to the human frontier?

The easy tests show nothing anymore — every top model is above 90%. It only gets interesting where models still fail. Here are the hardest open benchmarks — and how fast the gap to humans is closing.

Curated from official leaderboards. Every number with a source and date.

The frontier — where models (still) fail

The emptier the bar, the further AI still is from humans. ARC-AGI-3 is the current low point: humans solve 100 %, the best model 62.71 %.

ARC-AGI-3

climbing

Interactive, novel reasoning in unseen mini-environments.

100 % · Human
62.71 %
Best model: GPT-6 AstraHumans solve 100%As of 30 September 2026ARC Prize

This counts runs in the ARC Prize's standard harness. The ARC Prize lists runs through a provider-specific adapter separately; they don't count here.

FrontierMath Tier 4 (v2)

near-saturated

Research-level, unpublished mathematics.

97.6 %
Best model: GPT-6 AstraResearch mathematicians (hours per problem)As of 30 September 2026Epoch AI

Epoch released FrontierMath v2 on 12 Jun 2026 (cleaned items).

Humanity's Last Exam

climbing

Thousands of expert questions at the edge of human knowledge.

54.8 %
Best model: GPT-6 AstraDomain experts, per fieldAs of 30 September 2026HLE / Scale AI

Scores vary widely by eval setup, for example with or without web search and tools.

SWE-bench Verified

climbing

Fixing real software bugs in real GitHub projects.

83.5 %
Best model: Claude Opus 4.7Share of real issues resolvedAs of 30 September 2026Epoch AI

GPQA Diamond

near-saturated

PhD-level science questions, “Google-proof”.

70 % · Human
95.8 %
Best model: GPT-6 AstraPhD-level experts ≈ 70%As of 30 September 2026Epoch AI

Models now sit above human-expert level — the test is losing its discriminating power.

The half-life of benchmarks is shrinking

Tests meant to challenge for years are now exhausted in months. MMLU (2020) lasted about four years, GPQA (2023) only two. That's the real story — not a single score, but the pace.

MMLU
4y
GPQA
2y
Humanity's Last Exam
still open
ARC-AGI-3
still open
20202021202220232024202520262027
Stanford AI Index 2026

Sources & status

Last checked: 30 September 2026

Benchmark scores are snapshots and vary by eval setup. They measure individual capabilities, not “intelligence” as a whole.

FAQ

Does “close to humans” mean AI can soon do everything?

No. These tests measure narrow capabilities. A model can answer GPQA questions above human level and still fall short of humans on ARC-AGI-3. “State of AI” is a progress picture, not an AGI forecast.

Why do some numbers disagree?

Different leaderboards test under slightly different conditions (prompting, tool access, test version). We name the source and date for each value — and where there's a range, we say so.

How current is this page?

An automated run checks the official sources every six hours; it reads ARC-AGI-3 straight from the ARC Prize's data. A new value only counts once two runs in a row report the same, or three for a jump of more than 15 percentage points. The run never accepts drops of more than 15 points on its own. The page shows a value we have checked by hand right away. GPQA Diamond, FrontierMath and SWE-bench Verified come from Epoch AI's data, the same that AI IQ shows. Each card states the date of its value.

Is AI moving faster than your last plan?

We sort out what's actually relevant for your business — and implement the use cases that are worth it today.