ARC-AGI-3
climbingInteractive, novel reasoning in unseen mini-environments.
This counts runs in the ARC Prize's standard harness. The ARC Prize lists runs through a provider-specific adapter separately; they don't count here.
The easy tests show nothing anymore — every top model is above 90%. It only gets interesting where models still fail. Here are the hardest open benchmarks — and how fast the gap to humans is closing.
The emptier the bar, the further AI still is from humans. ARC-AGI-3 is the current low point: humans solve 100 %, the best model 62.71 %.
Interactive, novel reasoning in unseen mini-environments.
This counts runs in the ARC Prize's standard harness. The ARC Prize lists runs through a provider-specific adapter separately; they don't count here.
Research-level, unpublished mathematics.
Epoch released FrontierMath v2 on 12 Jun 2026 (cleaned items).
Thousands of expert questions at the edge of human knowledge.
Scores vary widely by eval setup, for example with or without web search and tools.
Fixing real software bugs in real GitHub projects.
PhD-level science questions, “Google-proof”.
Models now sit above human-expert level — the test is losing its discriminating power.
Tests meant to challenge for years are now exhausted in months. MMLU (2020) lasted about four years, GPQA (2023) only two. That's the real story — not a single score, but the pace.
Last checked: 30 September 2026
Benchmark scores are snapshots and vary by eval setup. They measure individual capabilities, not “intelligence” as a whole.
No. These tests measure narrow capabilities. A model can answer GPQA questions above human level and still fall short of humans on ARC-AGI-3. “State of AI” is a progress picture, not an AGI forecast.
Different leaderboards test under slightly different conditions (prompting, tool access, test version). We name the source and date for each value — and where there's a range, we say so.
An automated run checks the official sources every six hours; it reads ARC-AGI-3 straight from the ARC Prize's data. A new value only counts once two runs in a row report the same, or three for a jump of more than 15 percentage points. The run never accepts drops of more than 15 points on its own. The page shows a value we have checked by hand right away. GPQA Diamond, FrontierMath and SWE-bench Verified come from Epoch AI's data, the same that AI IQ shows. Each card states the date of its value.
We sort out what's actually relevant for your business — and implement the use cases that are worth it today.
These tools cover related ground.