Artificial Analysis Rejigs Intelligence Index With v4.2, Fable 5.1 Tops, GPT-6 Astra Placed Second

There had been some eyebrows raised when GPT-6 Astra had not shown a jump over GPT 5.6 Sol on the Artificial Analysis Intelligence Index, but the company has gone ahead and announced a new version of the index just a day later.

Artificial Analysis has rolled out version 4.2 of its widely-followed Intelligence Index, describing the update as an interim step ahead of a bigger v5 release the team has been building towards for months. The benchmarking outfit says it is “accelerating elements” of that upcoming v5 launch because the pace of frontier model releases over the past few weeks made it necessary to update the index sooner rather than later.

With v4.2, Anthropic’s Claude Fable 5.1 (max with fallback) tops the leaderboard with a score of 57, followed closely by OpenAI’s GPT-6 Astra (max) at 55. Anthropic’s Claude Opus 5 (max) and Claude Fable 5 (with fallback) round out the top four, scoring 54 and 53 respectively, meaning Anthropic effectively occupies three of the top four spots on the new index.

Meta’s Muse Spark 1.3 (max) comes in at 53, tied with Claude Fable 5, followed by GPT-5.6 Sol (max) and Grok 4.6 (high) at 51 each, and Kimi K3 (max) at 50. Z AI’s GLM-5.3 (max), Google’s Gemini 3.8 Flash (high), and OpenAI’s GPT-5.6 Terra (max) also feature in the upper half of the table.

What changed in v4.2

Artificial Analysis says the index now leans more heavily on realistic, complex, real-world tasks and on private test sets designed to make the benchmark harder to game. The two big additions are:

  • AA-Briefcase: an in-house evaluation with a private, held-out test set that tests models on realistic agentic knowledge work spread across multi-week projects, each involving many linked tasks and thousands of input source files. It combines rubric-based and pairwise grading to judge verifiable task success, analytical quality, and presentation quality.
  • GDP.pdf: built by Surge AI, this evaluation checks single-turn professional document reasoning across 100 PDFs spanning ten domains, requiring models to pull together evidence scattered across 4,592 pages of text, tables, charts, footnotes and exclusions. Responses are checked against 1,275 expert-written criteria, and the headline “All-pass Rate” only credits a task when every single criterion is met.

GPQA Diamond has been dropped from the index after Artificial Analysis said the benchmark had become saturated, meaning most frontier models were already scoring close to the ceiling on it, making it less useful for telling top models apart.

The company has also doubled the proportion of the index made up of private, held-out data, from 20% in v4.1 to 40% in v4.2. This held-out share includes AA-Briefcase, AA-Omniscience, and the solution set for CritPt, and Artificial Analysis says it plans to push this figure higher still in v5, specifically to cut down on labs optimising their models to game public test sets rather than genuinely improving capability.

On the grading side, Artificial Analysis says it has added a grading system prompt to AA-LCR v1.1 and fixed errors in its answer keys to sharpen scoring accuracy. GDPval-AA v2 and AA-Briefcase have had their sampling improved and their Elo scale re-anchored for more stable ratings as new models get added over time, while SciCode’s grading sandboxes have been made more robust so that correct but slow code no longer gets marked as a failure.

Anthropic and OpenAI lead, but on cost efficiency it’s a four-way race

Beyond the top-line rankings, Artificial Analysis points out that Anthropic and OpenAI are the two clearly leading labs on the new index, with Meta placed third, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google.

Looking at the Intelligence Index against cost per task, Artificial Analysis notes that Anthropic, OpenAI, Meta, and Z AI are the four labs currently sitting on the updated Cost per Task Pareto frontier, meaning models from these four are the ones offering the best trade-off between capability and cost at various price points. Claude Fable 5.1 sits at the top-right end of that frontier — highest capability, but also the most expensive per task among the group — while Meta’s GLM-5.3-Flash and OpenAI’s GPT-5.6 Luna variants anchor the cheaper end.

On token efficiency, GPT-6 Astra stands out as being more efficient than almost every other model near the intelligence frontier, meaning it needs fewer output tokens to reach its score. Artificial Analysis says Claude Fable 5.1 and Gemini 3.8 Flash are, by contrast, the most token-hungry among models that score at least 25 on the index.

Latest Rejig

This isn’t the first time Artificial Analysis has rejigged its flagship index. In June, Artificial Analysis had similarly released the v.1 version of the index which had caused a bit of a shuffle in the pecking order. With AI progressing rapidly, it’s perhaps expected that the indexes used to measure its progress will also evolve every few months.

Posted in AI