Meta Releases Muse Spark 1.2, Jumps To Score Of 54 On Artificial Analysis Intelligence Index

After being in the wilderness for over a year, Meta seems to now be quickly iterating on its AI models.

The company has released Muse Spark 1.2, its third model launch in four months, and the new model has scored 54 on the Artificial Analysis Intelligence Index, a benchmark that aggregates results across nine different evaluations. The score puts Muse Spark 1.2 in a tie for third place among US labs, alongside SpaceXAI’s Grok 4.5, and just behind the current frontier of Claude Opus 5, Claude Fable 5, GPT-5.6 Sol and Kimi K3.

The pace of releases is notable given where Meta’s AI division was standing a little over a year ago. The company’s last major model before this current stretch was Llama 4, a launch that turned messy after it emerged that the benchmark numbers put out at the time had been assembled using different versions of the model for different tests. Former Meta Chief Scientist Yann LeCun later said the team had fudged the numbers to look better than they were, a revelation that reportedly upset CEO Mark Zuckerberg enough to sideline the entire GenAI organization. Meta went on to acquire a large stake in Scale AI, brought in its CEO Alexandr Wang to run what is now called Meta Superintelligence Labs, and laid off hundreds of engineers in the process.

That rebuild is now producing results that are being verified by outside benchmarking rather than numbers Meta is simply putting out about itself. Muse Spark 1.2 arrives roughly a month after Muse Spark 1.1, which itself had scored 51 on the index, and follows Muse Spark 1.0’s debut score of 43 back in April. The jump from 43 to 54 in under four months is one of the sharper improvement curves any lab has shown this year.

Meta Muse Spark 1.2 Benchmarks

The 3-point gain over Muse Spark 1.1 is concentrated almost entirely in agentic capability, an area Artificial Analysis had flagged as the model’s clearest weakness at the time of the 1.1 launch. On GDPval-AA v2, the benchmark measuring agentic performance on real-world knowledge work, Muse Spark 1.2’s Elo rating jumped 260 points, from 1371 to 1631. That puts it fifth among every model Artificial Analysis has tested, ahead of Claude Opus 4.8 (max) at 1588, though still trailing Claude Opus 5 (max) at 1852, GPT-5.6 Sol (max) at 1730, and Kimi K3 at 1685.

Terminal-Bench 2.1 rose two points to 80 percent, and Tau3-Bench Banking gained two points as well, moving to 27 percent. Scientific reasoning was more of a mixed bag. CritPt improved three points to 18 percent, but SciCode dropped two points to 56 percent and Humanity’s Last Exam slipped a point to 44 percent, suggesting the gains this cycle came from deliberate investment in agentic workflows rather than a broad uplift across every category.

One of the more unusual shifts is in AA-Omniscience, Artificial Analysis’s hallucination benchmark. The score there rose from 18 to 22, but the reason behind it is abstention rather than accuracy. The model’s hallucination rate fell 10 points, from 38 to 28 percent, but that’s largely because Muse Spark 1.2 is now declining to answer questions far more often, with its attempt rate dropping from 82 to 67 percent. Accuracy on questions it does attempt actually fell slightly, from 41 to 38 percent. The model has effectively been tuned to say “I don’t know” more frequently, which lowers hallucinations on paper while also lowering its raw hit rate.

The context window remains unchanged from Muse Spark 1.1 at 1 million tokens, and the model is available at launch through Meta’s first-party API.

Meta Muse Spark 1.2 Pricing

Pricing is unchanged from the previous release, at $1.25 per million input tokens and $4.25 per million output tokens, with cached input discounted down to $0.15 per million tokens. Despite holding pricing flat, the cost per Intelligence Index task has gone up, from $0.29 with Muse Spark 1.1 to $0.40 with Muse Spark 1.2. Artificial Analysis attributes this entirely to the model using more tokens to complete each task rather than any pricing change, with input token usage up roughly 53 percent and output usage up around 36 percent, concentrated heavily in the GDPval-AA v2 evaluation.

Even with that increase, Muse Spark 1.2 remains one of the more cost-efficient options at its intelligence tier. Only Grok 4.5 (high) at $0.37 per task and GPT-5.6 Sol (medium) at $0.39 come in cheaper among models clustered near its score. GPT-5.6 Terra (max) costs $0.51 per task, Kimi K3 (max) costs $0.86, and GPT-5.5 (xhigh) runs $1.18, all for roughly comparable intelligence.

Zuckerberg has taken to personally announcing these launches, a habit that started with Muse Spark 1.1’s release in July, when he posted on X for the first time in three years. With Muse Spark 1.2 landing so soon after and closing much of the agentic gap that had held the previous version back, Meta looks intent on keeping up a release cadence that few labs have matched this year.

Posted in AI