The two rejigs in the Artificial Analysis Intelligence Index in two days have led to some interesting developments.
Research firm Semi Analysis has levelled one of the sharper accusations of the current AI cycle at Google and Meta, calling Gemini 3.8 Flash and Muse Spark 1.3 among the most “benchmaxxed” frontier models yet seen — systems that post competitive numbers on older public benchmarks but fall apart when the evaluation is even modestly refreshed.

The evidence, according to the firm, sits in the gap between Terminal Bench 2.1 and Terminal Bench 4.0. On the older benchmark, Gemini 3.8 Flash scores 89.4%, roughly on par with GPT-6 Astra at 88.4% and Anthropic’s Fable 5.1 at 91.4%. Muse Spark 1.3 lands at 88.8%, essentially tied with the field. Shift the test to Terminal Bench 4.0, however, and the picture inverts: Astra falls to 57.7% and Fable 5.1 to 55.8% — material drops of roughly 30-35 percentage points — while Gemini 3.8 Flash collapses to 19.1% and Muse Spark 1.3 to 33.3%. The Google model’s 70-point cliff is more than double the degradation seen at OpenAI or Anthropic.
Semi Analysis argues the mechanism is straightforward. Terminal Bench 2.1 is fully public, and while Google and Meta would not train on the tasks directly, they can buy data from reinforcement-learning environment startups designed to mimic those tasks as closely as possible. The net effect, the firm contends, is functionally identical to training on the benchmark itself. The problem is that the performance does not generalise: if the model had genuinely learned agentic reasoning, it should hold up on Terminal Bench 4.0. That it does not suggests the capability was narrower than the headline number implied.
The pattern allegedly repeats on DeepSWE, where Gemini 3.8 Flash sits in the “most efficient” corner of the cost-performance frontier. Semi Analysis hints that data-curve vendors may have shaped training tasks to mirror DeepSWE’s format, letting Google hit an attractive price-to-score ratio without building the underlying generalisation.
Meta’s response came directly from Alexandr Wang, the Scale AI founder now leading Meta Superintelligence Labs. In a public rebuttal, Wang called the argument “silly,” pointing out that GPT-5.6 Sol scored 88.8% on Terminal Bench 2.1 and just 37.3% on Terminal Bench 4.0 — a larger absolute drop than GPT-6 Astra’s — without attracting the same “benchmaxxed” label. The difference, Wang said, is that Terminal Bench 4.0 is simply a much harder, unsaturated evaluation, whereas Terminal Bench 2.1 is approaching saturation. He added that Meta does not claim Muse Spark 1.3 matches Astra or Fable 5.1 on raw capability, but rather that it is “significantly more cost-effective,” and that future Muse models will compete more directly with the top tier.
The exchange highlights a tension that has become central to the AI industry: public benchmarks are the primary language for comparing models, yet the moment they become public they become vulnerable to exactly this kind of optimisation. Semi Analysis notes that Terminal Bench 4.0 is useful signal today only because it was released two weeks ago; given that its tasks are similarly public, it expects the same labs to hill-climb it in short order. The firm’s proposed solution is more high-quality private benchmarks — evaluations whose tasks are not visible to model builders and therefore cannot be mimicked by data vendors.
Whether Muse Spark 1.3 and Gemini 3.8 Flash represent genuine cost-efficient engineering or narrower benchmark optimisation may not be fully settled until those private evaluations arrive. For now, the Terminal Bench 4.0 numbers have given the industry’s benchmark sceptics their most concrete talking point yet.