Google’s Gemini 3.7 Flash Tops Artificial Analysis’ Analyst Agent Benchmark, Beats Opus 5, GPT 5.6 Sol

Google’s Gemini 3.7 Flash Model isn’t quite at the frontier, but it’s beating the best in the world at some specific use-cases.

Artificial Analysis has published results for AA-AnalystAgent, its benchmark built specifically around the kind of spreadsheet and document work that business and data analysts do every single day, and Gemini 3.7 Flash has come out on top of the entire field. Running at high reasoning effort, the model scored 60% on the benchmark’s headline pass^5 metric, ahead of every other model Artificial Analysis tested, including systems that are generally considered more capable overall.

Google made sure the win got attention. The company posted about the result directly, noting that Gemini 3.7 Flash combined reasoning and speed to deliver the highest accuracy in the test while completing tasks 60% to 90% faster than the other top-performing models, and more than twice as fast, 2.4x to be specific, as its closest rival on accuracy.

What AA-AnalystAgent Actually Tests

AA-AnalystAgent is built around 80 questions spread across 14 business and scientific domains, and each question is tied to a folder of real source spreadsheets and documents rather than a clean, pre-formatted dataset. The idea, according to Artificial Analysis, is that arithmetic alone doesn’t make a good analyst. Interpreting which source to trust, deciding which caveats and exceptions apply, and settling on a defensible methodology matter just as much as getting the final number right, and that’s the kind of judgment the benchmark is trying to isolate.

The scoring method is the more interesting part. Instead of running each question once, Artificial Analysis runs every question five separate times per model and reports pass^5, the share of questions a model gets right on all five attempts. That bar is deliberately harsh. A model can solve a large chunk of questions once and still score poorly on pass^5 if its answers don’t hold up on a second or third try, which is closer to how analysts actually get evaluated at work. Nobody trusts a number that only checks out sometimes.

Where Gemini 3.7 Flash Landed Against The Field

Gemini 3.7 Flash’s 60% score put daylight between it and the rest of the leaderboard. Anthropic’s Claude Opus 5, running at max reasoning effort, came in second at 53.8%, making it the closest accuracy rival Google was likely referencing in its own framing of the results. OpenAI’s GPT-5.5 at xhigh effort followed at 50%, with Claude Fable 5, run with a fallback configuration, close behind at 48.8%.

GPT-5.6 Sol, OpenAI’s newer flagship, scored 47.5% at max effort, landing just below Fable 5 despite generally trading blows with it on other Artificial Analysis leaderboards, including the Intelligence Index, where the two models have been separated by a single point for weeks. xAI’s Grok 4.6 came in at 41.3%, and Moonshot’s Kimi K3, an open-weights model that has been closing the gap with closed labs on several other benchmarks this year, scored 38.8%.

Further down, Thinking Machines’ Inkling scored 23.8%, Claude 4.5 Haiku managed 15%, and Mistral Medium 3.5 came in at 12.5%. MiniMax-M3 and Nvidia’s Nemotron 3 Ultra rounded out the chart at 10% and 6.3% respectively, a reminder of just how steep the drop-off is once you move away from frontier-tier reasoning models on a task this demanding.

A Flash Model Beating Its Own Family’s Flagships

The result is notable partly because of what Gemini 3.7 Flash is supposed to be. Flash models sit below Google’s Pro-tier releases in the naming hierarchy, positioned as the faster, cheaper option rather than the frontier system. Gemini 3.7 Flash arrived just weeks after Gemini 3.6 Flash, and Artificial Analysis had already flagged it as sitting on the Pareto frontier of intelligence versus speed, meaning no other model available at the time offered a better combination of the two.

That earlier result showed the model scoring 56 on the general Artificial Analysis Intelligence Index, trailing GPT-5.6 Terra and Meta’s Muse Spark 1.2, both at 57, and sitting just ahead of Claude Sonnet 5 at 55. On raw general intelligence, in other words, Gemini 3.7 Flash isn’t the leader. On a benchmark built specifically around the messy, judgment-heavy work of pulling a reliable number out of real business documents, it is.

That distinction is important for how the AI labs are now competing with each other. General intelligence indices reward breadth. Task-specific agent benchmarks like AA-AnalystAgent reward a narrower kind of consistency, the ability to reach the same correct answer five times running on a task with genuine ambiguity in it, and that appears to be where Gemini 3.7 Flash’s architecture and tuning are currently paying off hardest.

The Wider Field Keeps Getting Tighter

The benchmark also lands at a moment when the gap between frontier labs has been shrinking across nearly every metric Artificial Analysis tracks. Anthropic’s Claude Fable 5 briefly held the top spot on the general Intelligence Index before US export control rules forced Anthropic to pull the model offline for foreign users, and even after access was restored, Kimi K3 has been closing in on both Fable 5 and GPT-5.6 Sol on that same index, at one point landing within two points of the top. Anthropic’s own Claude Opus 5 launch was explicitly pitched as a model that gets close to Fable 5’s capability while costing half as much to run.

Against that backdrop, a Flash-tier Gemini model topping a benchmark this specific and this demanding is a real result for Google, even if it doesn’t reset the broader hierarchy. Enterprise buyers evaluating AI for data and business analysis work aren’t shopping for the model that tops the most leaderboards. They’re shopping for the one that gets the number right, consistently, on the kind of spreadsheet their own analysts deal with every week. On that particular question, for now at least, Gemini 3.7 Flash has the strongest published claim.

Artificial Analysis notes that the question set and reference answers behind AA-AnalystAgent are being kept private, aside from a handful of published examples, specifically to limit contamination risk as labs train future models. The benchmark is being reported as a standalone leaderboard rather than folded into the company’s broader Intelligence Index, which suggests Artificial Analysis is treating analyst-style agentic work as its own category worth tracking separately going forward, rather than as one input among many.

Posted in AI