Back to Blog
Original

Terminal-Bench 4.0 Exposes Inflated AI Scores: Gemini 3.8 Flash Falls From 87.6% to 19.7%

Terminal-Bench 4.0 cut Gemini 3.8 Flash from 87.6% to 19.7% and Muse Spark 1.3 to 33.3% while GPT-6 Astra (59.6%) and Claude Fable 5.1 (55.1%) held the top. Inside the benchmaxxing debate and what it means for choosing AI models.

10 September 202611 min read
Terminal-Bench 4.0 Exposes Inflated AI Scores: Gemini 3.8 Flash Falls From 87.6% to 19.7%

Last Updated: September 10, 2026

Terminal-Bench 4.0 Exposes Inflated AI Scores: Gemini 3.8 Flash Falls From 87.6% to 19.7%

When Artificial Analysis upgraded its Intelligence Index to v4.3 on September 7, 2026 and replaced Terminal-Bench 2.1 with the much harder Terminal-Bench 4.0, several flagship model scores collapsed. According to Artificial Analysis's independent harness, Google's Gemini 3.8 Flash fell from 87.6% on Terminal-Bench 2.1 to 19.7% on Terminal-Bench 4.0, a 68-point drop. Meta's Muse Spark 1.3 fell from 84.3% to 33.3%. Only OpenAI's GPT-6 Astra, at 59.6%, and Anthropic's Claude Fable 5.1, at 55.1%, stayed near the top of the new board. The research firm SemiAnalysis calls the pattern benchmaxxing: a model that excels on one benchmark while being nowhere near as strong at general work. Meta's Chief AI Officer Alexandr Wang pushed back, arguing that a big drop can simply mean the new test is harder. Both things can be true, and the difference matters for anyone buying AI in 2026.

What happened when Artificial Analysis switched to Terminal-Bench 4.0?

Four frontier models launched within four days of each other: Claude Fable 5.1 on September 1, Gemini 3.8 Flash and Muse Spark 1.3 on September 2, and GPT-6 Astra on September 3. All four published launch tables built around the older Terminal-Bench revision, where several scored in the 80s and 90s. Then Artificial Analysis moved its index to Terminal-Bench 4.0 and re-ranked everyone under the harder test, and the leaderboard reshuffled violently.

Model (lab)Terminal-Bench 2.1Terminal-Bench 4.0Change
Claude Fable 5.1 (Anthropic)91.4%55.1%-36.3 pts
Gemini 3.8 Flash (Google)87.6%19.7%-67.9 pts
Muse Spark 1.3 (Meta)84.3%33.3%-51.0 pts
GPT-6 Astra (OpenAI)Not separately published59.6%New leader

Scores shown are the best Artificial Analysis harness result per model on each revision, using each model's highest reasoning effort. Claude Fable 5.1 led 2.1 at 91.4% (max effort with fallback) and scores 55.1% on 4.0 (xhigh effort). GPT-6 Astra, released after the benchmark work shifted to 4.0, tops the new revision at 59.6% (xhigh) with 59.1% at max effort, and per Artificial Analysis it beats Claude Fable 5.1 (52% at max) and GPT-5.6 Sol (40%) on the harder test.

The official Terminal-Bench leaderboard tells the same story under a different harness: GPT-6 Astra running in the Codex agent resolves 58.18% of trials, Claude Fable 5 in Claude Code resolves 44.55% at a reported $7.3k compute cost for the full run, and Gemini 3.8 Flash manages 19.09%. Four independent measurements, one consistent conclusion: the top of the old board did not survive contact with the new one, except for Astra and Fable 5.1.

Why is Terminal-Bench 4.0 so much harder than 2.1?

Terminal-Bench measures whether an AI agent can complete real, long-horizon work in a command-line environment: debugging software, running data pipelines, configuring systems, and passing automated verification. Version 4.0, released August 29, 2026 by the Harbor framework team behind the benchmark, is a deliberate hardening pass, not a cosmetic update.

According to the Terminal-Bench team's release notes, 4.0 is a 66-task suite spanning software, machine learning, science, operations, security, hardware, and media. The team removed 8 tasks: 2 because every latest-generation model solved them 5 out of 5 times (saturated), 2 over refusal issues, 2 because solutions were publicly visible, and 2 for unresolved quality problems. They fixed 19 task definitions with better instructions, environments, and verifiers. They also calibrated CPU, memory, and time allowances using the methodology from Anthropic's infrastructure-noise research, and set a flat 8-hour agent timeout so slow models are no longer rescued or punished by arbitrary per-task limits. Every trial now measures the same thing more precisely, and the ceiling got lower for everyone.

What is benchmaxxing, and why does SemiAnalysis keep using the word?

Benchmaxxing is what happens when a model posts an eyewatering number on one specific benchmark while being nowhere near as strong at general work. The term, popularized by the research firm SemiAnalysis and now standard shorthand in the evaluation community, describes optimization aimed at the test rather than the capability: tuned harnesses, cherry-picked effort levels, stitched-together comparison tables, and quiet reliance on an aging benchmark revision that the model happens to match well.

The launch week of September 2026 was a case study. The four vendors published DeepSWE numbers of 67.4% for Fable 5.1, 73.7% for Gemini 3.8 Flash, 75.4% for Muse Spark 1.3, and 74.1% for GPT-6 Astra, but as The Kaitchup documented, several competitor values in those tables were reused from other vendors' announcements and leaderboards rather than re-run under identical conditions. "An agentic benchmark score is a property of an entire evaluation system, not of the model alone," says Benjamin Marie of The Kaitchup. On the standardized DataCurve DeepSWE leaderboard, which freezes one harness for all submissions, Astra, Gemini 3.8 Flash, Opus 5, and GPT-5.6 Sol all cluster at roughly 73 to 74% with overlapping confidence intervals, and neither Fable 5.1 nor Muse Spark 1.3 has a public standardized submission yet.

Meta sits at the sharp end of this story. Muse Spark 1.3's launch coverage leaned hard on its Terminal-Bench 2.1 gain, from 80% to 85% and 86%, and Artificial Analysis initially scored its max variant at 62 on the Intelligence Index, second only to Anthropic's models. On the harder 4.0 revision the same model scores 33.3%. TechRepublic's independent writeup was already titled "independent tests are more mixed," and VentureBeat noted Meta's best results came from a limited-preview variant most developers cannot yet use.

Does the score crash prove the old scores were fake? Alexandr Wang says maybe not

Alexandr Wang, Chief AI Officer at Meta, publicly pushed back on the benchmaxxing framing around Terminal-Bench 4.0, arguing that a large score drop can simply mean the new benchmark is harder. That objection is not spin, and it deserves a fair hearing before you screenshot anyone's collapse.

Here is the steelman. Every model dropped on 4.0, including honest ones: Fable 5.1 fell 36 points and still leads the pack behind Astra. When a test removes saturated tasks, tightens verifiers, and recalibrates resources, absolute scores are supposed to fall, and a fall from 87.6% to 19.7% can partly reflect a harder yardstick rather than a worse model.

But the reordering is the tell. On Terminal-Bench 2.1, Gemini 3.8 Flash sat 3.8 points behind the best model, close enough to market as near-frontier. On 4.0 it sits 39.9 points behind the leader. A harder benchmark compresses honest models and exposes narrow ones, because harder, messier tasks have nowhere to hide harness polish. It is also relevant that Artificial Analysis specifically attributed Gemini 3.8 Flash's launch-week gains to agentic evaluations including Terminal-Bench 2.1, the exact test that was about to be retired. The drop does not prove deception by anyone. It proves the old number was not measuring general agentic strength.

How much can a harness or setup move a benchmark score?

More than most buyers assume, and this is the quiet part of the whole story. Published research puts the harness effect between 6 and 24 points on agentic benchmarks, which is larger than the gap between most competing frontier models.

  • The SWE-agent researchers improved the same GPT-4 Turbo model by 10.7 percentage points on SWE-bench Lite purely by changing the agent-computer interface, per their NeurIPS paper.
  • Harness-Bench, a 2026 systematic study, found a 23.8-point aggregate gap between harnesses over a shared pool of models and tasks.
  • Anthropic's own engineering team documented roughly 6 points of Terminal-Bench movement from infrastructure and resource-enforcement differences alone, which is exactly the noise the 4.0 calibration targets.

Against that backdrop, a launch table that mixes a freshly tuned first-party run with competitor numbers scraped from old leaderboards is marketing, not measurement. According to The Kaitchup's analysis, GPT-6 Astra's launch table reused Anthropic's entire Terminal-Bench comparator block and added Google's 19.1 for Gemini, meaning some "head-to-head" rows in vendor tables were never run head-to-head at all.

How should teams actually pick AI models when leaderboards disagree?

Treat every leaderboard as one data point in a portfolio of evidence, and never let a single benchmark make a procurement decision. The pattern that burned buyers this month was trusting one number on one revision of one test, harvested under one vendor's harness.

At Flowtivity we run a fixed evaluation protocol before recommending a model for client workflow automation: five representative tasks from the client's actual environment, three attempts each, identical prompts and tool access for every contender, and a frozen harness so nothing changes mid-bake-off. We score pass rate, cost per completed task, and wall-clock time, because a model that passes 90% of tasks at five times the token cost is a different business decision, not a different benchmark score. The Terminal-Bench data makes the case for us: on the official board, Astra's run cost about $3.3k in compute while Claude Fable 5's cost about $7.3k, and Gemini 3.8 Flash burned 17.2 billion tokens to finish near the bottom. Capability per dollar under your real conditions is the only leaderboard that pays invoices.

The practical checklist: triangulate across at least two independent leaderboards using the same benchmark revision, ignore any comparison table whose competitor numbers were imported from other vendors, weight tests with private task sets and versioned harnesses, and run the small pilot before the contract. One leaderboard told the market Gemini 3.8 Flash was a 89% terminal agent. A harder one said 19.7%. Your own tasks are the tiebreaker that matters.

The bottom line: Terminal-Bench 4.0 did not make the models worse. It made the measurement better, and better measurement is brutal for models whose scores were built on the old test's quirks. When GPT-6 Astra and Claude Fable 5.1 hold their lead on a harder, recalibrated benchmark while others fall 50 to 68 points, that is the system working. Trust leaderboards that change, version, and publish their noise, and distrust any vendor table that has never lost a comparison.

Want AI insights for your business?

Get a free AI readiness scan and discover automation opportunities specific to your business.