Back to Blog
Original

US vs China Frontier Models Compared: GPT-6 Astra, Fable 5.1, GLM 5.3, Kimi K3, DeepSeek V4, Qwen 3.8 (September 2026)

Ten frontier AI models from the US and China compared: benchmarks, token prices and workload routing for GPT-6 Astra, Claude Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, GLM 5.3, Kimi K3, DeepSeek V4 and Qwen 3.8.

3 September 202618 min read
US vs China Frontier Models Compared: GPT-6 Astra, Fable 5.1, GLM 5.3, Kimi K3, DeepSeek V4, Qwen 3.8 (September 2026)

Last Updated: September 3, 2026

Four frontier AI models shipped within 48 hours. OpenAI released GPT-6 Astra and Anthropic released Claude Fable 5.1 on September 3, 2026, one day after Meta shipped Muse Spark 1.3 and Google shipped Gemini 3.8 Flash on September 2. On leaked launch benchmarks, GPT-6 Astra leads every comparable test against the best Anthropic models, scoring 98.6 percent on ARC-AGI-3 where Claude Opus 5 scores 30.2 percent. But those Astra numbers are not independently verified, and broad API access only opens in the coming days. Claude Fable 5.1 holds the strongest verified and shipping results, Gemini 3.8 Flash is roughly 13 times cheaper on input tokens, and Muse Spark 1.3 brings an open-weights roadmap. The China frontier changes the math again: Kimi K3 ranks number 4 on the Artificial Analysis Intelligence Index, the highest position ever recorded by an open-weights model, DeepSeek V4 Pro statistically ties Claude Opus 4.7 on SWE-bench Verified at roughly one thirty-fourth the input price, and GLM 5.3 Flash charges $0.071 per million input tokens, about 140 times cheaper than Astra. This guide compares ten models from both camps on benchmarks, price and real business workloads.

The September 2026 lineup at a glance

Four models from four labs landed in two days, and they aim at different jobs. GPT-6 Astra is the OpenAI flagship for complex reasoning, coding, computer use, research and document creation, with a 1,050,000 token context window. Claude Fable 5.1 is the most capable Anthropic model for coding and knowledge work, available on day one across AWS, Google Cloud, Microsoft Azure and the Anthropic API. Muse Spark 1.3 is the Meta frontier push into long-horizon personal agents, rolling out in Muse Code and the Meta Model API. Gemini 3.8 Flash is the Google cost-efficient tier for production agents, accepting text, image, audio and video input.

ModelLabReleasedInput / output price per 1M tokens
GPT-6 AstraOpenAISep 3, 2026$10.00 / $50.00
Claude Fable 5.1AnthropicSep 3, 2026$10.00 / $50.00 ($0.25 cache reads)
Muse Spark 1.3MetaSep 2, 2026$1.25 / $4.25
Gemini 3.8 FlashGoogle DeepMindSep 2, 2026$0.75 / $3.75

According to OpenAI model documentation, GPT-6 Astra carries an April 30, 2026 knowledge cutoff, supports reasoning effort from low up to max, and takes text and image input. According to the Google DeepMind model card, Gemini 3.8 Flash offers a 1 million token context, a 64,000 token output limit and a March 2026 knowledge cutoff.

Is GPT-6 Astra really the best model in the world?

On the numbers circulating with the launch, yes, and by a wide margin. A benchmark comparison table leaked alongside the rollout shows GPT-6 Astra ahead of Claude Fable 5.1, Claude Fable 5 and Claude Opus 5 on every comparable benchmark, according to the NextBigFuture compilation. The gaps are largest on abstract reasoning: 98.6 percent on ARC-AGI-3 versus 30.2 percent for Opus 5, and 97.6 percent on FrontierMath Tier 4 versus 87.8 percent for Fable 5.1. Two caveats matter for buyers. The table is unofficial and awaits independent verification from evaluators like LMArena and Artificial Analysis, and the model itself is only in enterprise trusted access this week, with API and ChatGPT plan access arriving in the coming days.

BenchmarkGPT-6 AstraClaude Fable 5.1Claude Opus 5
ARC-AGI-3 (abstract reasoning)98.6%not scored30.2%
FrontierMath Tier 4 v297.6%87.8%73.2%
GPQA Diamond (graduate science)96.0%93.7%93.2%
Terminal-Bench Science 0.164.6%52.6%29.0%
DeepSWE v1.1 (agentic coding)74.1%67.4%68.8%
AutomationBench (business workflows)41.4%31.4%26.9%

Source: leaked launch comparison tables compiled by NextBigFuture, September 3, 2026. Astra scores are not yet independently verified.

Specs tell a similar flagship story. The Astra context window is 1,050,000 tokens with 128,000 token outputs, and OpenAI supports prompt caching at $1 per million cached input tokens, batch and flex processing at 50 percent of standard rates, and tool use spanning web search, file search, code interpretation, hosted shell and computer use. Note the limits: no realtime audio, no fine-tuning at launch, and prompts above 272,000 tokens are billed at 2x input rates.

Frontier benchmark comparison chart September 2026: GPT-6 Astra vs Claude Fable 5.1 vs Claude Opus 5 across ARC-AGI-3, FrontierMath, GPQA Diamond, Terminal-Bench Science, AutomationBench and DeepSWE

How strong is Claude Fable 5.1 on verified benchmarks?

Very strong, and these numbers are official and reproducible. According to the Anthropic announcement, Claude Fable 5.1 scores 52.6 percent on Terminal-Bench Science 0.1, more than doubling the 24.7 percent from Fable 5 and beating the 29.0 percent from Opus 5 and the 22.4 percent from GPT-5.6 Sol. On Terminal-Bench 4.0 it reaches 55.8 percent, or 60.9 percent when running as Mythos 5.1, the same model with lighter safeguards. It posts 1853 on GDPval-AA v2 knowledge work, 77.9 percent partial accuracy on OSWorld 2.0 computer use, 60.9 percent on Humanity's Last Exam without tools, and 73.4 percent on CursorBench 3.2.0 agentic coding.

BenchmarkFable 5.1Opus 5GPT-5.6 Sol
Terminal-Bench Science 0.152.6%29.0%22.4%
Terminal-Bench 4.055.8% (Mythos: 60.9%)52.3%37.3%
GDPval-AA v2 (knowledge work)185318241711
OSWorld 2.0 (partial)77.9%75.4%not tested
Humanity's Last Exam (no tools)60.9%56.6%not tested
AutomationBench31.4%26.9%19.6%
CursorBench 3.2.073.4%70.0%67.2%

Source: the official Anthropic Fable 5.1 and Mythos 5.1 announcement, September 3, 2026, with production safeguards enabled.

The scientific results are the eye-openers. Anthropic reports that Mythos 5.1 designed protein binders with nearly 50 percent hit rate across 12 targets, where 10 to 15 percent is typical in the field, and 10 times higher binding affinity than the best designs submitted to Adaptyv Bio protein design competitions on three targets. Fable 5.1 also produced a new elevation map of a third of Venus at 2 to 3 kilometer resolution from 30 year old Magellan radar data, and Mythos 5.1 wrote GPU kernels that sped up seven open-source biology models by up to 2.5 times. At investment firm Millennium, Fable 5.1 found the root cause of a rare internal crash that engineers and every other model had failed to explain for years.

"In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition," says Craig Falls, Head of Quantitative Research at Jane Street Capital.

Pricing is where Fable 5.1 sharpens its edge: token prices stay at $10 per million input and $50 per million output, but cache reads drop 75 percent to $0.25 per million. Anthropic estimates typical workloads get about 25 percent cheaper and highly agentic workloads up to roughly 45 percent cheaper than Fable 5. One naming note: Claude Mythos 5.1 is the identical model with lighter safeguards for vetted cybersecurity and life science organizations, initially US-based, under the Anthropic Cyber Verification Program and Life Sciences Verification Program.

What is Muse Spark 1.3 and who is it for?

Muse Spark 1.3 is the Meta frontier-class model for long-horizon agentic work, and it is available now in Muse Code and the Meta Model API. According to the Meta research blog, it is trained to sustain longer missions: it juggles multiple workflows in one long thread, asks clarifying questions when prompts are ambiguous, confirms before consequential actions, and keeps better track of what it has already learned. Meta engineers measured it using roughly 20 percent fewer tool calls and roughly 25 percent fewer tokens than Muse Spark 1.2, with stronger resistance to prompt injection. Artificial Analysis titled its coverage "Muse Spark 1.3: Meta reaches the frontier." Two gaps at launch: Meta published no benchmark table, and a max reasoning mode is still finishing safety testing. Model trackers including Benchable, LMMarketCap and Kingy.ai list Meta Model API pricing at $1.25 per million input tokens and $4.25 per million output tokens, with cache reads around $0.15. The roadmap includes an open-weights release, which matters if you want frontier capability you can host yourself.

Is Gemini 3.8 Flash good enough for production agents?

For high-volume agent workloads, yes, and the price is the point. According to the Google DeepMind model card published September 2, 2026, Gemini 3.8 Flash builds on 3.7 Flash with advances in software engineering and agentic knowledge workflows, supports customizable effort levels for quality, cost and latency, and takes text, image, audio and video input with a 1 million token context and 64,000 token outputs. According to the Codersera pricing guide, it costs $0.75 per million input tokens and $3.75 per million output tokens, about 13 times cheaper on input than either flagship. The honest limits: it is a Flash-tier model, not the Google flagship, its knowledge cutoff is March 2026, and the model card notes hallucination risk and occasional slowness. Treat it as your volume lane, not your hardest-problems lane.

The China frontier: GLM 5.3, Qwen 3.8, Kimi K3 and DeepSeek V4

China's four frontier labs all shipped within ten weeks, and they changed two things at once: the capability gap and the price gap. Kimi K3 from Moonshot AI (July 16, 2026) is a 2.8 trillion parameter mixture-of-experts model with a 1 million token context, multimodal input and MIT-licensed open weights, hosted at $3 per million input and $15 per million output with cached input at $0.30. DeepSeek V4 offers a 1.6T/49B-active Pro (GA August 13) and a 284B/13B Flash, both MIT-licensed with 1M context. Qwen 3.8 Max is the Alibaba production flagship (August 3): 2.4T total parameters, 95B active, native image and video input, at $2/$6 per million. GLM 5.3 from Z.ai (August 18) targets long-horizon coding agents at $1.09/$3.43 with a 1.31 million token context. On flash variants: GLM 5.3 Flash runs $0.071/$0.238 per million as a native multimodal model, Qwen 3.8 Flash runs $0.15/$0.47 with open weights, and DeepSeek V4 Flash runs $0.22 to $0.44 input. Kimi has no K3 Flash yet.

ModelLabReleasedInput / output per 1M tokens
GLM 5.3Z.aiAug 18, 2026$1.09 / $3.43
GLM 5.3 FlashZ.aiAug 26, 2026$0.071 / $0.238
Qwen 3.8 MaxAlibabaAug 3, 2026$2.00 / $6.00
Qwen 3.8 FlashAlibabaopen weights$0.15 / $0.47
Kimi K3Moonshot AIJul 16, 2026$3.00 / $15.00
DeepSeek V4 ProDeepSeekAug 13, 2026 GA$0.66 to $1.32 / $1.98 to $3.96
DeepSeek V4 FlashDeepSeekJul 31, 2026$0.22 to $0.44 / $0.66 to $1.32

Source: OpenRouter and pricepertoken listings, Kingy.ai, and Codersera citing DeepSeek api-docs. DeepSeek ranges are off-peak to peak; off-peak applies to roughly 79 percent of the week, all weekend.

China benchmarks: what is verified and what is claimed?

The verified layer is stronger than most Western buyers expect. According to the Artificial Analysis Intelligence Index, Kimi K3 scores 57 and ranks number 4 of 189 models, the highest position an open-weights model has ever recorded, roughly tied with Claude Opus 4.8 at about 56 and behind only Claude Fable 5 at about 60 and GPT-5.6 Sol at about 59. GLM 5.2 scores 51 and DeepSeek V4 Flash 0731 scores 50, tied with Gemini 3.6 Flash. Qwen 3.8 Max is the outlier: Alibaba officially claims it trails only Claude Fable 5, but as of this writing no third-party evaluation exists. On coding, DeepSeek V4 Pro tops LiveCodeBench at 93.5, posts a Codeforces ELO of 3206 versus 3168 for GPT-5.5, and statistically ties Claude Opus 4.7 on SWE-bench Verified at 80.6 versus 80.8. The V4 Flash 0731 checkpoint scores 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE. Kimi K3 holds number 1 on the Arena frontend coding leaderboard at 1,679 Elo across 483,895 blind votes.

ModelAA Intelligence IndexStatus
Claude Fable 5~60 (#1)reference point
GPT-5.6 Sol~59 (#2)reference point
Kimi K357 (#4 of 189)verified, open-weights record
Claude Opus 4.8~56reference point
GLM 5.251verified
DeepSeek V4 Flash 073150verified
Qwen 3.8 Maxnot yet scoredvendor claims only

Source: Artificial Analysis index compilation, August 3, 2026. GPT-6 Astra and Claude Fable 5.1 postdate this table and await scoring.

US vs China price gap chart September 2026: flagship and flash tier input prices per million tokens for GPT-6 Astra, Fable 5.1, Kimi K3, Qwen 3.8 Max, GLM 5.3, DeepSeek V4 and flash variants, with Artificial Analysis index panel

One pricing mechanic is unique to DeepSeek: time-of-day billing. Off-peak rates are exactly half of peak, and off-peak covers every hour except 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, about 79 percent of the week including all weekends. Cache hits are nearly free: $0.022 per million on V4 Pro and $0.007 per million on V4 Flash. For batch workloads that can run at night, DeepSeek is effectively untouchable on price.

Price comparison: what a million tokens costs

List prices split the field into two camps. GPT-6 Astra and Claude Fable 5.1 both charge $10 per million input tokens and $50 per million output tokens, while Gemini 3.8 Flash charges $0.75 and $3.75. Cache economics decide the flagship contest: Astra cached input costs $1 per million with cache writes at $12.50, while Fable 5.1 cache reads cost $0.25. Because long agent runs re-read their own context constantly, that 75 percent cache discount is why Anthropic estimates up to 45 percent savings on agentic workloads. Muse Spark 1.3 sits between the two camps: trackers list $1.25 input and $4.25 output per million tokens with cache reads near $0.15, roughly 8 times cheaper on input than the flagships but above Gemini 3.8 Flash. One tracker, llm-stats.com, shows a "starts at" tier of $0.10 input and $0.20 output, which may reflect a batch or discounted tier and is not confirmed on Meta's own pages.

ModelInputCached inputOutput
GPT-6 Astra$10.00$1.00 (writes $12.50)$50.00
Claude Fable 5.1$10.00$0.25$50.00
Gemini 3.8 Flash$0.75not listed at launch$3.75
Muse Spark 1.3$1.25$0.15$4.25

Token price comparison September 2026: input, output and cache read prices per million tokens for GPT-6 Astra, Claude Fable 5.1 and Gemini 3.8 Flash

We see the same economics in our own agent stack at Flowtivity, where an OpenClaw-based agent runs daily research and outreach workflows. On long research sweeps, cache reads make up most of the token bill, so a $0.25 cache read price changes unit economics more than any headline benchmark. Routing high-volume extraction to a Flash-tier model and reserving flagship calls for synthesis cut our per-report model spend by roughly 80 percent.

Which model should your business actually use?

Match the model to the workload, not the leaderboard. Use GPT-6 Astra for your hardest reasoning and research once API access opens, especially work that needs its 1.05 million token context. Use Claude Fable 5.1 for agentic coding and knowledge work you need shipping today: it is on every major cloud, its verified numbers lead, and its cache pricing makes long runs affordable. Use Gemini 3.8 Flash for high-volume production agents where cost per task decides viability. Watch Muse Spark 1.3 if an open-weights frontier model fits your compliance or self-hosting plans. A practical starting mix for a growing business: Flash for volume, Fable 5.1 for coding and documents, and a waitlist spot for Astra. For cost-driven and self-hosted lanes, add the China frontier: DeepSeek V4 Flash or GLM 5.3 Flash for budget agents, Kimi K3 or DeepSeek V4 Pro for open-weights frontier work. The MIT weights on DeepSeek, Kimi and GLM also mean you can self-host entirely, which sidesteps data residency questions that some Australian and US buyers have with Chinese-hosted endpoints.

Decision guide: which AI model for which job, covering GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash and Muse Spark 1.3

Is there a Gemini Cyber model?

No, and it is worth killing this rumor early. There is no Gemini Cyber in Google DeepMind published model cards or Gemini API release notes, and searches for the name return nothing. The confusion likely comes from two real programs with similar names: Claude Mythos 5.1, the Anthropic lighter-safeguard tier for vetted cybersecurity and life science organizations, and the OpenAI enterprise Trusted Access for Cyber program attached to the GPT-6 Astra rollout. If you see "Gemini Cyber" in a comparison chart, treat the source as unreliable.

What to watch in the next two weeks

Three events will settle the real leaderboard. First, independent evaluators: once GPT-6 Astra reaches general API access, expect LMArena and Artificial Analysis to publish verified scores within days, and that will confirm or break the leaked tables. Second, Mythos access: Anthropic is coordinating with the US government to widen Claude Mythos 5.1 availability beyond initial US organizations, which matters for security teams. Third, Muse Spark open weights: an open-weights frontier release from Meta would reset build-versus-buy math for self-hosted AI. Fourth, Qwen 3.8 Max verification: it is the only model in this comparison with zero third-party scores, so its Artificial Analysis debut will move the China ranking. And if Moonshot ships a Kimi K3 Flash, the budget tier gets another serious entrant. Our working assumption until then: Astra probably leads on raw capability, Fable 5.1 is the best verified model you can deploy today, and Flash is the value pick that most production agents actually need.

Frequently asked questions

Is GPT-6 Astra available now?

Partially. OpenAI began the rollout on September 3, 2026 through its enterprise Trusted Access Program, with API access and Plus, Pro, Business and Enterprise plans following in the coming days.

Is GPT-6 Astra better than Claude Fable 5.1?

On leaked numbers, Astra leads every comparable benchmark. Those figures await independent verification, while Fable 5.1 results are official and the model ships today on all major clouds.

What is the cheapest capable AI model in September 2026?

The cheapest capable APIs are now Chinese: GLM 5.3 Flash at $0.071 per million input and $0.238 per million output, with DeepSeek V4 Flash at $0.22/$0.66 off-peak. The cheapest frontier-class model is DeepSeek V4 Pro at $0.66/$1.98 off-peak.

Are Chinese AI models as good as US models now?

On verified independent numbers, close. Kimi K3 ranks number 4 on the Artificial Analysis Intelligence Index at 57, roughly tied with Claude Opus 4.8. DeepSeek V4 Pro statistically ties Opus 4.7 on SWE-bench Verified. Qwen 3.8 Max claims to trail only Fable 5 but is not yet independently verified.

Which model is best for agentic coding today?

Claude Fable 5.1, with 55.8 percent on Terminal-Bench 4.0, 73.4 percent on CursorBench 3.2.0 and $0.25 cache reads that cut long agent run costs by up to about 45 percent.

Researched and written by AJ Awan, founder of Flowtivity. Former EY management consultant, TOGAF certified enterprise architect, 9+ years advising enterprises on technology strategy. Sources: OpenAI model documentation, the Anthropic Fable 5.1 announcement, the Google DeepMind Gemini 3.8 Flash model card, the Meta research blog, the NextBigFuture leaked benchmark compilation, the Codersera Gemini 3.8 Flash pricing guide, and Muse Spark 1.3 pricing from the Benchable, LMMarketCap and Kingy.ai trackers. China model data: OpenRouter, DeepSeek api-docs via Codersera, Kingy.ai, qwen.ai and the Artificial Analysis index compilation published on cnblogs. Leaked GPT-6 Astra scores are unofficial and will be updated when independent evaluations publish.

Want AI insights for your business?

Get a free AI readiness scan and discover automation opportunities specific to your business.