Last Updated: August 15, 2026
China's two flagship open-weights labs shipped competing models within 24 hours of each other: DeepSeek released V4-Pro-0813 on August 13, 2026, and Zhipu released GLM-5.3 on August 14, 2026. Both claim the open-weights coding crown using entirely different philosophies. DeepSeek built a new 1.6-trillion-parameter Mixture-of-Experts architecture and publishes a 96.40% SWE-bench Verified score, the highest open-weight result on the board. Zhipu kept the same 743-billion-parameter base as GLM-5.2 and extracted all of its gains from post-training, claiming a 50% coding improvement plus emergent cybersecurity skills that found 2,436 real vulnerabilities across 269 open-source projects. We run GLM-5.2 daily as the engine behind our own operations agent and have benchmarked DeepSeek's V4 family locally, so this comparison breaks down what each lab actually measured, what the numbers hide, and which model fits which team, with the Australian cost angle included.
Key Takeaways
- 24-hour showdown: V4-Pro-0813 went GA on August 13, GLM-5.3 launched August 14, 2026. Both claim the open-weights coding lead.
- Different philosophies: V4-Pro is a new 1.6T-parameter MoE architecture with 49B active per token. GLM-5.3 is the same 743B base as GLM-5.2 with extended post-training on real engineering workflows.
- Shared benchmarks: On self-reported runs, GLM-5.3 edges CyberGym (84.5% vs 83.3%) and AutomationBench (48.2 vs 31.8). V4-Pro owns SWE-bench Verified at 96.40%.
- Not comparable numbers: GLM-5.3 reports Terminal-Bench 3.0 (28.3) while DeepSeek reports Terminal-Bench 2.1 (87.9). Different benchmark versions, no cross-lab verification yet.
- Price split: GLM-5.3 sells as a subscription from US$18/month. V4-Pro sells per-token at US$0.435/$0.87 per million, rising to peak rates of US$1.32/$3.96 from August 16.
- Openness: V4-Pro weights are MIT-licensed and live now. GLM-5.3 weights arrive in roughly two weeks after security review.
What Is DeepSeek V4-Pro-0813?
V4-Pro-0813 is the general-availability build of DeepSeek's flagship, released August 13, 2026 alongside the DeepSeek Harness. It is a 1.6-trillion-parameter Mixture-of-Experts model with about 49 billion parameters active per token, trained on more than 32 trillion tokens using hybrid CSA+HCA attention, manifold-constrained hyper-connections and the Muon optimizer. It accepts roughly one million tokens of input and can output up to 384,000, and its weights are open under the MIT license for commercial use and self-hosting. According to DeepSeek's API site, the official release features "significantly enhanced agent capabilities" with native OpenAI Responses API support and Codex integration, and the company reports a 500-request concurrency limit, positioning it for intensive reasoning workloads.
What Is GLM-5.3?
GLM-5.3 is Zhipu's flagship, released August 14, 2026 through the GLM Coding Plan and ZCode. The headline: it shares its base model with GLM-5.2, and every gain comes from post-training alone. According to The Decoder, Zhipu says GLM-5.3 is the most powerful open-weights coding model, with the biggest jumps in agent-based tasks. On Z.ai's internal Code Bench it improves 50% over GLM-5.2, and Zhipu claims programming and agent capabilities on par with Claude Fable 5. Training tasks were expanded from isolated coding problems to full engineering workflows, with some tasks equivalent to several days of senior-engineer work using real compute clusters, storage and internal codebases. Z.ai's documentation lists a one-million-token context with 128,000-token maximum output.
How Do the Benchmarks Stack Up?
This is where honest analysis matters, because almost every number floating around is self-reported by the lab that owns the model. The two companies also report different benchmark versions: DeepSeek publishes Terminal-Bench 2.1 while Zhipu publishes Terminal-Bench 3.0, which is a harder suite, so 87.9 and 28.3 are not comparable scores. Only two benchmarks appear in both companies' materials: CyberGym, where GLM-5.3 scores 84.5% against V4-Pro's 83.3%, and AutomationBench, where GLM-5.3 scores 48.2 against V4-Pro's 31.8. Independent verification has not landed yet. The table below separates each lab's own claims.
| Benchmark | GLM-5.3 (Z.ai-run) | V4-Pro-0813 (DeepSeek-run) |
|---|---|---|
| SWE-bench Verified | Not published | 96.40% (#1 open-weight, #2 overall) |
| SWE-bench Pro | Not published (5.2 scored 62.1) | 55.4% |
| Terminal-Bench (different versions) | 28.3 on TB 3.0, up from 4.6 | 87.9% on TB 2.1, up from 72.1 |
| DeepSWE v1.1 | 66.9, up from 46.2 | Not published in GA table |
| CyberGym (shared) | 84.5%, self-reported SOTA | 83.3% |
| ExploitBench | 54.4%, up from 24.4% | Not published |
| AutomationBench (shared) | 48.2, up from 26.2 | 31.8 |
| Toolathlon-Verified | Not published | 74.1% |
| Agents' Last Exam (CLI) | 28.5, up from 23.8 | Not published |
| Humanity's Last Exam | Not published | 60.0% with tools, 42.7% without |
| DSBench FullStack / Hard | Not published | 71.1% / 67.2% |
| GDPval-AA v2 (44 occupations) | 1769 points | Not published |
Two credibility notes cut both ways. According to SCMP's analysis, V4-Pro-0813 trails OpenAI's Terra and Moonshot's Kimi K3 on the independent Artificial Analysis Intelligence Index (53.0), even while leading CyberGym in its own table. On the GLM side, The Decoder notes some GLM-5.3 figures are Z.ai-run evaluations with independent leaderboard entries still pending. In short: believe the direction, not the decimals, until third-party boards update.
Why Is Everyone Talking About GLM-5.3's Cybersecurity Skills?
The security story is GLM-5.3's genuine differentiator. Zhipu trained the model with data and environments built to find software vulnerabilities, and according to the Z.ai blog, the model "began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains." Working with security teams in China, it found 2,436 vulnerabilities across 269 open-source projects, some up to 40 years old, all documented in a public registry at cvd.z.ai. VentureBeat reports it already surfaced a serious vulnerability in Cursor. The nuance: on ExploitBench, which measures deeper exploitation, GLM-5.3's 54.4% trails Anthropic's Mythos at 78% and OpenAI's GPT-5.6 Sol at 76.5%, so its edge sits at the discovery and validation end of the chain. That matters for defensive teams doing code review and audit work more than for red-team operators.
How Do the Specs and Pricing Compare?
The commercial split is the widest gap between the two. V4-Pro sells per-token with full transparency: US$0.435 input and US$0.87 output per million tokens, cache hits at US$0.003625. From August 16, 2026 at 16:00 UTC, DeepSeek switches to peak and off-peak rates, with peak at US$1.32/$3.96 and off-peak at US$0.66/$1.98 per million tokens. GLM-5.3 sells as a subscription: the GLM Coding Plan costs US$18 (Lite), US$80 (Pro) or US$168 (Max) per month, dropping to US$12.60, US$56 and US$117.60 on annual billing, with a per-token API listed as coming soon. A third-party preview endpoint prices it at US$0.06 input and US$0.22 output per million tokens, but that is an evaluation channel, not a production commitment.
| Dimension | GLM-5.3 | V4-Pro-0813 |
|---|---|---|
| Released | August 14, 2026 | August 13, 2026 |
| Parameters | 743B base (same as GLM-5.2) | 1.6T total, 49B active MoE |
| Context / max output | 1M / 128K | 1M / 384K |
| Weights | Open in ~2 weeks after security review | MIT, live now |
| API pricing | Subscription $18-168/mo, API coming soon | $0.435/$0.87 per M now, peak $1.32/$3.96 from Aug 16 |
| Harness fit | ZCode, Claude Code, OpenCode | DeepSeek Harness (dsh), Codex, Responses API |
| Standout strength | Vulnerability discovery (CyberGym 84.5%) | SWE-bench Verified 96.40% |
| Reasoning control | Multiple thinking modes | Non-think / Think / Think Max |
| Concurrency | Not published | 500 requests |
What Does This Mean in Australia?
Two practical angles for Australian teams. First, DeepSeek's new peak windows, 01:00-04:00 UTC and 06:00-10:00 UTC, convert to 11am-2pm and 4pm-8pm AEST, which is the core of the local working day. Any AU team pointing agents at V4-Pro during business hours should budget at peak rates, which are triple the current input price, and shift batch workloads to mornings or overnight. Second, GLM's flat subscription sidesteps token maths entirely: for a solo developer or a small dev team, US$18-80 per month for frontier-class coding is now cheaper than a single consulting hour. That price collapse is the real story of this 24-hour showdown for growing businesses: frontier coding capability has moved from enterprise procurement to a line item smaller than a Netflix family plan.
Our Take After Running Both Families
We have direct skin in this game. Our operations agent Flowbee runs on GLM-5.2, the exact base model GLM-5.3 post-trains on, so a claimed 50% coding jump on the same foundation is a free upgrade path we will test the day weights or API land. We have also run DeepSeek's V4-Flash locally on DGX Spark hardware and published our own benchmark notes, and we installed DeepSeek Harness in its first 48 hours. The pattern we see across both labs: raw model quality has converged enough that the differentiators are now workflow fit, cost shape and openness. Pick the model that matches how you actually work: subscription simplicity and security tooling with GLM-5.3, or open weights, giant outputs and per-token control with V4-Pro.
Which Should You Pick?
Pick V4-Pro-0813 if you self-host or need the strongest published software-engineering scores, maximum output length for long agent runs, or transparent per-token economics with the option to cache aggressively. Pick GLM-5.3 if you want a flat monthly coding plan, work in Claude Code or OpenCode, or your team touches security review and vulnerability discovery, where its CyberGym lead and its documented real-world vulnerability finds are unmatched in open weights. If you are building agent infrastructure rather than just using it, the pragmatic answer is both: V4-Pro inside DeepSeek Harness for the runtime experiment, GLM-5.3 as the daily coding driver, and revisit in two weeks when GLM's weights drop and independent leaderboards catch up.
Frequently Asked Questions
Are GLM-5.3 and V4-Pro really open source?
V4-Pro-0813 is fully open now under MIT, including commercial use. GLM-5.3's weights are promised roughly two weeks after launch, pending security review of its offensive capabilities. Zhipu has consistently released prior GLM weights, so the delay reads as caution, not retreat.
Which model is cheaper for heavy daily coding?
For a solo developer or small team, GLM's US$18-80 monthly plan is almost certainly cheaper than per-token billing at any serious volume. At API scale with strong cache hit rates, V4-Pro's cache-hit price of US$0.003625 per million tokens can undercut everything, but the new peak rates from August 16 penalise business-hours work in Australia.
Is GLM-5.3 better than Claude?
Zhipu claims parity with Claude Fable 5 on coding and agent tasks, and its CyberGym score of 84.5% edges Anthropic's Mythos at 83.8% in Zhipu's own evaluations. Independent verification is pending, and on ExploitBench Anthropic still leads by more than 23 points. Treat parity claims as directional until leaderboards update.
Can I run these models locally?
V4-Pro-0813 weights are downloadable now under MIT, though at 1.6 trillion parameters practical local inference needs serious multi-GPU hardware or disk-streaming setups. GLM-5.3 at 743B is lighter on paper and its weights are expected by late August 2026. For typical local rigs, the smaller siblings, V4-Flash and GLM's compact variants, remain the realistic options.
What happens on August 16?
DeepSeek's peak and off-peak API pricing begins at 16:00 UTC on August 16, 2026. Off-peak rates are US$0.66 input and US$1.98 output per million tokens, and peak rates are US$1.32 and US$3.96. In AEST, peak covers 11am-2pm and 4pm-8pm, so Australian business-hours usage lands almost entirely in the expensive windows.
Flowtivity builds and operates AI agent systems for growing businesses across Australia and beyond. We test every model we recommend on our own infrastructure first. Reach out via flowtivity.ai.