Models & benchmarks
28 articles
A new frontier model lands most weeks and almost every launch claims a benchmark win. These articles audit those claims: what a model costs per million tokens, where it fails, how open weights compare to hosted APIs, and which numbers matter once a model is inside a workflow. Comparisons are run on hardware we own, and the method is always stated.
Latest
Showing 1 to 20 of 28
Meta's Byte Latent Transformer Explained: Why Byte-Level Models Could Replace Tokenization
Meta's Byte Latent Transformer removes the tokenizer, matches Llama 3 at 8B scale with up to 50% fewer inference FLOPs, and Fast BLT cuts memory bandwidth by over 50% again.
GLM 5.3 Flash vs DeepSeek V4.1 Flash: The Best Model for a 256GB Dual DGX Spark Cluster
GLM 5.3 Flash beats DeepSeek V4.1 Flash on a 256GB dual DGX Spark cluster, not on raw quality but on fit: a documented two-Spark recipe at 29-70 tok/s with 1M context, while DeepSeek V4.1 Flash needs three to four boxes. Full comparison, tok/s math, and the serving recipe.
Recurrent Looped Transformer Explained: What Infinite Reasoning Depth Actually Means
The Recurrent Looped Transformer (RLT) re-runs one 48-layer decoder stack recurrently, so its compute path grows to t × 48 blocks while per-token cost stays flat. Announced by Yifan Zhang on September 12, 2026, it drew 386,100 views in a day and ships no benchmarks yet. Here is the architecture, the honest reading of infinite reasoning depth, and what it means for AI agents.
DeepSeek V4.1 Flash Benchmarks: Open-Weights Model Beats GPT-5.6 Sol at Agentic Coding
DeepSeek V4.1 Flash (released 10 September 2026) beats GPT-5.6 Sol and Claude Opus-5.0 on DeepSWE, AutomationBench, Agent's Last Exam and CyberGym. 552B open-weights MoE with 1M-token context and $0.30 per million peak input pricing. Full benchmark tables, pricing maths and what it means for Australian teams.
Terminal-Bench 4.0 Exposes Inflated AI Scores: Gemini 3.8 Flash Falls From 87.6% to 19.7%
Terminal-Bench 4.0 cut Gemini 3.8 Flash from 87.6% to 19.7% and Muse Spark 1.3 to 33.3% while GPT-6 Astra (59.6%) and Claude Fable 5.1 (55.1%) held the top. Inside the benchmaxxing debate and what it means for choosing AI models.
US vs China AI Models Compared: GPT-6 Astra, Fable 5.1, GLM 5.3, Kimi K3, DeepSeek V4, Qwen 3.8
Ten frontier AI models from the US and China compared: benchmarks, token prices and workload routing for GPT-6 Astra, Claude Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, GLM 5.3, Kimi K3, DeepSeek V4 and Qwen 3.8.
GLM-5.3-Flash vs DeepSeek V4 Flash on 2× DGX Spark: Real-World DeepSWE Benchmark Results
We benchmarked GLM-5.3-Flash NVFP4 against DeepSeek V4 Flash 0731 on a two-unit DGX Spark cluster using the DeepSWE coding benchmark. DeepSeek solved 2 of 5 tasks with a 1M-token context window at 41 to 66 tok/s. GLM-5.3-Flash solved 0 of 5, capped at 24K context by a serving-stack kernel bug, and decoded 3 to 4 times slower. Full results, root causes, and what it means for self-hosting coding agents.
Nvidia's $12.9 Billion Hugging Face Deal and the Nemotron Plan, Explained
Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion. What the deal covers, how the Nemotron Coalition plans to reach state of the art, and what it means for businesses building on open AI models.
DeepSeek V4-Flash-Vision-Exp: Multimodal Agents Near Opus 4.8 at V4 Flash Prices
DeepSeek's experimental V4-Flash-Vision-Exp adds vision to the bargain agent model: 83.9 Terminal Bench 2.1, 59.3 DeepSWE, close to Opus 4.8 on multimodal agent benchmarks, same price as V4 Flash. Benchmarks, API usage, image billing and dsh harness 0.1.1 support explained.
GLM-5.3 vs DeepSeek V4-Pro: The 24-Hour Showdown
GLM-5.3 and V4-Pro-0813 shipped 24 hours apart, both claiming the open-weights coding crown. Deep research, self-reported benchmarks separated honestly, pricing in AUD context and decision infographics.
DeepSeek V4-Flash Beats Its Own Pro Model: Agent Benchmarks That Just Changed the Game
DeepSeek V4-Flash just scored 82.7 on Terminal Bench 2.1, beating V4-Pro-Preview by 14.7%. At $0.14 per million input tokens, this is the most cost-effective agent model on the market.
Kimi K2.7 Code vs MiniMax M3: Open-Source AI Coding Models Compared
MiniMax M3 vs Kimi K2.7 Code compared head-to-head. Full benchmark table, cost analysis, and local deployment guide from dual DGX Spark testing.

GLM-5.2: The Open-Source AI Model Beating GPT-5.5 at 1/6th the Cost
GLM-5.2 is the strongest open-weight coding model with 81.0 on Terminal-Bench 2.1. Full review with benchmarks, architecture, and deployment options.

HY3 vs The Open Source Field: Is Tencent's 295B Model the Best Value in AI?
Tencent's HY3 delivers frontier-adjacent performance at $0.14 per million input tokens with a 5.4% hallucination rate. We compare it against GLM-5.2, DeepSeek V4, Kimi K2.6, and proprietary models on benchmarks, cost, and reliability.

Grok 4.5 vs Fable 5: The Cost of Intelligence Just Collapsed
Grok 4.5 delivers near-frontier performance at 80-90% lower cost than Fable 5. We break down the benchmarks, pricing, hallucination risks, and what it means for businesses building with AI in 2026.

Qwen-AgentWorld: The AI That Learns by Simulating Reality
Alibaba Qwen team released the first language world model covering 7 agent environments in one model. It beats GPT-5.4 and Claude Opus 4.8 on environment simulation - and makes agents better in the process.

Kimi K2.7 Code Review: Open-Source 1T Parameter Model Cuts Reasoning Tokens 30%
Moonshot AI's Kimi K2.7 Code is an open-source 1 trillion parameter coding model that reduces reasoning token usage by 30% while posting double-digit benchmark gains over K2.6.

Google's DiffusionGemma: The Model That Writes Entire Paragraphs at Once
Google just dropped DiffusionGemma, a 26B open model that generates text like an AI image generator, not a typewriter. 1000+ tokens per second, Apache 2.0 license, and it can actually solve Sudoku. Here's what it means for builders.

We Ran DeepSWE at 1M Context vs 262K. The Results Surprised Us.
Real-world A/B benchmark running DeepSWE tasks on DeepSeek V4 Flash at 1M vs 262K context. The 1M run was 3x faster but produced identical results. Here is what we learned about local LLM agent benchmarks.

We Ran DeepSeek V4 Flash at 1M Context on Two NVIDIA DGX Sparks. Here is What Happened.
Real-world benchmarks running DeepSeek V4 Flash (284B MoE) across two NVIDIA DGX Sparks with tensor parallelism over 200Gbps RoCE. 41 tok/s at 1 million token context, 3x faster than single-node. Includes how to run your AI agent for free.

One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.