Articles
191 articles on running a business with AI
Start here
All guidesLatest
Showing 1 to 20 of 191
Jev by TypeSafe AI: Is the 200x Faster Decision Model Too Good to Be True?
Is Jev too good to be true? A claim-by-claim audit of TypeSafe AI's 200x faster, 400x cheaper decision model, with real pricing math and HN skeptic pushback.
Meta's Byte Latent Transformer Explained: Why Byte-Level Models Could Replace Tokenization
Meta's Byte Latent Transformer removes the tokenizer, matches Llama 3 at 8B scale with up to 50% fewer inference FLOPs, and Fast BLT cuts memory bandwidth by over 50% again.
OpenAI's Agents API Explained: Cloud Agents on the Managed Codex Harness
OpenAI's Agents API runs the open source Codex harness as a managed cloud service: durable sessions, automatic context compaction, multi-agent orchestration, and optional sandboxes. How the architecture works, what it costs, and when to choose it over the Agents SDK.
GLM 5.3 Flash vs DeepSeek V4.1 Flash: The Best Model for a 256GB Dual DGX Spark Cluster
GLM 5.3 Flash beats DeepSeek V4.1 Flash on a 256GB dual DGX Spark cluster, not on raw quality but on fit: a documented two-Spark recipe at 29-70 tok/s with 1M context, while DeepSeek V4.1 Flash needs three to four boxes. Full comparison, tok/s math, and the serving recipe.
Recurrent Looped Transformer Explained: What Infinite Reasoning Depth Actually Means
The Recurrent Looped Transformer (RLT) re-runs one 48-layer decoder stack recurrently, so its compute path grows to t × 48 blocks while per-token cost stays flat. Announced by Yifan Zhang on September 12, 2026, it drew 386,100 views in a day and ships no benchmarks yet. Here is the architecture, the honest reading of infinite reasoning depth, and what it means for AI agents.
What Is Recursive Self-Improvement? 5 Autonomy Levels That Separate Real AI Self-Improvement From Hype
A 33-author survey (arXiv 2609.11873) maps recursive self-improvement across five autonomy levels and a Headroom-Closed Index that shows where LLMs stall. Here is what it means, where it works today, and how to audit self-improving agent claims.
Alibaba Open Code Review: The Open Source AI Reviewer That Out-Engineered Claude Code
Alibaba's open source AI code reviewer scored 33.90% precision versus Claude Code's 7.23% with the same underlying model on AACR-Bench, at roughly one ninth of the token cost. How the hybrid architecture works, benchmark numbers, our first-hand install test, and how to deploy it.
Harness Engineering Explained: What Meta's Auto-RecSys Means for AI in Business
Harness engineering, the craft of building memory, scripts, and playbooks around an AI model, cut operational failures roughly 87 percent in Meta's Auto-RecSys. Here is what happened and how any business can apply it.
Anthropic's 2030 Economy Report: What It Means for Trades, Construction and Small Business
Anthropic's September 2026 report models three AI futures for 2030: GDP up 1.6% to 32.4%, unemployment from 4.6% to nearly 12%, and rising wages for trades and construction while knowledge work automates. Here is what it means for Australian trade, construction and allied health businesses.
Google's Procedural Graphs, Explained Simply: The Self-Evolving Playbook for AI Agents
Google's Procedural Graphs paper gives LLM agents editable what-to-do-next knowledge as a graph, evolved automatically from execution feedback. First in 21 of 24 benchmarks.
DeepSeek V4.1 Flash Benchmarks: Open-Weights Model Beats GPT-5.6 Sol at Agentic Coding
DeepSeek V4.1 Flash (released 10 September 2026) beats GPT-5.6 Sol and Claude Opus-5.0 on DeepSWE, AutomationBench, Agent's Last Exam and CyberGym. 552B open-weights MoE with 1M-token context and $0.30 per million peak input pricing. Full benchmark tables, pricing maths and what it means for Australian teams.
Terminal-Bench 4.0 Exposes Inflated AI Scores: Gemini 3.8 Flash Falls From 87.6% to 19.7%
Terminal-Bench 4.0 cut Gemini 3.8 Flash from 87.6% to 19.7% and Muse Spark 1.3 to 33.3% while GPT-6 Astra (59.6%) and Claude Fable 5.1 (55.1%) held the top. Inside the benchmaxxing debate and what it means for choosing AI models.
Anthropic Researcher Quits Over AI Extinction Risk: What Business Leaders Should Do
Anthropic researcher Jacob Coxon quit with an extinction warning and Anthropic's own alignment lead put decade scale human extinction odds above 10 percent. Verified statements, context, and a practical governance playbook for businesses.
Spring 2026: The Economy Won't Wait, and Neither Should Your Business
Spring 2026 is the window to stop deliberating about the economy and deploy AI agents. RBA rate data, adoption statistics, real client outcomes, and the Flowtivity Discover, Design, Deploy methodology.
Drafted.ai Review 2026: Free AI House Plans vs the Open Source Alternatives
Drafted.ai gives away AI floor plans with free CAD and BIM export. We reviewed the beta and mapped the open source equivalents: HouseGAN++, Graph2Plan, Sweet Home 3D with a new MCP server, and IfcOpenShell. Here is the honest comparison and a $0 stack you can self-host.
Obscura: The Rust Headless Browser Built for AI Agents (Tested)
Obscura is an open-source Rust headless browser for AI agents and web scraping: 34MB of memory, a single binary, V8 JavaScript and Chrome DevTools Protocol compatibility. We benchmarked it live on a production VPS and traced how it inspired Cloudflare's agent-first Kitesurf browser.
The 15/80/5 Method: Let AI Agents Own the 80%
Gary Vaynerchuk's 15-80-5 rule, expanded into a full operating methodology for AI agents: humans direct the first 15%, agents execute the 80%, humans finish the last 5%. Includes stage-by-stage mechanics, handoff gates, failure modes and first-hand production data.
US vs China AI Models Compared: GPT-6 Astra, Fable 5.1, GLM 5.3, Kimi K3, DeepSeek V4, Qwen 3.8
Ten frontier AI models from the US and China compared: benchmarks, token prices and workload routing for GPT-6 Astra, Claude Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, GLM 5.3, Kimi K3, DeepSeek V4 and Qwen 3.8.
Meta Muse Spark 1.3 Benchmarks: What the Release Means for AI Agents
Meta released Muse Spark 1.3 on September 2, 2026: 75.4% on DeepSWE 1.1, 98.5% on long-context MRCR, a 1M token window and 25% fewer tokens. What the benchmarks, the agentic capabilities and Meta closed-model pivot mean for AI agents and the AI industry.
One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.