AI agents
61 articles
Agents are the part of AI that changes how work gets done rather than how text gets written. These articles cover the frameworks and harnesses we run in production, what breaks when an agent meets a real business process, and how to tell a genuine capability gain from a demo. Written for people building or buying agent systems, not for people reading about them.
Latest
Showing 1 to 20 of 61
Jev Use Cases: 12 Places It Beats a Text LLM, and 4 Where It Cannot
Real Jev use cases from week one: email triage, slop detection, browser agents, lead scoring. How Jev's parallel scoring differs from text-output LLMs, plus four limits including the poker eval failure.
OpenAI's Agents API Explained: Cloud Agents on the Managed Codex Harness
OpenAI's Agents API runs the open source Codex harness as a managed cloud service: durable sessions, automatic context compaction, multi-agent orchestration, and optional sandboxes. How the architecture works, what it costs, and when to choose it over the Agents SDK.
Alibaba Open Code Review: The Open Source AI Reviewer That Out-Engineered Claude Code
Alibaba's open source AI code reviewer scored 33.90% precision versus Claude Code's 7.23% with the same underlying model on AACR-Bench, at roughly one ninth of the token cost. How the hybrid architecture works, benchmark numbers, our first-hand install test, and how to deploy it.
Harness Engineering Explained: What Meta's Auto-RecSys Means for AI in Business
Harness engineering, the craft of building memory, scripts, and playbooks around an AI model, cut operational failures roughly 87 percent in Meta's Auto-RecSys. Here is what happened and how any business can apply it.
Google's Procedural Graphs, Explained Simply: The Self-Evolving Playbook for AI Agents
Google's Procedural Graphs paper gives LLM agents editable what-to-do-next knowledge as a graph, evolved automatically from execution feedback. First in 21 of 24 benchmarks.
Obscura: The Rust Headless Browser Built for AI Agents (Tested)
Obscura is an open-source Rust headless browser for AI agents and web scraping: 34MB of memory, a single binary, V8 JavaScript and Chrome DevTools Protocol compatibility. We benchmarked it live on a production VPS and traced how it inspired Cloudflare's agent-first Kitesurf browser.
OpenClaw 2.0: What the 16,000 Pull Request Release Changes for AI Agents
OpenClaw 2.0 is version 2026.8.1, the largest release in project history with 933 contributors and 16,000+ pull requests. Here is what changed and whether businesses should deploy it.
What Is WebMCP? Making Your Website Agent-Ready in 2026
WebMCP is the proposed W3C standard that turns websites into structured toolkits for browser AI agents. How it works, MCP vs WebMCP, the 2026 business case, and a six-step plan to make your site agent-ready.
Headlong: The AI Agent That Never Stops Thinking (Laude Institute Deep Dive)
Deep dive into Headlong, Laude Institute open source Bash microharness for persistent AI agents that think continuously. Includes cost breakdown, self-repair case study, and comparison with reactive harnesses like OpenClaw.
How AI Agents Solved the Reporting Equation Businesses Face
AI agents solve the reporting equation by replacing manual data preparation (which consumes 80% of reporting effort per Gartner) with an autonomous perceive-reason-act loop, turning reporting from a chore into a continuous source of decisions.
Why DeepSeek Harness Is the New Meta for AI Agents
DeepSeek Harness hit 183K GitHub stars in 9 days, 100K in two days at 2,100 stars per hour, with 10,000+ community plugins. Why dsh is the OpenClaw moment all over again, one level deeper: the harness, not the model, is the product.
I Built a Fireflies.ai Clone With the dsh Harness: 109 Agent Steps at 48 tok/s on DeepSeek V4 Flash
Personal experience running DeepSeek's dsh harness: one prompt, 109 agent steps, 48 tok/s on V4 Flash 0731, a working Fireflies.ai clone plugin, and the approval-gate behaviour that earned my trust. Includes session metrics, cost math, V4 Flash vs Opus comparison and secure VPS + Tailscale self-hosting steps.
TrueForge vs DeepSeek Harness vs Claude Managed Agents: 2026 Comparison
TrueFoundry's open-source TrueForge harness claims up to 75% cheaper agent runs than Claude Managed Agents. We break down the benchmark, compare it with DeepSeek Harness, Codex CLI and deepagents, and explain what it means for businesses choosing agent infrastructure.
Coding From the Fireplace: When AI Agents Run Your Tests
How remote AI agent development with Claude enables coding from anywhere while autonomous agents run regression tests on self-improvement loops. Real-world experience with phone-based development and 24/7 AI testing.
From Loops to Graphs: The Next Paradigm in AI Agent Engineering
Graph engineering is replacing loop-based AI agents. Learn the 5-stage methodology, decision matrix, typed edges framework, and cost/performance tradeoffs. Includes visual infographics, code examples, and benchmark data from GraphRAG-Bench.

Hermes Agent vs OpenAI Codex vs Claude Cowork: The Coding Agent Showdown
Hermes Agent vs OpenAI Codex vs Claude Code compared. Which coding agent fits your workflow? Based on real dual DGX Spark deployment experience.
ZCode The Open-Source Coding Agent Harness Chasing Cursor and Claude Code
ZCode is Z.ai's open-source coding agent harness for GLM-5.2. How it compares to Cursor and Claude Code for agentic coding workflows.

Google Open Knowledge Format: How Plain Markdown Files Are Becoming the Brain of AI Agents
Google OKF is a plain markdown format for AI agent knowledge exchange. Implementation guide, competitive analysis, and practical business applications.

Loop Engineering: The Feedback Cycle That Turns AI Agents Into Reliable Workers
A great harness isn't enough. The real engine behind every reliable AI agent is the loop - observe, decide, act, verify. Here's how to engineer loops that actually ship work, the five patterns to choose from, the failure modes to avoid, and how it all connects back to harness engineering. Now updated with peer-reviewed research from PACT 2025 showing 3.54x performance gains from agentic loops in compiler optimization.

Agentic Operations: A Practical Guide for Australian Businesses (2026)
Agentic operations is the next shift in business automation - autonomous AI agents that plan, execute, and learn. Here is how Australian businesses can build them practically.

One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.