AI agents
64 articles
Agents are the part of AI that changes how work gets done rather than how text gets written. These articles cover the frameworks and harnesses we run in production, what breaks when an agent meets a real business process, and how to tell a genuine capability gain from a demo. Written for people building or buying agent systems, not for people reading about them.
Latest
Showing 1 to 20 of 64
How we built our own booking system with Claude (and ended up copying Calendly)
We replaced Calendly with our own booking system, built with Claude in about ten days. What we built, how Claude helped, what a real phone test changed, and what independent AI checks caught before launch.

How we rebuilt our design system with Claude Opus 5.5
A two-day case study: AI site audit, four concepts judged by four agents, a new palette and ribbon system, live on 25 Sept 2026. Stages, tokens and costs.

What we learned running Claude dynamic workflows
11 multi-agent runs, roughly 17 million tokens, 3 usage-limit stops. What worked, what broke and how we designed around it with Claude Opus 5.5.

Jev Use Cases: 12 Places It Beats a Text LLM, and 4 Where It Cannot
Real Jev use cases from week one: email triage, slop detection, browser agents, lead scoring. How Jev's parallel scoring differs from text-output LLMs, plus four limits including the poker eval failure.
OpenAI's Agents API Explained: Cloud Agents on the Managed Codex Harness
OpenAI's Agents API runs the open source Codex harness as a managed cloud service: durable sessions, automatic context compaction, multi-agent orchestration, and optional sandboxes. How the architecture works, what it costs, and when to choose it over the Agents SDK.
Alibaba Open Code Review: The Open Source AI Reviewer That Out-Engineered Claude Code
Alibaba's open source AI code reviewer scored 33.90% precision versus Claude Code's 7.23% with the same underlying model on AACR-Bench, at roughly one ninth of the token cost. How the hybrid architecture works, benchmark numbers, our first-hand install test, and how to deploy it.
Harness Engineering Explained: What Meta's Auto-RecSys Means for AI in Business
Harness engineering, the craft of building memory, scripts, and playbooks around an AI model, cut operational failures roughly 87 percent in Meta's Auto-RecSys. Here is what happened and how any business can apply it.
Google's Procedural Graphs, Explained Simply: The Self-Evolving Playbook for AI Agents
Google's Procedural Graphs paper gives LLM agents editable what-to-do-next knowledge as a graph, evolved automatically from execution feedback. First in 21 of 24 benchmarks.
Obscura: The Rust Headless Browser Built for AI Agents (Tested)
Obscura is an open-source Rust headless browser for AI agents and web scraping: 34MB of memory, a single binary, V8 JavaScript and Chrome DevTools Protocol compatibility. We benchmarked it live on a production VPS and traced how it inspired Cloudflare's agent-first Kitesurf browser.
OpenClaw 2.0: What the 16,000 Pull Request Release Changes for AI Agents
OpenClaw 2.0 is version 2026.8.1, the largest release in project history with 933 contributors and 16,000+ pull requests. Here is what changed and whether businesses should deploy it.
What Is WebMCP? Making Your Website Agent-Ready in 2026
WebMCP is the proposed W3C standard that turns websites into structured toolkits for browser AI agents. How it works, MCP vs WebMCP, the 2026 business case, and a six-step plan to make your site agent-ready.
Headlong: The AI Agent That Never Stops Thinking (Laude Institute Deep Dive)
Deep dive into Headlong, Laude Institute open source Bash microharness for persistent AI agents that think continuously. Includes cost breakdown, self-repair case study, and comparison with reactive harnesses like OpenClaw.
How AI Agents Solved the Reporting Equation Businesses Face
AI agents solve the reporting equation by replacing manual data preparation (which consumes 80% of reporting effort per Gartner) with an autonomous perceive-reason-act loop, turning reporting from a chore into a continuous source of decisions.
Why DeepSeek Harness Is the New Meta for AI Agents
DeepSeek Harness hit 183K GitHub stars in 9 days, 100K in two days at 2,100 stars per hour, with 10,000+ community plugins. Why dsh is the OpenClaw moment all over again, one level deeper: the harness, not the model, is the product.
I Built a Fireflies.ai Clone With the dsh Harness: 109 Agent Steps at 48 tok/s on DeepSeek V4 Flash
Personal experience running DeepSeek's dsh harness: one prompt, 109 agent steps, 48 tok/s on V4 Flash 0731, a working Fireflies.ai clone plugin, and the approval-gate behaviour that earned my trust. Includes session metrics, cost math, V4 Flash vs Opus comparison and secure VPS + Tailscale self-hosting steps.
TrueForge vs DeepSeek Harness vs Claude Managed Agents: 2026 Comparison
TrueFoundry's open-source TrueForge harness claims up to 75% cheaper agent runs than Claude Managed Agents. We break down the benchmark, compare it with DeepSeek Harness, Codex CLI and deepagents, and explain what it means for businesses choosing agent infrastructure.
Coding From the Fireplace: When AI Agents Run Your Tests
How remote AI agent development with Claude enables coding from anywhere while autonomous agents run regression tests on self-improvement loops. Real-world experience with phone-based development and 24/7 AI testing.
From Loops to Graphs: The Next Paradigm in AI Agent Engineering
Graph engineering is replacing loop-based AI agents. Learn the 5-stage methodology, decision matrix, typed edges framework, and cost/performance tradeoffs. Includes visual infographics, code examples, and benchmark data from GraphRAG-Bench.

Hermes Agent vs OpenAI Codex vs Claude Cowork: The Coding Agent Showdown
Hermes Agent vs OpenAI Codex vs Claude Code compared. Which coding agent fits your workflow? Based on real dual DGX Spark deployment experience.
ZCode The Open-Source Coding Agent Harness Chasing Cursor and Claude Code
ZCode is Z.ai's open-source coding agent harness for GLM-5.2. How it compares to Cursor and Claude Code for agentic coding workflows.

One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.