Field reports
10 articles
First-hand reports from running this work: what we built, what it cost, what broke and what the numbers were.
Latest
Showing 1 to 10 of 10
Obscura: The Rust Headless Browser Built for AI Agents (Tested)
Obscura is an open-source Rust headless browser for AI agents and web scraping: 34MB of memory, a single binary, V8 JavaScript and Chrome DevTools Protocol compatibility. We benchmarked it live on a production VPS and traced how it inspired Cloudflare's agent-first Kitesurf browser.
GLM-5.3-Flash vs DeepSeek V4 Flash on 2× DGX Spark: Real-World DeepSWE Benchmark Results
We benchmarked GLM-5.3-Flash NVFP4 against DeepSeek V4 Flash 0731 on a two-unit DGX Spark cluster using the DeepSWE coding benchmark. DeepSeek solved 2 of 5 tasks with a 1M-token context window at 41 to 66 tok/s. GLM-5.3-Flash solved 0 of 5, capped at 24K context by a serving-stack kernel bug, and decoded 3 to 4 times slower. Full results, root causes, and what it means for self-hosting coding agents.
I Built a Fireflies.ai Clone With the dsh Harness: 109 Agent Steps at 48 tok/s on DeepSeek V4 Flash
Personal experience running DeepSeek's dsh harness: one prompt, 109 agent steps, 48 tok/s on V4 Flash 0731, a working Fireflies.ai clone plugin, and the approval-gate behaviour that earned my trust. Includes session metrics, cost math, V4 Flash vs Opus comparison and secure VPS + Tailscale self-hosting steps.
We Tested 10 AI Vision Models on Real Construction Plans: The Models Were Never the Problem
We tested 10 frontier AI vision models on real architectural drawings for construction takeoffs. Four of six models achieved within 10% accuracy. A median of four models hit 0.03% error. Total cost: $15. The bottleneck was never the models.
"I Just Ask My AI to Do It": What a Real AI Discovery Call Taught Me About Where Most Companies Actually Are
After sitting down with a Sydney construction firm to map out their AI strategy, I realised most companies are stuck between bottom-up experimentation and top-down ambition. Here's what a real discovery call revealed about where Australian businesses actually stand with AI adoption.

Running a 284B AI Model on Your Desk: Our Real-World DSpark Deployment Log
We deployed DeepSeek V4 Flash with DSpark speculative decoding on 2x NVIDIA DGX Spark boxes. 49 tok/s, 1M token context, 6-way concurrency, zero API bill. Real numbers, real bugs, real fixes.

We Ran DeepSWE at 1M Context vs 262K. The Results Surprised Us.
Real-world A/B benchmark running DeepSWE tasks on DeepSeek V4 Flash at 1M vs 262K context. The 1M run was 3x faster but produced identical results. Here is what we learned about local LLM agent benchmarks.

We Ran DeepSeek V4 Flash at 1M Context on Two NVIDIA DGX Sparks. Here is What Happened.
Real-world benchmarks running DeepSeek V4 Flash (284B MoE) across two NVIDIA DGX Sparks with tensor parallelism over 200Gbps RoCE. 41 tok/s at 1 million token context, 3x faster than single-node. Includes how to run your AI agent for free.

We Ran DeepSWE on Local Models. Here's What Actually Happened.
We tested DeepSeek V4 Flash, AEON-27B, and Step 3.7 Flash against the DeepSWE benchmark on DGX Spark hardware. All three scored zero. The story behind that zero is what matters.

Step 3.7 Flash Review: We Tested StepFun's 198B Model on a DGX Spark
StepFun released Step 3.7 Flash on May 29, 2026. We deployed it on our NVIDIA DGX Spark within 24 hours. 100% tool call success rate, SWE-Bench PRO 56.3, ClawEval 67.1 (first place). Here is our first-hand review with benchmark comparisons, local deployment guide, and DGX Spark performance data.

One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.