Articles
191 articles on running a business with AI
Page 4
Showing 61 to 80 of 191
Kimi K2.7 Complete Review: Benchmarks, Cost, and Local Inference
Kimi K2.7 achieves GPT-5.5-class performance at a fraction of the cost. Full review with benchmarks, local inference setup, and competitive analysis.

HY3 vs The Open Source Field: Is Tencent's 295B Model the Best Value in AI?
Tencent's HY3 delivers frontier-adjacent performance at $0.14 per million input tokens with a 5.4% hallucination rate. We compare it against GLM-5.2, DeepSeek V4, Kimi K2.6, and proprietary models on benchmarks, cost, and reliability.

Grok 4.5 vs Fable 5: The Cost of Intelligence Just Collapsed
Grok 4.5 delivers near-frontier performance at 80-90% lower cost than Fable 5. We break down the benchmarks, pricing, hallucination risks, and what it means for businesses building with AI in 2026.

Why a Custom AI Just Beat Every Frontier Model and What It Means for Your Business
A custom-trained model from Thinking Machines Lab and Bridgewater just outperformed GPT, Claude, and Gemini on financial tasks at 13.8x lower cost. Here's what it means for the future of AI in business.

Microsoft Just Open-Sourced the OS for AI Agents: Inside the Agent Governance Toolkit
Microsoft's Agent Governance Toolkit brings OS-like security, identity, and reliability to autonomous AI agents. One pip install, any framework, sub-millisecond enforcement.

Running a 284B AI Model on Your Desk: Our Real-World DSpark Deployment Log
We deployed DeepSeek V4 Flash with DSpark speculative decoding on 2x NVIDIA DGX Spark boxes. 49 tok/s, 1M token context, 6-way concurrency, zero API bill. Real numbers, real bugs, real fixes.

Palantir CEO Alex Karp Just Called Out OpenAI and Anthropic: Here's What He Said
Alex Karp went on CNBC and accused AI frontier labs of overselling, overcharging, and extracting enterprise IP. Here's why every Australian business should pay attention.

Nvidia's Qwen3.6-27B-NVFP4: 27B Parameter AI Now Runs on Consumer Blackwell GPUs
Nvidia's NVFP4 quantization of Alibaba's Qwen3.6-27B cuts memory by 2.5x with under 1% accuracy loss. The 19.7GB model runs on RTX PRO 6000 and DGX Spark, delivering up to 2,000+ tokens/sec with vLLM and MTP speculative decoding.

Structured Knowledge Extraction: How Hypergraphs Transform Unstructured Data
Every business sits on mountains of unstructured text. A new generation of LLM-powered extraction frameworks turns documents into databases. This guide covers structured extraction, why hypergraphs beat knowledge graphs, and real use cases for Australian businesses.

Loop Engineering: The Feedback Cycle That Turns AI Agents Into Reliable Workers
A great harness isn't enough. The real engine behind every reliable AI agent is the loop - observe, decide, act, verify. Here's how to engineer loops that actually ship work, the five patterns to choose from, the failure modes to avoid, and how it all connects back to harness engineering. Now updated with peer-reviewed research from PACT 2025 showing 3.54x performance gains from agentic loops in compiler optimization.

Qwen-AgentWorld: The AI That Learns by Simulating Reality
Alibaba Qwen team released the first language world model covering 7 agent environments in one model. It beats GPT-5.4 and Claude Opus 4.8 on environment simulation - and makes agents better in the process.

Cua: The Open-Source Framework Giving AI Agents Full Computer Access
Cua is an MIT-licensed infrastructure framework backed by Y Combinator that lets AI agents control desktop applications across macOS, Linux, and Windows — no APIs required. Here's how it works, how it compares to Browser-Use and OpenCUA, and why it matters for business automation in 2026.

Agentic Operations: A Practical Guide for Australian Businesses (2026)
Agentic operations is the next shift in business automation - autonomous AI agents that plan, execute, and learn. Here is how Australian businesses can build them practically.

Microsoft SkillOpt Explained: How to Train AI Agent Skills (2026 Guide)
SkillOpt treats markdown skill documents as trainable parameters, lifting GPT-5.5 accuracy by +23.5 points without fine-tuning. Here is what it is, how it works, and why it matters.

OpenClaw vs Hermes Agent: 2026 Comparison (Updated June)
Updated June 2026: Honest comparison of OpenClaw and Hermes Agent covering multi-model orchestration, pricing, memory systems, and real business use cases. Both are open source - the right choice depends on your needs.

The Interoperability Thesis: How OpenClaw Turned AI Agents From Hype Into Infrastructure
The AI industry bottleneck is not model intelligence - it is interoperability. OpenClaw built the protocol layer that makes agentic operations realistic, with natural reinforcement learning baked in.

Kimi K2.7 Code Review: Open-Source 1T Parameter Model Cuts Reasoning Tokens 30%
Moonshot AI's Kimi K2.7 Code is an open-source 1 trillion parameter coding model that reduces reasoning token usage by 30% while posting double-digit benchmark gains over K2.6.

Google's DiffusionGemma: The Model That Writes Entire Paragraphs at Once
Google just dropped DiffusionGemma, a 26B open model that generates text like an AI image generator, not a typewriter. 1000+ tokens per second, Apache 2.0 license, and it can actually solve Sudoku. Here's what it means for builders.

We Ran DeepSWE at 1M Context vs 262K. The Results Surprised Us.
Real-world A/B benchmark running DeepSWE tasks on DeepSeek V4 Flash at 1M vs 262K context. The 1M run was 3x faster but produced identical results. Here is what we learned about local LLM agent benchmarks.

We Ran DeepSeek V4 Flash at 1M Context on Two NVIDIA DGX Sparks. Here is What Happened.
Real-world benchmarks running DeepSeek V4 Flash (284B MoE) across two NVIDIA DGX Sparks with tensor parallelism over 200Gbps RoCE. 41 tok/s at 1 million token context, 3x faster than single-node. Includes how to run your AI agent for free.

One email, most weeks
What actually changed in AI and what to do about it. No pitch, and you can leave in one click.