Skip to content

ArticlesAnalysis

Claude Can Now Build Your Evals and Hillclimb Your App Against Them

Anthropic's new claude-api skill commands build evals and hillclimb apps against them: four eval properties, overfitting guards, and a support benchmark that cut cost about 5x while accuracy rose.

Claude Can Now Build Your Evals and Hillclimb Your App Against Them
On this page
  1. What did Anthropic actually ship?
  2. What makes an eval trustworthy? Four properties
  3. Adversarial sampling: the eval trap that flatters you
  4. Adversarial sampling in practice: a project audit example
  5. Graders: cheapest tool that fits, then prove it works
  6. What is hillclimbing and why is overfitting the mortal enemy?
  7. Inside a hillclimb round: what Claude actually does
  8. The 78 percent cost cut: a worked example
  9. The skill that improved itself: 66 to 88 percent
  10. What the first-mover blogs missed
  11. What hillclimbing will not fix
  12. What Australian growing businesses should do Monday morning
  13. Frequently asked questions
  14. Is the claude-api skill free?
  15. Do I need real production data to build an eval?
  16. How big should my eval set be?
  17. Does hillclimbing only improve prompts?
  18. Can hillclimbing make my product worse?

Last Updated: 29 September 2026

Key Takeaways

  • Anthropic shipped eval design and hillclimbing guidance inside the claude-api skill for Claude Code, announced 28 September 2026 and free to install.
  • Two commands drive the workflow: /claude-api build-eval interviews you and builds a graded evaluation, /claude-api hillclimb improves your app against it one change at a time.
  • A trustworthy eval has four properties: tasks mirror production, scores rise with stronger models, the best model sits well below 100 percent for headroom, and variance stays low.
  • The hillclimber makes one attributable change per round, never reads the held-out test set, and reverts any patch that only helps the train set as suspected overfitting.
  • In Anthropic's support benchmark, the loop cut cost to roughly one fifth while held-out accuracy rose from 78.6 percent to 90.5 percent.

Anthropic shipped one of the least flashy but most consequential updates for AI engineering this year: Claude Code can now design your evaluations, grade them, and hillclimb your application against them without fooling itself. The announcement from the ClaudeDevs account on X on 28 September 2026 drew about 3,400 likes and 5,000 bookmarks in under a day against roughly 271,000 views, a reaction profile practitioners reserve for things they have been doing painfully by hand.

The guidance lives inside the claude-api skill, an open-source install from Anthropic's GitHub skills repository. Two slash commands appear in Claude Code. /claude-api build-eval interviews you, builds an evaluation inside your codebase, and pauses for approval at each gate. /claude-api hillclimb improves your application against that evaluation one change at a time, with a held-out test set catching overfitting along the way.

Why this matters more than another model drop: the gap between an AI prototype and an AI product is almost never the model. It is whether anyone measured the thing. According to Anthropic's own engineering essays, scoring failures are among the most common ways an eval is misconfigured, and most teams shipping AI features still cannot answer the simplest board-level question: is the output good, and is it getting better? This announcement turns that discipline from consulting-grade effort into a two-command workflow.

What did Anthropic actually ship?

Claude Code now ships with two commands that turn eval design and hillclimbing from a written discipline into an executable one. According to Anthropic's claude.dev article "Automating eval design and hillclimbing with Claude", the guidance encodes internal eval principles into the claude-api skill so Claude Code applies them in your repository directly.

The ladder has two rungs. Build the eval first, climb second. And the scope of the second rung is yours to choose: before hillclimbing, Claude asks what it may change, from your system prompt to model and effort parameters, and what you are optimizing, performance or cost while performance holds.

Build-eval and hillclimb two-command workflow diagram
How it works: build-eval turns your production examples into a graded eval, hillclimb closes the loop between score and change, and a held-out test set guards every step.

The interesting part is what each rung refuses to do without you: no inputs get sampled until you approve them, no grader ships before you have watched it score, and no patch survives if the held-out set does not confirm it.

What makes an eval trustworthy? Four properties

Anthropic's post distills eval design into four checks, and each one targets a specific way teams lie to themselves. Getting these right is the difference between an eval that guides engineering and one that generates unfixable arguments.

First, tasks must mirror production. The instinct when building an eval is to grab cases that are easy to generate or easy to grade. That produces an eval measuring nothing you care about. The build-eval workflow samples inputs in a deliberate order: production transcripts first, after asking about data retention and sensitive fields, then bug reports and support tickets, then five to ten cases you write by hand, then synthetic cases anchored to your real examples. Before running anything, it renders a review page showing every input and waits for your confirmation.

Second, performance should rise with stronger models and more thinking. If a more capable model at higher effort scores the same or worse, your eval is broken, not the model. Culprits are usually an ambiguous task or a miscalibrated grader. This is a free diagnostic: monotonicity doubles as a sanity check on the test itself.

Third, there must be headroom at the frontier. The most capable model at the highest effort should score well below 100 percent, otherwise you cannot reliably judge how changes impact performance. Anthropic's tell for a broken eval: a task that fails every run regardless of replicates, which usually means the task is impossible or ambiguous. The fix is their two-expert standard: a good task is one where two domain experts would reach the same verdict, and everything the grader checks is stated in the task.

Fourth, low run-to-run variance. High variance comes from ambiguous tasks, graders that flip verdicts on identical output, effort not applied consistently, or leftover environment state, a file or git history, handing the agent the answer. Variance is noise, and noise buries small improvements.

That fourth property deserves more airtime than it gets, because it sounds like an enterprise concern. It is not. Small teams tune prompts by vibes, enterprises drown in MLOps platforms, and there is a workable middle: a 20 to 50 case eval in a spreadsheet, a grader you have personally watched score a dozen transcripts, and a rule that no change ships unless the held-out score moves. Anthropic's own cost-reduction example used 44 tickets. That is the entire scale of rigor required.

Adversarial sampling: the eval trap that flatters you

The most novel section of the post is the one fast-moving readers will skip, and it quietly invalidates most internal evals: how cases get chosen.

Model capability is jagged. If you pick eval cases because today's model fails them, you are sampling the valleys of that specific model's capability surface, and the eval ends up measuring that model's failure fingerprint rather than what is intrinsically hard for your application. Swap in the next model version and your carefully collected hard cases evaporate, because they were never hard, just holes in a particular model.

The fix sounds soft but is sharp: include a case only if a human can articulate why it is hard before it goes in. And there is a mirror-image trap on the data side: users sometimes try what they expect to work, so a task distribution drawn strictly from user traffic skews easy. Bug reports and support tickets carry the honest distribution; the happy path flatters you.

Adversarial sampling trap in AI eval design
How it works: sampling the cases a model fails today measures that model's holes, not task difficulty, and the eval breaks at the next model swap.

Adversarial sampling in practice: a project audit example

Say you run a document-heavy workflow, invoices in, structured data out, and you want an eval for your extraction agent. The valley-sampling path: run the current model on 100 invoices, collect the 12 it got wrong, and build the eval from those. You have just memorized one model's failure fingerprint. The next model version, tuned differently, passes all 12 on raw competence and tells you nothing new. The human-judgment path: pick invoices a finance professional calls hard, mixed VAT rules, a scanned receipt at an angle, a line item that references a footnote, and keep them because two domain experts agree on the expected output and the grader's checks are stated in the task. That eval survives model swaps, because it measures your task, not a model's temporary edges.

Graders: cheapest tool that fits, then prove it works

After inputs come verdicts. The skill proposes the cheapest grader that fits the output space, in escalating order, per Anthropic's claude.dev article.

  • Programmatic verification: exact match, a label from a fixed set, schema-valid JSON, or tests that pass. Deterministic, free, and boring, which is exactly right when output space is constrained.
  • LLM-as-judge: for open-ended output with many valid answers but clear quality criteria. A second model reads the input, the output, and a rubric written as checkable claims, not a 1-to-5 vibe scale, and returns a score with reasoning. With a baseline available, the judge reads both answers in random order without knowing which is the baseline and picks the better one, which kills position bias. You pick the judge model, and it should not be the model under test.

Then the part almost everyone skips: Claude grades a handful of cases and asks whether you would have scored any of them differently. Anthropic recommends reading a sample of scored transcripts before believing your evaluator, precisely because scoring failures are among the most common ways an eval is misconfigured. During the baseline run, the skill runs the grader twice on the same output and reports if the verdict changed, checks for timeouts, API errors, and cut-off answers so infrastructure noise does not pass as model variance, and warns if the baseline already sits at 95 percent or higher, pointing the hillclimb at cost or latency instead.

Why so paranoid about the grader? Because a miscalibrated grader poisons everything downstream: baselines look wrong, the hillclimb chases phantom regressions, and the final report lies with confidence intervals attached. An eval with a wrong grader is worse than no eval, because it produces confident, wrong conclusions.

What is hillclimbing and why is overfitting the mortal enemy?

Once the eval is trustworthy, hillclimbing is the improvement loop: run, read failures, make one change, re-run, keep or revert. According to Anthropic, it suits surfaces that are cheap to modify and easy to attribute. Prompts, skills, tool descriptions, model choices, and effort levels are cheap to change and revert. Open-ended harness rewrites are not, and an unscoped objective like "improve the agent" is a common stall.

The scarier failure is overfitting to the eval, and Anthropic gives it the clearest treatment published outside a research paper. Even a well-designed eval never exactly matches your production distribution, so systems that score well can still fail production. Their example: an eval task benefits from OCR while production rarely does, the hillclimb adds an OCR tool to your harness, the benchmark score rises, and production does not improve at all. The eval leaked into the system.

The countermeasures, straight from the article:

  • Split the cases. A train set the hillclimber may read, and a test set it never sees. If train scores improve while test stays flat, that is the overfitting tell, and the skill reverts the patch.
  • Never paste failures into the prompt. If the hillclimber reads failing transcripts, the fix goes into behavior, never the eval answers into context.
  • Keep the answers structurally out of the model's reach so it cannot reward-hack by finding them directly.
Hillclimb loop with train test split and revert rules
How it works: each round proposes one patch, measures it against train and test, keeps only what both confirm, and reverts anything else.

One more scope rule worth stealing: pick a surface where the metric is directly coupled to the change. Anthropic's example is skill triggering, where the metric is trigger rate and the modified surface is the skill description. Attribution is what makes each round meaningful; without it you are just vibrating the system and calling the wobble progress.

Inside a hillclimb round: what Claude actually does

The execution discipline is where this stops being advice and starts being leverage, because the loop automates the exact behaviors senior engineers do manually and often skip.

You choose the surfaces Claude may modify: system prompt, skills or instruction files, tool descriptions, model choice, effort level and other API parameters, or harness code. You state the goal, better performance, or lower cost while performance holds. Then the loop runs with its checksheets enforced every round:

Before round one, Claude verifies the eval's noise floor is smaller than the smallest improvement you would act on. If it is not, it says so and suggests more repetitions or cases rather than spending rounds on changes too small to measure. With a cost goal, it starts from the common cost drivers Anthropic documented separately: prompt caching, auditing the prompt for compatibility with the selected model, and model and effort selection.

Each round, Claude reads the previous round's train transcripts and proposes one change as a patch. It aims the change at a root cause, rewriting the section that causes failure or adding a missing rule, rather than rewording a line, and sizes it so its effect can show above the eval's noise. Then the decision rule: train and test both improve, keep the patch. Regression, revert. Train improves but test is flat, suspected overfitting, revert.

On stall, after two or three flat rounds, or early if no single fix could gain more than the noise floor, Claude sorts every remaining train failure by cause without editing anything. Anthropic calls this reflection step out explicitly, and it catches eval ambiguity, grader bugs, and harness errors that no single patch would have fixed. Tasks that never improve despite obvious content gaps are flagged as defective tasks or graders rather than wasted on more rounds.

At completion, Claude leaves the code at the version that did best on the test set for your goal, reports test result against baseline with confidence intervals, and, if the gain is within noise, says so and recommends against merging. An AI that tells you not to merge is worth more than one that always claims victory.

The 78 percent cost cut: a worked example

The numbers from Anthropic's internal customer-support benchmark, cross-referenced with their separate cost-reduction writeup, show what this loop does on a deliberately modest eval.

Setup: 44 support tickets, 30 for the hillclimbing search, 14 held out that the search never sees. Baseline: Opus 4.8 at default high effort, 74.4 percent decision accuracy on search tickets, 4.6 cents per ticket in token cost.

  • Prompt hygiene first: the hillclimb removed mandatory tool-call rituals, a scratchpad step, and contradictory rules. No model change. Cost drops first, because prompts accrete dead weight that every call pays for.
  • Model step down one: Opus 5.5 at low effort cleared the accuracy bar at 87.8 percent and cut cost to 1.9 cents per ticket, less than half the start. Part of that comes from Opus 5.5 pricing, with input and output tokens 20 percent cheaper than Opus 4.8 and cache reads 60 percent cheaper, and because it cleared the bar the loop kept stepping down.
  • Model step down two: Sonnet 5 at low effort scored about the same, 88.9 percent, at about half Opus's cost, roughly 1 cent per ticket.
  • One targeted prompt edit: routing rules and a refund-cap cross-reference lifted Sonnet 5 to 98.9 percent at the same cost.

On the 14 held-out tickets: 90.5 percent final accuracy against 78.6 percent for the original setup, at about one fifth of the cost. One automated afternoon, measured against a holdout, delivered a 5x cost reduction with 12 points of accuracy on top. Every team running a support agent should read that as an invoice for measurement work they have not done.

Claude hillclimb support benchmark results infographic
At a glance: the hillclimb cut cost about 5x while held-out accuracy rose from 78.6 to 90.5 percent across four moves.

What the example quietly demonstrates is sequencing: the cheapest changes came first, model swaps only happened after prompt hygiene, and the most expensive resource, rounds, was never wasted on edits the noise floor could not measure. That ordering is a strategy, not a lucky run.

The skill that improved itself: 66 to 88 percent

The second example is the best one, because the subject was the hillclimber's own house: the claude-api skill, which teaches developers the current Claude API, had drifted from the API's actual behavior. According to the article, Anthropic built an eval from their documentation and the skill started at 66 percent.

Given access to docs and SDKs, the hillclimber found the skill was missing coverage of eight API features. Adding sections: 74 percent. Fixing C# and Java type-table errors: 77 percent. Then the score stalled, and the reflection step earned its keep: failures were bucketed by root cause, which surfaced something a single-fix round would have missed. The skill's content was present and correct, but Claude kept writing older API shapes from training priors, like extended thinking with a fixed token budget that recent Opus models reject in favor of adaptive thinking, or superseded versions of the web search and web fetch tools.

The fix was a mapping table near the top of the skill that routes Claude from the forms it remembers to the current ones, plus moving C# and Java warnings against fixed-budget thinking above the examples that trigger the wrong reflex. That lifted performance to 80 percent. Rewording one flawed task, whose grader wanted a chain of at least three catches when the task asked for one, and fixing a grader whose instructions contradicted the docs, where testing the real API proved the docs right, brought performance to roughly 88 percent.

Skill self-improvement hillclimb from 66 to 88 percent diagram
How it works: content gaps first, then a no-edit reflection round, then a mapping table that corrects stale priors and two grader fixes.

The generalizable insight for anyone maintaining prompts, skills, or RAG systems: stale knowledge is rarely a knowledge addition problem. It is a reflex-correction problem. The model does not need the new fact dropped into the document; it needs a route from the wrong fact it reliably retrieves to the right one. Mapping tables that name the old pattern and point at the current one outperformed eight whole feature sections, and that matches what we see maintaining our own agent skills, where a correction placed above the triggering example beats one buried in the appendix.

What the first-mover blogs missed

Aggregator coverage of the launch is mostly announcement replication. Read the actual article and three structural points got less airtime than they deserve.

Reflection rounds make no edits. When the score stalls, the loop shifts to pure analysis: bucket every remaining failure by cause, then act on the biggest bucket. It is the same discipline a senior consultant applies by refusing to react to the loudest ticket. Most failed improvement programs fail during the edit, not the analysis, and Anthropic hard-coded the pause.

The privacy default is offline by design. Result pages are static files that open locally and load nothing from the network. Eval data and production transcripts stay on your machine. For regulated workloads and client-confidential engagements, that is not a nicety, it is the difference between adoptable on day one and stuck in procurement.

Targeted reflection beats scattered fixes. In the skill-improvement example, the highest-leverage edit was not new content but a routing table correcting where Claude's priors go wrong. The lesson travels to every RAG and prompting system in production: diagnose failure by class, treat the class, and stop patching instances.

What hillclimbing will not fix

Fairness requires stating the scope, because a technique this useful accretes magical expectations fast.

  • Garbage eval in, confident garbage out. The headroom and variance checks are real, but the workflow still trusts your task distribution. Production failures that never appear in eval cases stay invisible no matter how many rounds run.
  • Not a research loop. It tunes surfaces you name toward a goal you state. It will not invent a new product interaction that makes the eval obsolete, and per Anthropic, open-ended objectives are exactly where it stalls.
  • Eval scores are not business metrics. 98.9 percent on a ticket benchmark is not the same as customers being happier. Somebody still has to establish that the eval proxies the outcome a customer pays for.
  • Blind spots are structural. If the eval undersamples a failure class, hillclimbing undersamples fixing it. Distribution coverage is a human judgment call the loop cannot make for you.

Think of it as taking the measurement discipline previously locked inside frontier labs and pricing it at a skill install. The remaining gap is judgment: choosing what to measure, and checking that the grader matches reality. That part is still yours.

What Australian growing businesses should do Monday morning

The teams reading this with 11 to 200 employees do not have an eval platform budget, and they do not need one. The free, reproducible core of this workflow is not vendor-locked: a small eval set drawn from real tickets, a grader you have personally watched work, a train/test split, one change per round with revert rules. A three-person team can run exactly this discipline with a spreadsheet and a copy of Claude Code.

The quick-start path: install the claude-api skill from Anthropic's GitHub, run /claude-api build-eval on the workflow that costs you the most support hours, approve or replace the sampled inputs with five to ten cases from your worst tickets, and watch the grader score a dozen transcripts yourself before trusting any number. Then set an explicit goal. If quality already looks fine, ask for cost reduction at performance parity, which Anthropic notes is a strong objective even when an eval is near saturation. One caution from our own deployments: budget the API spend for rounds in advance and let the noise-floor check size your eval, or the loop will happily spend a week of inference credits proving what a dozen transcripts would have told you.

Anthropic's quiet launch says something bigger than the feature: the competitive edge in applied AI is shifting from prompt craft to measurement discipline. Teams that measure will quietly eat the teams that demo, and the tooling to be on the right side of that line is now a free skill away.

Frequently asked questions

Is the claude-api skill free?

Yes, the skill is open source on Anthropic's skills GitHub repository and installs into Claude Code. You pay only the API costs of running eval and hillclimb rounds, which the workflow is explicitly designed to keep small, and the generated result pages are static files that load nothing from the network.

Do I need real production data to build an eval?

No, but it helps. The build-eval workflow prefers production transcripts, bug reports, and tickets, and asks about retention and sensitive data before sampling them. If you only have a few real examples, the skill generates synthetic cases anchored to them, and five to ten hand-written cases are explicitly part of the recommended mix.

How big should my eval set be?

Small beats absent. Anthropic's worked example used 44 tickets, split 30 for the search and 14 held out. Size should be driven by variance rather than vanity: if the noise floor sits above the smallest improvement you would act on, add repetitions or cases, which the skill checks before round one.

Does hillclimbing only improve prompts?

No. Allowed surfaces include the system prompt, skills and instruction files, tool descriptions, model choice, effort level, other API parameters, and even harness code. The guidance favors surfaces that are cheap to modify and attributable, which is why text is the common target and open-ended harness edits are a stall risk.

Can hillclimbing make my product worse?

Overfitting is the hazard the design attacks hardest. The held-out test set is never read, patches that only help train are reverted automatically, failure content is never pasted into prompts, and the final report recommends against merging if the gain is within noise. Residual risk lives in how well your eval matches true production traffic, which remains a human judgment call.


AJ Awan is the founder of Flowtivity, an AI automation consultancy on the Gold Coast, Australia, and a former EY management consultant. He deploys governed agent systems for growing businesses and holds to one review rule: if a feature cannot be evaluated, it cannot be trusted in production.

  • Anthropic
  • claude code
  • AI Evaluation
  • hillclimbing
  • LLM-as-judge
  • AI engineering

One email a month, no noise

Practical AI notes for Australian businesses. Unsubscribe anytime.

One good place to start

What would you like to take off your plate?

Bring a process that feels repetitive or harder than it needs to be. We’ll help you find a practical first step.

Book a free consult

A free 1-hour conversation with AJ. No pressure, no pitch.