September 28, 2026 — The same day Sonnet 5.5 shipped, Anthropic published the missing half of a production agent stack: how to measure the app without lying to yourself. Lance Martin’s Automating eval design and hillclimbing with Claude (12 min) adds two commands to the claude-api skill:
/claude-api build-eval
/claude-api hillclimb
@ClaudeDevs and Martin posted the same hour. X summaries (including Grok’s “Anthropic Launches Tools to Automate AI Evaluations”) compressed it to two slash commands. The article is longer than that: four eval-design rules, adversarial sampling, grader validation, harness overfitting, and two worked climbs — cost on a support bench and quality on the claude-api skill itself (66% → ~88%).
explainx.ai already covered /claude-api hillclimb as a cost-search tool in Claude Platform cost and effort. This post is the eval-design companion: when the bench is junk, hillclimbing just overfits faster.
TL;DR
| Question | Answer |
|---|---|
| What shipped? | Eval principles + build-eval + hillclimb in the claude-api skill |
| Who wrote it? | Lance Martin, Sep 28, 2026, claude.dev |
| build-eval? | Interviews you, samples cases, you approve inputs and grader, runs a baseline with CI |
| hillclimb? | One patch per round; train/test split; revert if test is flat or scores drop |
| Headline result? | Support bench: 78.6% → 90.5% held-out, ~⅕ cost (Opus 4.8 high → Sonnet 5 + prompt) |
| Skill dogfood? | claude-api skill eval 66% → ~88% over 24 rounds |
| Vs plugin eval? | plugin eval = A/B your plugin; this = improve the app |
Four properties of a eval that is not theater
Martin’s Figure 1 is the acceptance test. If your suite fails these, do not hillclimb yet.
| Property | What it means | Failure mode |
|---|---|---|
| Mirrors production | Task mix is what you actually ship | Easy-to-grade toys, not real tickets |
| Scales with model/effort | Stronger model / higher effort should score higher | Ambiguous tasks or a broken grader |
| Headroom at the frontier | Best model at max effort is well below 100% | Saturated bench; you cannot see deltas |
| Low variance | Tight error bars; same output → same grade | Leftover git/files handing the agent the answer |
Impossible tasks fail every replicate. Good tasks are ones two domain experts would score the same, and everything the grader checks is in the prompt.
That is the same discipline as evals for engineers and PMs (write goldens before you buy a platform) and Terminal-Bench (scored, repeatable, not a demo).
Do not sample only where today’s model dies
Adversarial sampling (Figure 2): if you pick cases because this model failed them, you measure that model’s valleys, not what is hard for the job. The next model’s curve is smoother — your “hard set” becomes noise.
Martin’s rule: pick hard cases because a human can say why they are hard before you include them. Use production bugs. Do not trust traffic blindly — users try what they expect to work, so raw logs skew easy. LangChain’s Viv Trivedy quoted the production-mirror line and pointed at trace → eval loops at LangChain scale — same idea as Jev agent evals.
/claude-api build-eval — what it actually does
In Claude Code the skill interviews you, writes the eval in-repo, and stops for approval.
Case order:
- Production transcripts (after retention / PII questions)
- Bug reports and tickets
- Five to ten handwritten cases
- Synthetics anchored on your codebase (and on a few real examples you provide)
You get a local HTML page of every input. You confirm they are representative. Illustration in the article: an inbox-router with 24 emails tagged billing / easy / ambiguous.
Grader, cheapest first:
| Output shape | Grader |
|---|---|
| Constrained (labels, JSON schema, tests) | Programmatic |
| Open-ended but checkable | LLM-as-judge with a rubric of claims, not a 1–5 vibe scale |
| You have a baseline | Blind pairwise judge (random order, judge ≠ model under test) |
Claude grades a handful and asks whether you would score them differently. Then it tells you cases × repeats × models, runtime, runs a baseline, and prints a score + confidence interval. Artifacts: cases, grader, runner, one JSON line + full transcript per case, a static local page (no network). Ask for a chart; it builds another static page.
Diagnostics on the baseline:
- Grader stability — same output graded twice
- Plumbing — timeouts, API errors, truncated answers (not “model variance”)
- Headroom — ~95%+ → warn: climb cost/latency, not quality
Hillclimbing without fooling yourself
Hillclimb surfaces that are cheap to edit: prompts, skills, tool descriptions. Open-ended harness rewrites are a stall risk. The metric should couple to the surface (e.g. skill trigger rate vs skill description). A strong default objective is cost at parity when quality is already saturated.
How evals leak into the harness (Figure 5)
| Eval quirk | Bad harness “fix” |
|---|---|
| Tasks need OCR | Add an OCR tool you never use in prod |
Tasks live in /app | always cd /app && pytest |
| Distinctive phrasing | Prompt tuned to those strings |
| Failures you already read | One patch per failure |
| Public repo with answers | curl the gold (reward hack) |
Three counters the skill applies:
- Train / test split — climber may read train; never test
- Never paste failures into the prompt
- Keep answers out of reach
/claude-api hillclimb — the loop
You choose the edit surface: system prompt, skills, tool descriptions, model / effort / API params, or harness code. You pick performance vs cost at hold.
Before round 1: noise vs smallest shippable delta. If noise is larger, add repeats or cases.
Each round: one patch aimed at the root of a train failure, not a reworded line. Run eval.
| Train | Test | Action |
|---|---|---|
| Up | Up | Keep |
| Up | Flat | Overfit → revert |
| Either down | — | Revert |
Stall for 2–3 rounds (or if no single fix beats noise): bucket remaining train failures by cause. Catch ambiguous cases, grader bugs, harness errors. Only legitimate failures go to more rounds.
Finish: code left at best test-set version. Report test vs baseline with CIs. If gain is within noise, do not merge.
Inbox-router schematic: variant that defines queues + tie-break wins 0.875 train and test; a variant that adds two worked examples is reverted (train up, test flat).
Example 1 — cost: ⅕ the bill, better held-out accuracy
Internal customer support bench: 44 tickets, 30 for search, 14 held out.
| Step | Setup | Train accuracy | Cost / ticket |
|---|---|---|---|
| Baseline | Opus 4.8, high effort | 74.4% | 4.6¢ |
| After prompt audit | Drop mandatory tool rituals, scratchpad, contradictions | — | — |
| Model step | Opus 5.5, low effort | 87.8% | 1.9¢ |
| Cheaper tier | Sonnet 5, low | 88.9% | ~1¢ |
| Prompt | Routing rules + refund-cap cross-ref | 98.9% | ~same 1¢ |
Held-out (never seen in search): final config 90.5% vs original 78.6%, about one-fifth the cost.
Part of the save is Opus 5.5 pricing (tokens 20% cheaper than 4.8, cache reads 60% cheaper) — but the climber did not stop there. It stepped down a tier once 5.5 cleared the bar, then fixed the prompt on Sonnet. That is the Sonnet 5.5 story in miniature: same list price, fewer tokens, lower effort if the eval allows it.
Grok/X attributed the 90.5% / ⅕-cost line to “Vox.” The primary numbers are in Martin’s article. Treat social recaps as pointers, not citations.
Example 2 — the claude-api skill eating its own cooking
Eval derived from API docs. Skill started at 66%. Hillclimber got docs + SDKs.
| Round cluster | What changed | Score |
|---|---|---|
| Baseline | — | 66% |
| Coverage | Eight missing features | 74% |
| Types | C# / Java type-table errors | 77% |
| After stall | Map trained priors → current API (adaptive thinking, current web tools) | 80% |
| Eval bugs | Task asked for one error type, grader wanted three; grader vs docs | ~88% |
The stall reflection is the underrated feature: no edit, only sort failures. Content can be present while Claude writes old API shapes. A table at the top of the skill from remembered forms to current ones is the fix — same class of problem plugin eval was built for when a new model quietly breaks skills.
What people are asking
“Is this just another LLM-as-judge toy?”
No. Programmatic first. Judge only when the output space is open. You read scored transcripts. Grader is re-run on identical output. That is closer to Hamel/Shreya eval practice than to a dashboard that always says 4.2/5.
“Will this rewrite my agent harness into a benchmark monster?”
Only if you allow harness code as an edit surface and skip the test split. Default the climber onto prompts and skills. SkillOpt automated skill text; this loop is supervised and revert-happy.
“I already run plugin eval. Do I need this?”
Both. Plugin eval answers does this plugin help. build-eval answers is my product eval honest. hillclimb answers can I buy the same quality cheaper. Order: goldens → plugin eval if you ship plugins → build-eval for the app → hillclimb.
“Can I point it at LangSmith / production traces?”
The skill asks for traces. It does not replace your observability vendor. Export a sample you are allowed to retain, then let build-eval interview you. Do not dump PII into the repo because a slash command asked nicely.
Commands to run today
# You need an eval for a real problem (router, support, skill, API codegen)
/claude-api build-eval
# You already have an eval and a goal (quality or cost@parity)
/claude-api hillclimb
Pair with /claude-api migrate this project to claude-sonnet-5-5 from the Sonnet 5.5 building guide if you are still on Sonnet 5 strings, then re-run the eval — new models change both quality and token mix.
Credits in the article: Misha Khalman (skill), reviews from Khalman, Michael Segner, Matt Bell, Matt Thanabalan.
Related reading
- Claude Platform: cost, caching, effort, hillclimb
- claude plugin eval
- Claude Sonnet 5.5 building guide
- What a task costs on Opus 5.5
- AI evals for engineers and PMs
- LangChain Jev agent evals
- Terminal-Bench 2.0
- Primary: Automating eval design and hillclimbing
Commands, worked-example scores, and overfitting rules follow Lance Martin’s September 28, 2026 claude.dev article and the claude-api skill. CLI flags can change — run the skill in current Claude Code before you treat a number as a SLA.
