explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Four properties of a eval that is not theater
  • /claude-api build-eval — what it actually does
  • Hillclimbing without fooling yourself
  • /claude-api hillclimb — the loop
  • Example 1 — cost: ⅕ the bill, better held-out accuracy
  • Example 2 — the claude-api skill eating its own cooking
  • What people are asking
  • Commands to run today
  • Related reading
← Back to blog

explainx / blog

Claude Code Can Build Evals and Hillclimb Them: /claude-api build-eval

Claude Code, AI Evals, Anthropic, Agent Skills

Anthropic’s Lance Martin added /claude-api build-eval and hillclimb to Claude Code on Sep 28, 2026 — production-shaped evals, held-out tests, 90.5% vs 78.6%.

Sep 29, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Claude Code Can Build Evals and Hillclimb Them: /claude-api build-eval

September 28, 2026 — The same day Sonnet 5.5 shipped, Anthropic published the missing half of a production agent stack: how to measure the app without lying to yourself. Lance Martin’s Automating eval design and hillclimbing with Claude (12 min) adds two commands to the claude-api skill:

text
/claude-api build-eval
/claude-api hillclimb

@ClaudeDevs and Martin posted the same hour. X summaries (including Grok’s “Anthropic Launches Tools to Automate AI Evaluations”) compressed it to two slash commands. The article is longer than that: four eval-design rules, adversarial sampling, grader validation, harness overfitting, and two worked climbs — cost on a support bench and quality on the claude-api skill itself (66% → ~88%).

explainx.ai already covered /claude-api hillclimb as a cost-search tool in Claude Platform cost and effort. This post is the eval-design companion: when the bench is junk, hillclimbing just overfits faster.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What shipped?Eval principles + build-eval + hillclimb in the claude-api skill
Who wrote it?Lance Martin, Sep 28, 2026, claude.dev
build-eval?Interviews you, samples cases, you approve inputs and grader, runs a baseline with CI
hillclimb?One patch per round; train/test split; revert if test is flat or scores drop
Headline result?Support bench: 78.6% → 90.5% held-out, ~⅕ cost (Opus 4.8 high → Sonnet 5 + prompt)
Skill dogfood?claude-api skill eval 66% → ~88% over 24 rounds
Vs plugin eval?plugin eval = A/B your plugin; this = improve the app

Four properties of a eval that is not theater

Martin’s Figure 1 is the acceptance test. If your suite fails these, do not hillclimb yet.

table · 3 cols
PropertyWhat it meansFailure mode
Mirrors productionTask mix is what you actually shipEasy-to-grade toys, not real tickets
Scales with model/effortStronger model / higher effort should score higherAmbiguous tasks or a broken grader
Headroom at the frontierBest model at max effort is well below 100%Saturated bench; you cannot see deltas
Low varianceTight error bars; same output → same gradeLeftover git/files handing the agent the answer

Impossible tasks fail every replicate. Good tasks are ones two domain experts would score the same, and everything the grader checks is in the prompt.

That is the same discipline as evals for engineers and PMs (write goldens before you buy a platform) and Terminal-Bench (scored, repeatable, not a demo).

Do not sample only where today’s model dies

Adversarial sampling (Figure 2): if you pick cases because this model failed them, you measure that model’s valleys, not what is hard for the job. The next model’s curve is smoother — your “hard set” becomes noise.

Martin’s rule: pick hard cases because a human can say why they are hard before you include them. Use production bugs. Do not trust traffic blindly — users try what they expect to work, so raw logs skew easy. LangChain’s Viv Trivedy quoted the production-mirror line and pointed at trace → eval loops at LangChain scale — same idea as Jev agent evals.

/claude-api build-eval — what it actually does

In Claude Code the skill interviews you, writes the eval in-repo, and stops for approval.

Case order:

  1. Production transcripts (after retention / PII questions)
  2. Bug reports and tickets
  3. Five to ten handwritten cases
  4. Synthetics anchored on your codebase (and on a few real examples you provide)

You get a local HTML page of every input. You confirm they are representative. Illustration in the article: an inbox-router with 24 emails tagged billing / easy / ambiguous.

Grader, cheapest first:

table · 2 cols
Output shapeGrader
Constrained (labels, JSON schema, tests)Programmatic
Open-ended but checkableLLM-as-judge with a rubric of claims, not a 1–5 vibe scale
You have a baselineBlind pairwise judge (random order, judge ≠ model under test)

Claude grades a handful and asks whether you would score them differently. Then it tells you cases × repeats × models, runtime, runs a baseline, and prints a score + confidence interval. Artifacts: cases, grader, runner, one JSON line + full transcript per case, a static local page (no network). Ask for a chart; it builds another static page.

Diagnostics on the baseline:

  • Grader stability — same output graded twice
  • Plumbing — timeouts, API errors, truncated answers (not “model variance”)
  • Headroom — ~95%+ → warn: climb cost/latency, not quality

Hillclimbing without fooling yourself

Hillclimb surfaces that are cheap to edit: prompts, skills, tool descriptions. Open-ended harness rewrites are a stall risk. The metric should couple to the surface (e.g. skill trigger rate vs skill description). A strong default objective is cost at parity when quality is already saturated.

How evals leak into the harness (Figure 5)

table · 2 cols
Eval quirkBad harness “fix”
Tasks need OCRAdd an OCR tool you never use in prod
Tasks live in /appalways cd /app && pytest
Distinctive phrasingPrompt tuned to those strings
Failures you already readOne patch per failure
Public repo with answerscurl the gold (reward hack)

Three counters the skill applies:

  1. Train / test split — climber may read train; never test
  2. Never paste failures into the prompt
  3. Keep answers out of reach

/claude-api hillclimb — the loop

You choose the edit surface: system prompt, skills, tool descriptions, model / effort / API params, or harness code. You pick performance vs cost at hold.

Before round 1: noise vs smallest shippable delta. If noise is larger, add repeats or cases.

Each round: one patch aimed at the root of a train failure, not a reworded line. Run eval.

table · 3 cols
TrainTestAction
UpUpKeep
UpFlatOverfit → revert
Either down—Revert

Stall for 2–3 rounds (or if no single fix beats noise): bucket remaining train failures by cause. Catch ambiguous cases, grader bugs, harness errors. Only legitimate failures go to more rounds.

Finish: code left at best test-set version. Report test vs baseline with CIs. If gain is within noise, do not merge.

Inbox-router schematic: variant that defines queues + tie-break wins 0.875 train and test; a variant that adds two worked examples is reverted (train up, test flat).

Example 1 — cost: ⅕ the bill, better held-out accuracy

Internal customer support bench: 44 tickets, 30 for search, 14 held out.

table · 4 cols
StepSetupTrain accuracyCost / ticket
BaselineOpus 4.8, high effort74.4%4.6¢
After prompt auditDrop mandatory tool rituals, scratchpad, contradictions——
Model stepOpus 5.5, low effort87.8%1.9¢
Cheaper tierSonnet 5, low88.9%~1¢
PromptRouting rules + refund-cap cross-ref98.9%~same 1¢

Held-out (never seen in search): final config 90.5% vs original 78.6%, about one-fifth the cost.

Part of the save is Opus 5.5 pricing (tokens 20% cheaper than 4.8, cache reads 60% cheaper) — but the climber did not stop there. It stepped down a tier once 5.5 cleared the bar, then fixed the prompt on Sonnet. That is the Sonnet 5.5 story in miniature: same list price, fewer tokens, lower effort if the eval allows it.

Grok/X attributed the 90.5% / ⅕-cost line to “Vox.” The primary numbers are in Martin’s article. Treat social recaps as pointers, not citations.

Example 2 — the claude-api skill eating its own cooking

Eval derived from API docs. Skill started at 66%. Hillclimber got docs + SDKs.

table · 3 cols
Round clusterWhat changedScore
Baseline—66%
CoverageEight missing features74%
TypesC# / Java type-table errors77%
After stallMap trained priors → current API (adaptive thinking, current web tools)80%
Eval bugsTask asked for one error type, grader wanted three; grader vs docs~88%

The stall reflection is the underrated feature: no edit, only sort failures. Content can be present while Claude writes old API shapes. A table at the top of the skill from remembered forms to current ones is the fix — same class of problem plugin eval was built for when a new model quietly breaks skills.

What people are asking

“Is this just another LLM-as-judge toy?”

No. Programmatic first. Judge only when the output space is open. You read scored transcripts. Grader is re-run on identical output. That is closer to Hamel/Shreya eval practice than to a dashboard that always says 4.2/5.

“Will this rewrite my agent harness into a benchmark monster?”

Only if you allow harness code as an edit surface and skip the test split. Default the climber onto prompts and skills. SkillOpt automated skill text; this loop is supervised and revert-happy.

“I already run plugin eval. Do I need this?”

Both. Plugin eval answers does this plugin help. build-eval answers is my product eval honest. hillclimb answers can I buy the same quality cheaper. Order: goldens → plugin eval if you ship plugins → build-eval for the app → hillclimb.

“Can I point it at LangSmith / production traces?”

The skill asks for traces. It does not replace your observability vendor. Export a sample you are allowed to retain, then let build-eval interview you. Do not dump PII into the repo because a slash command asked nicely.

Commands to run today

text
# You need an eval for a real problem (router, support, skill, API codegen)
/claude-api build-eval

# You already have an eval and a goal (quality or cost@parity)
/claude-api hillclimb

Pair with /claude-api migrate this project to claude-sonnet-5-5 from the Sonnet 5.5 building guide if you are still on Sonnet 5 strings, then re-run the eval — new models change both quality and token mix.

Credits in the article: Misha Khalman (skill), reviews from Khalman, Michael Segner, Matt Bell, Matt Thanabalan.

Related reading

  • Claude Platform: cost, caching, effort, hillclimb
  • claude plugin eval
  • Claude Sonnet 5.5 building guide
  • What a task costs on Opus 5.5
  • AI evals for engineers and PMs
  • LangChain Jev agent evals
  • Terminal-Bench 2.0
  • Primary: Automating eval design and hillclimbing

Commands, worked-example scores, and overfitting rules follow Lance Martin’s September 28, 2026 claude.dev article and the claude-api skill. CLI flags can change — run the skill in current Claude Code before you treat a number as a SLA.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 29, 2026

Claude Sonnet 5.5 Is Live: Building Guide, Migration, and Claude Code Defaults

Claude Sonnet 5.5 joined the Claude 5.5 family on September 28, 2026 with unchanged list prices, native 1M context, thinking on by default, and a claude.dev playbook from Addy Osmani. This post distills when to pick Sonnet over Opus 5.5, breaking API changes, and how Claude Code v2.1.284 maps the sonnet alias — plus the IceSolst reference-image credit debate on X.

Sep 26, 2026

Claude Code Stops Hard-Stopping Mid-Task at the 5-Hour Reset

For most of 2026, hitting your Claude Code 5-hour usage limit mid-task meant an abrupt hard stop — generation cut off mid-response, context lost, work half-done. Anthropic has now fixed that specific failure mode. Here's what changed, why it mattered, and what it means for long-running agent sessions.

Sep 26, 2026

Claude Plugin Developer Portal: How to Submit, Track, and Ship

Anthropic shipped a dedicated developer portal for Claude plugins on September 26, 2026 — one place to submit a plugin, track its review, and see who's actually installing it. Here is exactly what it does, the two submission paths, and a step-by-step walkthrough from "I have an MCP server" to a live, analytics-tracked plugin.