Stop "Fixing on Instinct"──Claude Code Gains build-eval / hillclimb to Automate Eval Creation and Self-Improvement
On September 28, Anthropic added build-eval and hillclimb to Claude Code's claude-api skill. They semi-automate eval creation and self-improvement, with built-in safeguards against overfitting, replacing "tuning by instinct" with tuning by the numbers.
When trying to improve the quality of AI-powered features, many teams tend to stop at "I tweaked the prompt and it kind of feels better now." On September 28, 2026, Anthropic added a mechanism that turns this "fixing on instinct" into "fixing by the numbers" to the claude-api skill for Claude Code. The additions are two commands, /claude-api build-eval and /claude-api hillclimb, which let Claude Code itself semi-automatically handle building evals and improving against those evals.
First, Build Your Own Grading Criteria
build-eval is a command that assembles a test set and a grading method inside your codebase. Rather than diving straight into building, it interviews you and gathers the materials in a fixed order.
- Production conversation logs (after confirming how data retention and sensitive information are handled)
- Bug reports and support tickets
- Hand-written cases (5 to 10)
- Synthetic cases tied to real examples
The collected materials are presented as a list (review page), and Claude waits for your approval before moving on. For the grading method, it uses programmatic checks such as exact match or schema validation when the output has a fixed form, and LLM-as-judge measured against "verifiable facts" when the output is free-form. It also inspects grading variance and environmental noise, and warns you that "this eval is too easy" if the initial score exceeds roughly 95%.
Change One Thing at a Time, Keep Only What Improves
hillclimb is a command that improves your app against the eval you built. As the name hill climbing suggests, it repeatedly changes just one thing per round, adopts the change if the result improves, and reverts it if the result gets worse. You can specify what may be changed, and the main targets are the following.
- System prompts, skills, and instruction files
- Tool descriptions
- The model used, the effort level, and other API parameters
- Harness (app-side control) code
Safeguards So You Don't Fool Yourself
The heart of this workflow is that it questions overfitting (over-optimizing to the test cases alone) on its own. hillclimb randomly splits the eval into a "train" set and a "test" set, and if the train score rises but the test score stays flat, it judges this to be "an apparent improvement only" and rolls the change back. Naturally, it also reverts any change that simply lowers the score. If the score stalls for two or three rounds, instead of continuing with incremental tweaks, it analyzes the remaining failures by cause and changes its approach. In other words, a brake against moving forward on a "sense that things got better" is built into the procedure.
The Demo's Shift Toward "Cheaper and Smarter"
In the customer support example Anthropic showed, the model configuration itself was swapped out between before and after optimization. A setup that had been running an expensive top-tier model at high effort was replaced with one using a lighter model at low effort, and the accuracy actually went up.
| Item | Before tuning | After tuning |
|---|---|---|
| Model / effort | Opus 4.8, high | Sonnet 5, low |
| Accuracy on new tickets | 74.4% | 90.5% |
| Cost per ticket | ~4.6 cents | ~1 cent (about one-fifth) |
Anthropic also applied this method to the claude-api skill itself, explaining that it raised the eval pass rate from 66% to about 88% over 24 rounds. Measurement overturned the assumption that "using a top-tier model makes it smart."
What Changes on the Ground in Development
Until now, tuning prompts and selecting models tended to become a person-dependent "craft." build-eval and hillclimb distill that craft into a reproducible procedure: "build an eval → try one change at a time → record what is kept and what is reverted." As a result, you can make decisions such as swapping a model for a cheaper configuration while preserving accuracy, based on numbers rather than gut feeling. Use requires Claude Code v2.1.259 or later, and a good starting point is to build an eval from a small number of real cases in your own repository and run hillclimb for a few short rounds.
Things to Watch Out for Behind the Convenience
At the same time, there are caveats that come with automation. First, because build-eval uses production logs as material, you yourself need to correctly judge how data retention and sensitive information are handled (the procedure builds in confirmation steps, but the judgment itself is a human responsibility). Second, because optimization advances toward "what is being measured," a lax eval can send optimization off in the wrong direction. That is precisely why Anthropic also emphasizes building the grading criteria, for example by warning you when the initial score is too high. In exchange for the ease of fixing things by the numbers, you could argue that the responsibility for deciding "what counts as correct" actually weighs more heavily on the human side.
References: Automating eval design and hillclimbing with Claude (claude.dev Blog) / Anthropic adds eval and hillclimb commands to Claude Code (metatalks.ai) / Anthropic Built a Better Way to Improve AI Agents (The Neuron) / Claude Code Helps Build App Evaluations: Version 2.1.259 or Later Required (vibecoding.tech)


