> ## Content Index
> Fetch the complete content index at: https://www.metatalks.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Anthropic adds eval and hillclimb commands to Claude Code for automated app tuning
- URL: https://www.metatalks.ai/anthropic-claude-code-eval-hillclimb-commands/
- Published: 2026-09-30T10:41:00.000Z
- Updated: 2026-09-30T10:40:59.000Z
- Author: Al
- Tags: News, Agentic AI, #newswire

**The commands carry a machine-learning routine into app development: Claude builds a test set from real examples, then changes the app step by step and keeps only what improves results on held-back cases.**

Anthropic added the two commands, /claude-api build-eval and /claude-api hillclimb, to its claude-api skill for Claude Code and described them in a September 28 [post on its developer blog](https://claude.dev/blog/automating-eval-design-and-hillclimbing/?ref=metatalks.ai). On the company's own customer-support benchmark, the setup the tuning produced scored 90.5% on tickets it never saw, against 78.6% for the original, at about a fifth of the cost.

The build-eval command starts by interviewing the user and gathers inputs in a fixed order: production transcripts, after asking about data retention and sensitive information, then bug reports and support tickets, hand-written cases and, last, synthetic ones. Claude lists every input and waits for confirmation before going on.

It then proposes the cheapest grader that fits: a code check where the output is constrained, and a second model scoring each answer against checkable claims where many answers are valid. Claude runs the grader twice on the same output to see whether the verdict changes, and reports a baseline score with a confidence interval.

For hillclimb, the user decides what Claude may change, from the system prompt and tool descriptions to the model, effort level and harness code, and whether the goal is better results or lower cost at the same quality. Claude splits the cases at random into a training set and a test set and makes one change per round. It reverts a change when results regress, and also when the training score rises while the test score stays flat, a common warning sign of overfitting. If the final gain is within noise, it recommends against merging.

The support benchmark had 44 tickets, 14 of them held out. Starting from Opus 4.8 at high effort and 4.6 cents a ticket, Claude stripped contradictory rules and mandatory tool calls from the prompt, then moved to Opus 5.5 and finally to Sonnet 5, both at low effort. Tuned the same way, the claude-api skill itself rose from 66% to about 88% on an evaluation set built from Anthropic's documentation.

![Chart of the claude-api skill's evaluation score rising over successive hillclimb rounds](https://storage.ghost.io/c/ef/d4/efd46b24-40c6-4ee1-b2e7-34720b3b26fd/content/images/2026/09/anthropic-claude-api-hillclimb-fig9.png)

Evaluation score of Anthropic's claude-api skill over successive hillclimb rounds, from 66% to about 88%. Source: [Anthropic](https://claude.dev/blog/automating-eval-design-and-hillclimbing/?ref=metatalks.ai)

The human work shifts from writing prompts and code to defining success: choosing the examples and approving the grader. Anthropic's post counts faulty scoring among the most common ways an eval goes wrong, and a metric the tuning can game or overfit makes that human step matter more, not less.