---
title: "Evaluating and improving an AI app, explained with an exam · How-to · Consonance"
description: "Two commands of the claude-api skill: build-eval builds an evaluation inside your codebase, with your approval, and hillclimb improves your app against it, one change at a time, keeping examples aside to catch overfitting."
canonical: "https://consonance.fyi/en/mode-d-emploi/automating-eval-design-and-hillclimbing"
lang: "en"
published: "2026-10-03"
---

# Evaluating and improving an AI app, explained with an exam

_Claude series · sheet no. 3_

## In one sentence

Two commands of the claude-api skill: build-eval builds an evaluation inside your codebase, with your approval, and hillclimb improves your app against it, one change at a time, keeping examples aside to catch overfitting.

## The image: an exam

| In everyday life | In Claude Code |
| --- | --- |
| an exam | an eval: a signal on how your app performs a task |
| revising, one thing at a time | hillclimbing: one change per round · `/claude-api hillclimb` |
| an exam that looks like real life | tasks that mirror production |
| the best student gets the best mark | a stronger model, more thinking: a better score |
| nobody scores 100% | the best model stays well below 100% |
| the same paper, the same mark | low variance from run to run |
| a teacher who writes trick questions for one student | cases picked only because today's model fails them |
| a grader who marks the same paper twice | the grader runs twice on the same output |
| learning past papers by heart | overfitting: better on the eval than in production |
| a mock exam never seen | a test set the hillclimber never reads |

## To try

1. Update Claude Code: the claude-api skill ships inside it.

   ```
   claude update
   ```
2. Build the eval inside your codebase. Claude interviews you and waits for your approval of the examples and the grader.

   ```
   /claude-api build-eval
   ```
3. Improve your app against it, towards your goal (performance, or cost while performance holds).

   ```
   /claude-api hillclimb
   ```

> **The trap** Cramming. If the train set improves while the test set stays flat, that is an overfitting warning sign: Claude reverts the change. Never paste failure content into the prompt.

## Sources

- [claude.dev · Automating eval design and hillclimbing with Claude](https://claude.dev/blog/automating-eval-design-and-hillclimbing/)

pace: [steady](https://consonance.fyi/media/lessons/automating-eval-design-and-hillclimbing/en-pose.mp4?v=12533728)

Independent summary, not affiliated with Anthropic. Claude is a trademark of Anthropic, PBC. The everyday comparisons are images to help understand, not quotes: every command and figure comes from the source.

## Next

- [How-to](https://consonance.fyi/en/mode-d-emploi.md)
- [previous sheet: Claude Code mods: two examples from the guide, and a real mod for recording your screen](https://consonance.fyi/en/mode-d-emploi/getting-started-with-claude-code-mods.md)
- [archives](https://consonance.fyi/en/archives.md)
- [method](https://consonance.fyi/en/methode.md)
- [Consonance](https://consonance.fyi/en/index.md)
- Languages: [Français](https://consonance.fyi/mode-d-emploi/automating-eval-design-and-hillclimbing.md) · [Tiếng Việt](https://consonance.fyi/vi/mode-d-emploi/automating-eval-design-and-hillclimbing.md)
- For machines: [llms.txt](https://consonance.fyi/llms.txt) · [RSS](https://consonance.fyi/en/rss.xml) · [sitemap.xml](https://consonance.fyi/sitemap.xml)
- HTML version: https://consonance.fyi/en/mode-d-emploi/automating-eval-design-and-hillclimbing
