Module 06 / 08

Evals

Measure model behavior so quality improves for reasons you can explain.

The point

Without evals, AI product work becomes vibes. Evals turn errors into a system you can improve.

Start with

Understanding check

You should be able to…

  • write examples that represent real product failures
  • separate unit evals, human review, and production monitoring
  • use error analysis to choose the next change
  • know when an eval is being gamed

Practice with an agent

Turn five failures into checks

beginner · starter repository, JSON fixtures, executable verification skill

Make

five fixtures from real failures and a command that reports which behaviors pass or fail

You know it works when

the checks catch the seeded regression and a fresh agent can run them without asking you how

Go deeper by building

Build an eval set for an LLM feature

intermediate · typescript or python, json fixtures, model API, simple report output

Make

a repeatable eval harness with examples, expected behavior, and scoring notes

You know it works when

compare two prompts or models and explain the regression you would ship or reject