Module 06 / 08
Evals
Measure model behavior so quality improves for reasons you can explain.
The point
Without evals, AI product work becomes vibes. Evals turn errors into a system you can improve.
Start with
Read next
Understanding check
You should be able to…
- write examples that represent real product failures
- separate unit evals, human review, and production monitoring
- use error analysis to choose the next change
- know when an eval is being gamed
Practice with an agent
Turn five failures into checks
beginner · starter repository, JSON fixtures, executable verification skill
Make
five fixtures from real failures and a command that reports which behaviors pass or fail
You know it works when
the checks catch the seeded regression and a fresh agent can run them without asking you how
Go deeper by building
Build an eval set for an LLM feature
intermediate · typescript or python, json fixtures, model API, simple report output
Make
a repeatable eval harness with examples, expected behavior, and scoring notes
You know it works when
compare two prompts or models and explain the regression you would ship or reject
