3 Picking evaluation metrics
This chapter covers
- What AI evaluation metrics are and what are their core components.
- Choosing between different metric types depending on evaluation goals and context.
- The main criteria that make a metric suitable for reliable decision making.
- Designing custom metrics for domain-specific needs and maintain them over time, as systems, data, and evaluation goals evolve.
“When a measure becomes a target, it ceases to be a good measure.”
— Marilyn Strathern
Once you have framed the evaluation by deciding what you are evaluating, why you are evaluating it, and which aspects of performance matter, the next step is to decide how those aspects will be measured. This is the role of metrics.
A metric is a measurement tool that quantifies some aspect of system behavior and turns abstract evaluation goals into observable evidence. Choosing metrics is one of the most important decisions in an evaluation because every metric captures only a particular aspect of performance and relies on specific assumptions. For example, accuracy measures how often a system produces the correct answer, while task success rate measures whether the system achieves its intended outcome, regardless of how it gets there. As a result, it is extremely rare for a single metric to capture everything you care about.