chapter two

2 Framing and scoping the evaluation

 

This chapter covers

  • Setting evaluation goals that answer meaningful questions and guide decisions.
  • Determining what components of an AI system you need to evaluate.
  • Defining the tasks to be evaluated and what success looks like.
  • Selecting the evaluation dimensions that are most important for your evaluation.
“I wisely started with a map.”

—J.R.R. Tolkien

Before you can measure anything about your AI system, you must decide what you are trying to learn from your evaluation. This step sounds obvious, yet it is where many AI evaluations fail. The reason is that teams often rush to metrics, benchmarks, or toolkits, without having agreed on a common direction and without having made their underlying assumptions and expectations explicit.

Framing and scoping an evaluation is about turning a vague question like “Is this system any good?” into a clear, shared understanding of what success means, for whom, and under which assumptions. It requires you, your team, and your stakeholders make deliberate choices about what is being evaluated, why it is being evaluated, which tasks and scenarios matter, and which dimensions of performance should be assessed. These decisions shape everything that follows: the metrics you choose, the data you collect, and the evaluation methods and tools you develop.

2.1 One room, many expectations

2.2 Defining evaluation goals

2.2.1 Setting expectations

2.2.2 Comparing

2.2.3 Deciding go/no-go

2.2.4 Assessing risk

2.2.5 Ensuring compliance

2.2.6 Diagnosing and debugging

2.2.7 Negotiating and prioritizing evaluation goals

2.3 Determining the objects of the evaluation

2.4 Defining the task to be evaluated

2.4.1 Problems with AI task definitions

2.4.2 Defining an AI task well

2.5 Choosing evaluation dimensions

2.6 References

2.7 Summary