welcome
Thank you for purchasing the MEAP edition of Evaluating AI Systems. I’m delighted to have you as an early reader, and I hope you’ll find the material useful in your work.
AI evaluation has always been important, but modern AI systems have made it both more critical and more difficult. With generative and agentic AI, we are increasingly building systems whose behavior is probabilistic, context-dependent, and difficult to predict. Often there is no simple ground truth against which we can compare their outputs. And even when we can measure performance, a good score doesn't necessarily tell us whether a system is reliable, safe, fair, useful, or ultimately fit for its intended purpose.
This book is my attempt to provide a practical and systematic way of dealing with these challenges.
Rather than treating evaluation as a collection of metrics, benchmarks, or tools, we will approach it as a process for generating evidence about AI systems and using that evidence to make better decisions. You will learn how to move from vague questions such as "Is this system good enough?" to concrete evaluation goals; decide what aspects of a system need to be evaluated; choose meaningful metrics and evaluation data; obtain trustworthy judgments from humans and AI; and interpret the resulting evidence in the context of the decisions you need to make.