chapter one

1 The importance of AI evaluation

 

This chapter covers

  • What AI evaluation means and entails.
  • What are AI systems and their main categories.
  • Why AI evaluation is challenging and harder than traditional software testing.
  • How to structure and execute AI evaluation projects in a systematic way.
“Quality is never an accident. It is always the result of intelligent effort.”

—John Ruskin

In late 2021, real estate giant Zillow stunned the market by announcing the abrupt shutdown of Zillow Offers, its AI-powered home-flipping business, and laying off 2,000 employees. The cause was a machine learning algorithm that predicted home prices too optimistically, leading Zillow to massively overpay for thousands of properties. The company ultimately wrote off over $300 million in unsellable inventory and its own CEO admitted that the algorithm could not be reliably fixed, calling the risk “too high” [1].

Two years later, another AI system made headlines, this time for discriminating against job applicants based on age. In August 2023, the tutoring company iTutor Group agreed to pay $365,000 to settle a lawsuit brought by the U.S. Equal Employment Opportunity Commission. According to the commission, the company’s AI-powered recruiting software automatically rejected female applicants over 55 and male applicants over 60, screening out more than 200 qualified candidates before the pattern was discovered [2].

1.1 What is AI evaluation

1.2 What is an AI system

1.3 Why AI evaluation has become critical

1.4 Why AI evaluation is challenging

1.5 A systematic approach to AI evaluation

1.6 The codex and the compass

1.7 References

1.8 Summary