Overview

1 The importance of AI evaluation

AI evaluation is presented as the disciplined practice of understanding how a system behaves in the real world, not just how well a model scores on benchmarks. The chapter opens with examples of costly failures in housing, hiring, and customer service to show that AI can cause serious harm when its behavior is not adequately assessed before and after deployment. Its central message is that quality in AI does not happen by accident; it comes from deliberate, intelligent effort to anticipate risk and measure performance in context.

The text defines AI systems broadly as software designed to perform tasks associated with human intelligence, while stressing an important distinction between a model and the full system around it. It also groups AI systems into three broad categories: predictive systems that forecast or classify, generative systems that create content, and agentic systems that take actions toward goals. Because real-world products often combine these categories, evaluation should focus on what the complete system is meant to do, how it is used, and whether its behavior is acceptable for that setting.

The chapter argues that AI evaluation has become critical because modern systems are more capable, more complex, and more consequential than traditional software, yet teams often understand them only partially. This creates an evaluation gap that can lead to both over-reliance and under-reliance on AI, while shallow metric-based testing misses issues such as fairness, robustness, explainability, security, and ethical safety. To address this, the chapter outlines a systematic evaluation workflow: define the goal and scope, choose meaningful metrics, gather suitable data, compute results carefully, interpret them in context, and use the findings to decide what to do next, ideally as an iterative and continuous process throughout development and deployment.

The difference between an AI model and an AI system. A model is a self-contained computational component that transforms inputs into outputs, while an AI system combines the model with s user interfaces, data pipelines, integration logic, human-in-the-loop mechanisms, and other functionalities and features to perform real-world tasks.
Three broad categories of AI systems: predictive systems that make forecasts or classifications, generative systems that create content, and agentic systems that take actions to achieve goals in an environment.
The six main steps of an AI evaluation project, typically executed in an iterative fashion. All six are essential, and their order matters.

Summary

  • AI capabilities have surged while evaluation practices lag, and high-profile failures show the cost of that gap.
  • A widening capability–comprehension gap creates two opposing risks: over-reliance and under-reliance.
  • AI evaluation means systematically assessing how systems perform under real-world complexity and unpredictability.
  • We should be evaluating AI systems, not just models, as the real-world impact of AI emerges from interfaces, data pipelines, human interaction, and policies working together.
  • AI Evaluation is difficult: unlike traditional software, AI is probabilistic, data-dependent, and prone to edge-case breakdowns and unforeseen behaviors.
  • Effective evaluation is deliberate and structured, and every project should iterate through six core activities:
    • framing and scoping
    • choosing metrics
    • assembling evaluation data
    • running and computing
    • interpreting and communicating
    • deciding and acting
  • Much of AI evaluation can be automated, but the most critical decisions, such as what to measure, how to weigh trade-offs, and what actions to take, still require human judgment.

FAQ

```html
What does AI evaluation mean in this chapter?AI evaluation is the systematic practice of understanding and measuring how an AI system behaves in its intended context of use. It goes beyond benchmark scores or accuracy metrics to ask whether the system reliably achieves its purpose in real-world conditions.
Why is evaluating an AI system different from evaluating traditional software?Unlike traditional software, AI systems are probabilistic, data-driven, and often behave differently across users, inputs, and deployment contexts. They can fail in unexpected ways, exhibit emergent behavior, and require evaluation across multiple dimensions such as fairness, robustness, and explainability.
What is the difference between an AI model and an AI system?An AI model is the trained computational component that produces predictions or outputs. An AI system includes the full product around the model, such as interfaces, data pipelines, integration logic, guardrails, and human oversight. Evaluation should focus on the whole system, not just the model.
What are the main categories of AI systems discussed in the chapter?The chapter describes three broad categories: predictive AI systems, which forecast, classify, or rank; generative AI systems, which create content such as text or images; and agentic AI systems, which can take actions and pursue goals in an environment.
Why has AI evaluation become more critical now?AI systems are being used in high-stakes domains like healthcare, law, hiring, and customer service, where failures can have serious consequences. At the same time, modern systems are more complex, less predictable, and harder to understand, making rigorous evaluation essential.
What makes AI evaluation especially challenging?Key challenges include context-dependent performance, poor or biased evaluation data, missing ground truth, open-ended outputs, human-AI interaction effects, emergent behavior, opacity, and the need to balance conflicting values such as accuracy, fairness, and efficiency.
Why is evaluation data so important?Metrics are only as good as the data they are based on. If evaluation data is outdated, biased, unrepresentative, or poorly aligned with the real use case, the results can give a false sense of confidence and miss critical failures.
What is meant by the “AI evaluation gap”?The AI evaluation gap is the mismatch between how capable modern AI systems are and how limited current evaluation practices often are. Many organizations still rely on narrow benchmark tests or one-time checks instead of thorough, context-aware evaluation.
What are the six steps in a systematic AI evaluation project?The six steps are: define what and why you are evaluating, select appropriate metrics, acquire suitable evaluation data, compute the metrics, interpret and communicate the results, and make a decision about next steps such as deployment or further testing.
Is AI evaluation a one-time task before deployment?No. The chapter emphasizes that evaluation should be iterative and continuous throughout the AI lifecycle. It should happen both offline before deployment and online after deployment, because data, user needs, and system behavior change over time.
```

pro $24.99 per month

  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose one free eBook per month to keep
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime

lite $19.99 per month

  • access to all Manning books, including MEAPs!

team

5, 10 or 20 seats+ for your team - learn more


choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • Evaluating AI Systems ebook for free
choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • Evaluating AI Systems ebook for free