Evaluating AI Systems

you own this product
Assessing effectiveness, reliability, and safety
  • MEAP began October 2026
  • Last updated October 2026
  • Publication in Spring 2027 (estimated)
  • ISBN 9781633434295
  • 375 pages (estimated)
  • printed in black & white

pro $24.99 per month

  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose one free eBook per month to keep
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime

lite $19.99 per month

  • access to all Manning books, including MEAPs!

team

5, 10 or 20 seats+ for your team - learn more


Look inside
AI systems fail in ways you’ve never seen before. A model scores high on benchmarks and still misleads users. An agent completes a task by taking harmful actions and then covering its tracks. A RAG pipeline retrieves the right documents and still returns a fabricated response. Evaluating AI Systems teaches you how to know—and prove—whether your LLM-based applications are ready to ship. In this practical and important book, author Panos Alexopoulos provides a data and decision-driven approach that starts with the questions you need to answer about your system and works backwards to the evidence you need to gather.

Evaluating AI Systems goes beyond how to gather metrics and use eval libraries. In it, you’ll learn to think systematically about evaluating the performance of AI-led applications and agents. You’ll start with the core issues: how to define your evaluation goals and determine which aspects of a system need to be evaluated to achieve them. Next, you’ll learn how to select meaningful metrics and acquire the appropriate evaluation data, along with the trustworthy judgments you need to calculate them.

Because evals without actions are meaningless, you’ll learn how to interpret evaluation results, communicate their meaning and limitations to different stakeholders, and turn evaluation evidence into decisions and improvements. You’ll see how evaluation continues after deployment, allowing you to monitor how your AI system behaves in the real world and identify when its behavior or performance requires attention.

Along the way, you’ll test the book’s systematic approach against some of the most challenging qualities of modern AI systems, such as evaluating robustness when systems encounter unexpected conditions, assessing whether their behavior is fair across relevant groups, and identifying security and privacy risks. Plus, explore the challenges of evaluating increasingly general and adaptable AI systems whose capabilities and behaviors may be difficult to anticipate in advance.

Throughout the book, examples and case studies spanning predictive, generative, and agentic AI demonstrate how the same evaluation principles can be adapted to different systems and technologies. By the end, you’ll have the principles, frameworks, and practical judgment to build evaluations that produce evidence you can act on.

what's inside

  • A decision-driven evaluation framework
  • Obtaining reliable human and AI-based evaluation judgments
  • Interpreting, communicating, and acting on evaluation results
  • Evaluating and monitoring AI systems in production
  • Generality, adaptability, and emergent behaviors

about the reader

For AI engineers, data scientists, machine learning practitioners, and technical leaders who build, deploy, or oversee AI systems. You will need to know the basics of machine learning and modern AI applications, such as LLMs, RAG, or agents.

about the author

Panos Alexopolous is Lead Semantic Data and AI solutions at Triply BV, in Amsterdam, Netherlands. An AI practitioner, author, and educator with 20 years of industry experience across diverse domains, he helps large organizations design, develop and deploy data management and AI solutions. Panos is the author of author of Semantic Modeling for Data and has designed and delivered over 20 master classes, tutorials, and courses on data and AI, including a highly popular course on Knowledge Graphs and Large Language Models. He regularly speaks at conferences, meetups, and podcasts as an advocate for grounded, responsible, and actionable data and AI practices.
choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • Evaluating AI Systems ebook for free
choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • Evaluating AI Systems ebook for free