Manning Early Access Program (MEAP)
Read chapters as they are written, get the finished eBook as soon as it’s ready, and receive the pBook long before it's in bookstores.
AI systems fail in ways you’ve never seen before. A model scores high on benchmarks and still misleads users. An agent completes a task by taking harmful actions and then covering its tracks. A RAG pipeline retrieves the right documents and still returns a fabricated response. Evaluating AI Systems teaches you how to know—and prove—whether your LLM-based applications are ready to ship. In this practical and important book, author Panos Alexopoulos provides a data and decision-driven approach that starts with the questions you need to answer about your system and works backwards to the evidence you need to gather.
Evaluating AI Systems goes beyond how to gather metrics and use eval libraries. In it, you’ll learn to think systematically about evaluating the performance of AI-led applications and agents. You’ll start with the core issues: how to define your evaluation goals and determine which aspects of a system need to be evaluated to achieve them. Next, you’ll learn how to select meaningful metrics and acquire the appropriate evaluation data, along with the trustworthy judgments you need to calculate them.
Because evals without actions are meaningless, you’ll learn how to interpret evaluation results, communicate their meaning and limitations to different stakeholders, and turn evaluation evidence into decisions and improvements. You’ll see how evaluation continues after deployment, allowing you to monitor how your AI system behaves in the real world and identify when its behavior or performance requires attention.
Along the way, you’ll test the book’s systematic approach against some of the most challenging qualities of modern AI systems, such as evaluating robustness when systems encounter unexpected conditions, assessing whether their behavior is fair across relevant groups, and identifying security and privacy risks. Plus, explore the challenges of evaluating increasingly general and adaptable AI systems whose capabilities and behaviors may be difficult to anticipate in advance.
Throughout the book, examples and case studies spanning predictive, generative, and agentic AI demonstrate how the same evaluation principles can be adapted to different systems and technologies. By the end, you’ll have the principles, frameworks, and practical judgment to build evaluations that produce evidence you can act on.
what's inside
A decision-driven evaluation framework
Obtaining reliable human and AI-based evaluation judgments
Interpreting, communicating, and acting on evaluation results
Evaluating and monitoring AI systems in production
Generality, adaptability, and emergent behaviors
about the reader
For AI engineers, data scientists, machine learning practitioners, and technical leaders who build, deploy, or oversee AI systems. You will need to know the basics of machine learning and modern AI applications, such as LLMs, RAG, or agents.
about the author
Panos Alexopolous is Lead Semantic Data and AI solutions at Triply BV, in Amsterdam, Netherlands. An AI practitioner, author, and educator with 20 years of industry experience across diverse domains, he helps large organizations design, develop and deploy data management and AI solutions. Panos is the author of author of Semantic Modeling for Data and has designed and delivered over 20 master classes, tutorials, and courses on data and AI, including a highly popular course on Knowledge Graphs and Large Language Models. He regularly speaks at conferences, meetups, and podcasts as an advocate for grounded, responsible, and actionable data and AI practices.
Introductory offer Save 50% for a limited time!
eBook
pdf, ePub, online
$47.99
$23.99
you save $24.00 (50%)
Introductory offer Save 50% for a limited time!
print
includes eBook
$59.99
$29.99
you save $30.00 (50%)
with subscription
free or 50% off
$24.99
pro $24.99 per month
access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!