Overview

4 Securing GenAI

The chapter explains that GenAI systems blur the line between instructions and data in a way traditional software does not, which makes prompt injection, exfiltration, and other abuses hard to prevent with ordinary validation alone. Because language models interpret all text in context, attackers can hide commands inside emails, documents, code, or other content and cause the model to follow them instead of the user’s intent. The result is that defenses must rely on layered guardrails, monitoring, and containment rather than assuming a clean technical boundary between what the model should do and what it should merely process.

Security risk rises as a system gains more capability. The chapter uses a staged insurance assistant to show how adding data access, model customization, write access, and autonomy each expands the attack surface and the possible damage. With access to live records, attackers can exfiltrate sensitive information; with write permissions, they can change records or redirect payments; and with agentic behavior, they can manipulate multi-step workflows, exploit trust between tools, and trigger cascading failures. The central message is that each new capability requires matching controls, because a simple chatbot and a fully autonomous agent face very different levels of exposure.

To manage these risks, the chapter recommends matching controls to the specific kind of exposure involved: secure the environment, vet models and supply-chain components, isolate sensitive data, scope credentials tightly, enforce policy outside the model, require human approval for high-impact actions, and sandbox agent execution. It also emphasizes strong logging, red-teaming, rate limits, secret management, tool registries, and continuous monitoring so organizations can detect misuse early and contain failures before they spread. Overall, the chapter presents GenAI governance as a practical discipline of aligning permissions, oversight, and technical boundaries with the real-world consequences of each system feature.

ClaimAssist v1 on the left is a simple policy chatbot. ClaimAssist v4 on the right can accept user uploaded images and change a claim status as “ready for adjuster review”.
Security Risk Factors. Each additional factor compounds the risk.
Real life prompt injection attack
Image prompt injection causing a faulty decision in ChatGPT-5.

Summary

Generative AI introduces security challenges that traditional software controls were not designed to address. Large language models cannot reliably distinguish between data they should process and instructions they should follow. When you connect these systems to sensitive data, let them take actions, or grant them autonomy to pursue goals, you create attack surfaces that traditional controls cannot protect. Firewalls, access controls, and input validation were not designed for systems that blur the line between data and instruction.

The six risk factors introduced in Section 4.2 determine how much exposure a GenAI system carries. Environment and model risks set the foundation: an insecure underlying platform or a compromised model undermines everything built on top. Untrustworthy input and data security risks compound: prompt injection against a system with data access turns every input manipulation into a potential breach. Action capability and agency take it to its worst form: from “the model said something wrong” to “the model did something wrong at scale, across multiple sessions, without anyone asking it to.”

The Lethal Trifecta (untrusted input, access to sensitive data, and a channel for external communication) is the condition that makes agentic systems most dangerous. Break any one element and the worst attacks fail. But useful AI assistants often need all three, making containment rather than prevention the realistic goal.

The controls at each layer follow a consistent architecture: match permissions to the requesting user’s identity, not a shared service account; enforce authorization decisions outside the model’s reasoning process; require human review proportional to consequence; maintain complete audit trails for actions, not just outputs; contain damage through sandboxing and rate limits; and monitor reasoning traces, not just network traffic.

The chapter leaves two different kinds of security problem. Identity, gateways, external policy enforcement, approval paths, traceability, and containment can be governed now, although none is foolproof. General prompt injection, trust in changing tool ecosystems, and assurance that an autonomous agent will remain within intent under adversarial conditions are not solved. Organizations should use the available controls, collect evidence that they operate, and constrain capabilities when those controls cannot keep the consequences acceptable.

FAQ

Why can’t an LLM reliably tell the difference between instructions and data?LLMs process both as plain text and infer meaning from context, rather than enforcing a technical boundary. That means they may treat malicious content inside data as if it were a legitimate instruction.
What is prompt injection in GenAI systems?Prompt injection is an attack where crafted input causes the model to ignore its intended task and follow attacker-supplied instructions instead. It can happen in direct chat prompts or indirectly through documents, emails, code, and other retrieved content.
Why are guardrails and input filters not enough to stop prompt injection?Guardrails can block many obvious attacks, but they are still filters, not hard boundaries. Attackers can bypass them with rephrasing, hidden text, encoding tricks, or multi-step attacks, so they must be paired with monitoring and containment controls.
What is the “risk ladder” described in this chapter?The risk ladder is a framework for understanding how GenAI exposure grows as systems gain more capabilities: environment, model, input, data access, ability to make changes, and agency. Each added factor increases attack surface and requires stronger controls.
How does data access increase GenAI security risk?Once an assistant can retrieve sensitive runtime data, a successful prompt injection can turn into a data breach. Instead of just generating bad output, the model may leak customer records, policies, source code, or other confidential information.
Why is write access more dangerous than read-only access?Read-only mistakes are usually limited to incorrect or leaked information, but write access can change real systems. A compromised assistant can update records, redirect payments, send emails, or corrupt data, which can create irreversible harm.
What is the confused deputy problem in AI assistants?It happens when an assistant uses a high-privilege shared account and accidentally gives users more access than they should have. The assistant becomes “confused” and performs actions or retrieves data on behalf of someone who should not be authorized.
Why is agentic AI a bigger security concern than a simple chatbot?Agentic systems can plan multi-step actions, choose tools, and pursue goals on their own. That autonomy creates new risks such as tool misuse, memory poisoning, cascading failures, and rogue behavior across sessions or agents.
What controls should organizations use for AI agents that can act on their own?Use external policy engines, tiered human approval, short-lived credentials, sandboxing, detailed logging, behavioral monitoring, and kill switches. These controls help limit damage even if the agent is manipulated or makes a bad decision.
What is the best way to reduce risk from shadow AI?Provide sanctioned tools that are easy to use, route traffic through monitored gateways, enforce SSO and data loss prevention, and educate employees about approved GenAI usage. If safe, convenient options exist, employees are less likely to use unsanctioned personal accounts.

pro $24.99 per month

  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose one free eBook per month to keep
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime

lite $19.99 per month

  • access to all Manning books, including MEAPs!

team

5, 10 or 20 seats+ for your team - learn more


choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • AI Governance ebook for free
choose your plan

team

monthly
annual
$49.99
$499.99
only $41.67 per month
  • five seats for your team
  • access to all Manning books, MEAPs, liveVideos, liveProjects, and audiobooks!
  • choose another free product every time you renew
  • choose twelve free products per year
  • exclusive 50% discount on all purchases
  • renews monthly, pause or cancel renewal anytime
  • renews annually, pause or cancel renewal anytime
  • AI Governance ebook for free