4 Securing GenAI
The chapter explains that GenAI systems blur the line between instructions and data in a way traditional software does not, which makes prompt injection, exfiltration, and other abuses hard to prevent with ordinary validation alone. Because language models interpret all text in context, attackers can hide commands inside emails, documents, code, or other content and cause the model to follow them instead of the user’s intent. The result is that defenses must rely on layered guardrails, monitoring, and containment rather than assuming a clean technical boundary between what the model should do and what it should merely process.
Security risk rises as a system gains more capability. The chapter uses a staged insurance assistant to show how adding data access, model customization, write access, and autonomy each expands the attack surface and the possible damage. With access to live records, attackers can exfiltrate sensitive information; with write permissions, they can change records or redirect payments; and with agentic behavior, they can manipulate multi-step workflows, exploit trust between tools, and trigger cascading failures. The central message is that each new capability requires matching controls, because a simple chatbot and a fully autonomous agent face very different levels of exposure.
To manage these risks, the chapter recommends matching controls to the specific kind of exposure involved: secure the environment, vet models and supply-chain components, isolate sensitive data, scope credentials tightly, enforce policy outside the model, require human approval for high-impact actions, and sandbox agent execution. It also emphasizes strong logging, red-teaming, rate limits, secret management, tool registries, and continuous monitoring so organizations can detect misuse early and contain failures before they spread. Overall, the chapter presents GenAI governance as a practical discipline of aligning permissions, oversight, and technical boundaries with the real-world consequences of each system feature.
ClaimAssist v1 on the left is a simple policy chatbot. ClaimAssist v4 on the right can accept user uploaded images and change a claim status as “ready for adjuster review”.
Security Risk Factors. Each additional factor compounds the risk.
Real life prompt injection attack
Image prompt injection causing a faulty decision in ChatGPT-5.
Summary
Generative AI introduces security challenges that traditional software controls were not designed to address. Large language models cannot reliably distinguish between data they should process and instructions they should follow. When you connect these systems to sensitive data, let them take actions, or grant them autonomy to pursue goals, you create attack surfaces that traditional controls cannot protect. Firewalls, access controls, and input validation were not designed for systems that blur the line between data and instruction.
The six risk factors introduced in Section 4.2 determine how much exposure a GenAI system carries. Environment and model risks set the foundation: an insecure underlying platform or a compromised model undermines everything built on top. Untrustworthy input and data security risks compound: prompt injection against a system with data access turns every input manipulation into a potential breach. Action capability and agency take it to its worst form: from “the model said something wrong” to “the model did something wrong at scale, across multiple sessions, without anyone asking it to.”
The Lethal Trifecta (untrusted input, access to sensitive data, and a channel for external communication) is the condition that makes agentic systems most dangerous. Break any one element and the worst attacks fail. But useful AI assistants often need all three, making containment rather than prevention the realistic goal.
The controls at each layer follow a consistent architecture: match permissions to the requesting user’s identity, not a shared service account; enforce authorization decisions outside the model’s reasoning process; require human review proportional to consequence; maintain complete audit trails for actions, not just outputs; contain damage through sandboxing and rate limits; and monitor reasoning traces, not just network traffic.
The chapter leaves two different kinds of security problem. Identity, gateways, external policy enforcement, approval paths, traceability, and containment can be governed now, although none is foolproof. General prompt injection, trust in changing tool ecosystems, and assurance that an autonomous agent will remain within intent under adversarial conditions are not solved. Organizations should use the available controls, collect evidence that they operate, and constrain capabilities when those controls cannot keep the consequences acceptable.
AI Governance ebook for free