The main security risks specific to LLM applications and the layered controls that reduce them, from input handling to tool permissions.
A new kind of attack surface
LLM applications accept natural language, follow instructions and increasingly take actions. That makes them vulnerable to risks that traditional input validation does not cover. The OWASP Top 10 for LLM Applications is a useful reference for the main categories.
Prompt injection
Prompt injection happens when text supplied by a user, or hidden in a document, web page or email the model reads, contains instructions that override the intended behaviour. Indirect injection through retrieved content is especially relevant to RAG and agents. There is no complete fix today, so design on the assumption that it can happen.
Sensitive data exposure
Models can reveal information from their prompt, retrieved documents or connected systems. Apply access control before retrieval so the model only sees what the current user is allowed to see. Avoid putting secrets in prompts, and filter outputs for personal or confidential data where appropriate.
Excessive agency
When an agent can call tools, the risk extends from wrong text to wrong actions. Give each tool the narrowest permissions it needs, validate parameters in code, and require human confirmation for payments, deletions, external messages and other irreversible steps.
Insecure output handling
Treat model output as untrusted input. If it is rendered in a browser, inserted into a database query or passed to a shell, escape and validate it as you would any user input.
Guardrails in layers
Useful layers include input classification to catch clear abuse, clear system instructions, retrieval restricted by permissions, output checks for policy and format, and tool-level authorisation. No single layer is enough; together they reduce both the likelihood and impact of an attack.
Testing and monitoring
Run adversarial tests (red teaming) before launch and after significant changes. Log prompts, retrieved content, tool calls and outputs so incidents can be investigated. Frameworks such as the NIST AI Risk Management Framework help structure ongoing governance.
Key takeaways
- Assume prompt injection is possible and limit what a successful attack can do.
- Enforce access control before retrieval, not after.
- Treat model output as untrusted input.
- Combine several guardrail layers and keep testing.
If you are connecting an AI system to internal data or tools, our integration team can help you design the access controls around it.
AI Integration & Automation- OWASP Top 10 for LLM ApplicationsOWASP GenAI Security Project · Industry standardgenai.owasp.org
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt InjectionGreshake et al. · 2023 · Research paperarxiv.org
- AI Risk Management FrameworkNIST · Frameworknist.gov