AI security is the practice of protecting LLM applications and AI agents from attacks that exploit how models read, reason over and act on untrusted text. A language model cannot reliably tell instructions from data, so every LLM app must be engineered on the assumption that the model will sometimes be manipulated, and what it can see and do decides how much damage follows.
Why AI demos fail in enterprise production lists prompt injection and permission leakage as two failure modes among twelve. This article goes deeper and treats LLM security as its own engineering discipline.
Why LLM security is different
Traditional application security separates code from data. SQL injection was solved, in principle, by parameterised queries that keep user input out of the instruction channel. LLMs have no such separation: the system prompt, the user's question, a retrieved PDF and a tool result arrive as one stream of tokens, and the model decides what counts as an instruction. Three consequences follow:
- You cannot filter your way to safety. Classifiers reduce noise, but natural language has unlimited ways to phrase an instruction. They are a signal, not a boundary.
- Permissions set the blast radius, not prompts. A manipulated chatbot over public documents is low risk. An agent with write access to email, tickets and payments is high risk however good its prompt.
- Output is untrusted input to the next system. Whatever the model produces must be validated before it reaches a browser, database, shell or another agent.
Securing AI agents is therefore mostly classic security engineering applied to a capable but easily persuaded component.
A threat model for LLM apps and AI agents
Start every review by drawing the system and marking where trust changes. This diagram covers a RAG assistant and a tool-calling agent.
[ User ] (authenticated via IdP)
|
==== TB1: user input (untrusted) ==========
v
+-------------------+
| App / API layer |-- logs, traces
+-------------------+
| |
| ==== TB2: retrieved content =====
| v
| [ Vector index / docs / email ]
| (may contain hostile text)
v
+-------------------+
| LLM (provider) |
+-------------------+
| tool call request
==== TB3: model output -> actions =========
v
+-------------------+
| Tool layer / MCP |-- policy, approval
+-------------------+
| scoped, on-behalf-of token
==== TB4: enterprise systems ==============
v
[ CRM ] [ Tickets ] [ Email ] [ Payments ]
Assets: confidential documents and their access rules, customer PII, credentials and tokens, the system prompt, the integrity of actions in downstream systems, the model budget and audit records.
Actors: users probing limits; malicious insiders; outsiders who never log in but can place text the system will read (an email, ticket or web page); compromised plugins or MCP servers; and the model acting on manipulated context.
Trust boundaries: TB1 is the familiar one. TB2 is the one teams forget: retrieved content is attacker-controllable whenever someone outside your perimeter can write to a source you index. TB3 is where a probabilistic suggestion becomes a real action. TB4 is where enterprise authorisation must be enforced, never delegated to the model.
Generative AI security risks: the main classes
The OWASP Top 10 for LLM Applications is a useful community reference for naming and prioritising these risks. The classes below overlap with it, framed for enterprise builds.
1. Direct and indirect prompt injection
What it is. Direct injection is a user typing instructions meant to override the system's rules. Indirect injection hides those instructions in content the system processes: a knowledge-base document, an inbound email, a web page an agent browses, a code comment. The user may be innocent; the attacker never touches your app.
Illustrative impact. An assistant summarising a supplier's email is steered into forwarding an internal thread, marking an invoice approved or adding a phishing link to its summary.
Mitigations. Assume injection will succeed sometimes and limit what a manipulated model can do: least-privilege tools, actions under the user's identity, human approval for writes, and no single context that mixes untrusted content with high-privilege tools. Tag untrusted content with delimiters and provenance. Red-team with hostile content in documents and emails, not just the chat box.
2. Sensitive data disclosure and PII leakage
What it is. The system reveals data the requester should not see: another customer's record, a restricted HR file, PII sent to a provider without approval, or secrets captured in logs.
Illustrative impact. A hospital's clinical-notes assistant gives a scheduler diagnosis details, or patient identifiers land in a tracing tool a wide engineering team can read.
Mitigations. Enforce access control at retrieval, not in the prompt. Minimise what enters the context. Redact PII before it reaches the model where the use case allows, and before logging. Confirm provider retention and training-use terms; use private endpoints and approved regions. Treat traces as sensitive data (see AI observability).
3. Insecure output handling
What it is. Model output goes to an interpreter unchecked: rendered as HTML, concatenated into SQL, executed as code or used as a URL the server fetches.
Illustrative impact. A chat UI renders raw HTML, so an injected instruction makes the model emit script that runs in the user's browser. A text-to-SQL feature produces a destructive statement.
Mitigations. Treat output like user input. Encode it for its destination; allow-list link and image domains. Run text-to-SQL on a read-only role over restricted views, allow-list statement types and add row limits and timeouts. Prefer schema-validated structured output. Never pass model output to eval or a shell outside a sandbox.
4. Excessive agency and over-permissioned tools
What it is. An agent has more functions, permissions or autonomy than its task needs: a generic "call any API" tool, an admin service account, or freedom to send, delete and pay without review.
Illustrative impact. An IT-ops agent built to restart a stuck service can also delete virtual machines because its cloud role was broad. One manipulated step becomes an outage.
Mitigations. Build narrow, typed, task-specific tools instead of pass-through ones. Separate read and write tools and gate writes. Use on-behalf-of authentication so the agent never exceeds the signed-in user's rights, with tokens scoped to specific resources. Cap iterations, tool calls and spend per task. What is MCP shows how tool definitions shape what an agent can attempt.
5. Supply chain: models, plugins and MCP servers
What it is. Third-party components are compromised, malicious or poorly built: model weights from an unverified source in a format that can execute code on load, a package in your agent framework, a plugin, or a community MCP server whose tool descriptions carry hidden instructions or whose update quietly changes behaviour.
Illustrative impact. A developer installs a convenient file-system MCP server. A later version adds a tool description nudging the model to read credential files and pass them to another tool.
Mitigations. Inventory models, datasets, packages and MCP servers with owners. Pin versions, verify hashes and prefer safe weight formats. Review MCP server code and tool descriptions before approval, run servers with minimal permissions in isolated containers and re-review on update. Scan dependencies in CI, as taught in our DevSecOps course. MCP vs API covers when a capability should be an MCP server at all.
6. Training and RAG data poisoning
What it is. Someone inserts or edits data so the system learns or retrieves something false or hostile. For most enterprises the realistic vector is RAG: anyone who can edit a wiki, upload to a shared drive or file a ticket can influence retrieval.
Illustrative impact. An insurer's claims assistant indexes a shared folder; someone adds a document with an outdated, more generous coverage rule, and staff start quoting it to customers.
Mitigations. Index only governed sources with owners. Record provenance, author and version per chunk and show citations. Separate authoritative from user-generated content in ranking. Keep an evaluation set that catches changed answers on critical questions.
7. Denial of wallet and cost abuse
What it is. Attackers or runaway loops burn model capacity: huge inputs, prompts that force long outputs, agents looping on tools, or a public endpoint used as a free LLM proxy.
Illustrative impact. A retailer's public shopping assistant is scripted by a third party to answer unrelated questions at scale, and the bill arrives before anyone notices.
Mitigations. Authenticate where possible. Set per-user and per-tenant rate limits and token quotas, cap input size, output tokens and agent iterations and alert on spend anomalies daily.
8. System-prompt leakage
What it is. Users extract the system prompt and anything embedded in it.
Illustrative impact. Serious only if it contains API keys, internal URLs or admin lists.
Mitigations. Assume the prompt is public. Keep secrets and authorisation logic out of it and enforce rules in code.
AI security controls by layer
No single control is enough; defence comes from layers.
| Control | What to implement | Main risks reduced |
|---|---|---|
| Identity and least privilege | Enterprise IdP sign-in; on-behalf-of token exchange so tools act as the user; short-lived, resource-scoped tokens; a separate workload identity per agent | Excessive agency, data disclosure, injection impact |
| Input filtering | Size limits, injection classifiers as signals, PII detection, provenance tags | Prompt injection, PII leakage, cost abuse |
| Output filtering | Schema validation, context-aware encoding, domain allow-lists, PII and secret scanning | Insecure output handling, exfiltration |
| Retrieval access control | Source ACLs as chunk metadata, filtered by user identity before ranking; permission tests in CI | Sensitive data disclosure |
| Human approval for writes | Approval screen showing the exact action and parameters for send, update, delete, pay; deterministic policy checks | Excessive agency, indirect injection |
| Sandboxing | Code execution and browsing in isolated containers without credentials, with restricted egress; MCP servers isolated by trust level | Insecure output handling, supply chain |
| Secrets management | Secrets in a vault or cloud secrets manager, injected at the tool layer, never in prompts or logs; rotation | Prompt leakage, supply chain |
| Logging and audit | Trace requests, retrievals, tool calls and approvals with user identity; redaction; immutable action audit | All (detection, forensics) |
| Rate limits and budgets | Per-user and per-tenant quotas, iteration caps, spend alerts | Denial of wallet |
| Red-teaming | Adversarial suites for direct and indirect injection, leakage and tool misuse, before launch and on major changes | All (verification) |
Identity is where most teams cut corners. With Microsoft Entra ID, an on-behalf-of flow lets your API exchange the user's token for a downstream token carrying the user's own permissions, so a SharePoint or Graph call made by the agent can never reach more than the user could. The Entra ID course covers app registrations, scopes and these flows.
Want to build these controls hands-on? Cloudsoft's AI, GenAI and Agentic AI course takes you from LLM fundamentals to agents that can survive a security review.
Illustrative scenario: containing injection in an email-triage agent
Consider a bank's operations team in a Hyderabad GCC handling a heavy flow of corporate-client email. They want an agent that classifies each email, extracts account references, drafts a reply and creates a service-desk ticket. Anyone on the internet can send this agent text, so every email crosses TB2. Suppose one hides an instruction to look up a different client's balance and include it in the reply, or to mark a payment-change request as verified.
The secure design does not depend on the model resisting that text. It contains the damage:
- Split privileges across steps. A reader step with no tools extracts category, account reference and requested action into a strict schema. Code then checks the account reference format and that it belongs to the sender's verified domain.
- Narrow, scoped tools. The drafting step can only call
get_ticket_history(account_id)for the validated account andcreate_draft_reply(ticket_id, body). No search, no send, no other clients' data; tokens are scoped to the operations queue. - Humans approve writes. Drafts go to an analyst's queue. Payment-instruction changes are never actioned by the agent; they follow the existing call-back verification process.
- Output checks. Code scans each draft for account numbers not belonging to the ticket's client, links outside an allow-list and PII patterns.
- Full tracing. Email, extracted schema, tool calls, draft and analyst decision share one trace ID. Flagged emails are sampled for review.
Replay the attack. The fooled reader has no tools; the drafter can only reach the validated account; a strange draft is caught by output checks, then by a human. The injection "worked" on the model but had nowhere to go. That is the goal of LLM security design.
AI security review checklist
- A threat model exists with assets, actors and trust boundaries, including every retrieved or tool-returned content source.
- Users authenticate through the enterprise identity provider; no anonymous access to internal data.
- Tools act on behalf of the user with scoped, short-lived tokens; no super-user service accounts.
- Each tool is narrow and typed; read and write tools are separate.
- Every write, send, delete or payment needs human approval or a deterministic policy check.
- Retrieval filters by user identity using source ACLs, with permission tests in CI.
- Untrusted content is provenance-tagged and never shares a context with high-privilege tools without a control between them.
- Model output is schema-validated and encoded for its destination; no raw HTML, SQL or shell.
- Code execution and browsing are sandboxed, without credentials, with restricted egress.
- No secrets or authorisation logic in prompts; secrets live in a secrets manager.
- Models, packages, plugins and MCP servers are inventoried, pinned, reviewed and scanned.
- PII is minimised and redacted in prompts, logs and traces per data classification.
- Rate limits, token quotas, iteration caps and spend alerts are live.
- A red-team suite for injection, leakage and tool misuse has run, with tracked results.
- An incident runbook explains how to disable tools, revoke tokens and roll back a prompt or model fast.
Governance basics: data classification and vendor review
Data classification. Map your existing classes (for example public, internal, confidential, restricted) to which model endpoints may receive them, what may be indexed and what may be logged. Some data should never enter an LLM pipeline at all. In India, align this with the Digital Personal Data Protection Act and your sector regulator, with privacy and compliance teams involved early.
Vendor and model review. For each provider, record retention and training-use terms, hosting regions, encryption, provider staff access, sub-processors, incident notification and data deletion. Apply the same review to MCP servers and plugins: they are code running with your permissions.
Inventory and change control. Keep a register of AI use cases with business and technical owners, data classes and the tools each agent can call; a fast approved path discourages shadow AI. Treat prompts, tool definitions, retrieval settings and model versions as production configuration: version them, review changes and rerun red-team and LLM evaluation suites before release.
Skills behind AI security
AI security spans application security, cloud security and AI engineering. The fundamentals come from a cyber security course, IAM and network isolation from cloud security training, and pipeline scanning from DevSecOps. Carrying all of this into a customer's environment and taking an agent to security sign-off is what Forward Deployed Engineers do; the Secure Banking AI Assistant project in FDE PRO is built around these controls.
FAQ
What is AI security?
AI security is the discipline of protecting AI systems, especially LLM applications and agents, from attacks and failures such as prompt injection, data leakage, insecure output handling, over-permissioned tools, supply chain compromise and cost abuse. It combines classic controls like identity, least privilege and audit with model-specific ones such as retrieval access control and output validation.
What is the difference between direct and indirect prompt injection?
Direct prompt injection is a user typing instructions meant to override the system's rules. Indirect prompt injection hides those instructions in content the system processes, such as a retrieved document, an inbound email or a web page an agent visits. Indirect injection is more dangerous because the attacker never needs access to your application.
Can prompt injection be fully prevented?
Not with current models. Because an LLM cannot reliably separate instructions from data, filters and careful prompts reduce the success rate but cannot eliminate it. The reliable approach is to limit impact with least-privilege tools, actions under the user's identity, human approval for writes, output validation and sandboxing.
How do you secure AI agents that call tools?
Give each agent narrow, typed tools rather than generic API access, run tool calls on behalf of the signed-in user with scoped short-lived tokens, separate read and write tools, require approval or policy checks for consequential actions, cap iterations and spend, and log every tool call with the user identity.
What is the OWASP Top 10 for LLM Applications?
It is a community-maintained OWASP project that names and describes the most important security risks for applications built on large language models, with examples and mitigations. It is a useful reference for threat modelling and security reviews and is updated periodically.
Are MCP servers a security risk?
They can be. An MCP server is code that exposes tools to an AI application, often with real credentials. A malicious or poorly built server can grant too much access or carry hidden instructions in tool descriptions. Review the code and tool definitions, pin versions, run servers with minimal permissions in isolation and re-review on every update.
Should the system prompt be kept secret?
Design as if it will leak. Keep instructions in the system prompt but never secrets, credentials, internal URLs or authorisation rules, and enforce access and business rules in code so a leaked prompt reveals nothing useful.
Who is responsible for LLM security in an enterprise?
It is shared. Security teams set policy, data classification and vendor review; platform teams provide identity, secrets, isolation and logging; and the engineers building each app or agent own its threat model, tool permissions, output handling and red-team tests. Each use case also needs a named business owner.
If you want to design LLM apps and agents that pass a real security review, not just a demo, Cloudsoft's Agentic AI training covers RAG, tool-calling agents, MCP and guardrails through hands-on labs, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.



