The agent security checklist: twenty questions before production
AI agent security checklist: twenty yes/no questions on identity, permissions, data flow, injection, sandboxing, supply chain and audit, each with a link.
This post is also available in Deutsch.
Before an agent goes to production, answer twenty questions. If any answer is “no” or “we do not know”, that is the work item. This AI agent security checklist condenses the security posts of this blog into seven areas: identity, permissions, data flow, injection, sandboxing, supply chain and audit. It is a threat model in question form, not a certification, and it will not replace a review by someone who knows your environment.
How to use it: print the list, answer each question per agent (not per platform), and write the evidence next to the answer, such as a config file, a policy rule or a log query. “Probably” is a “no”.
1. Identity
- Does every agent run under its own identity? Shared service accounts make it impossible to tell which agent did what. One identity per agent, and per environment.
- Is it clear on whose behalf the agent acts? An agent acting for a user should carry that user’s authority at most, never more. Record both the agent and the user in every call.
- Are tokens short-lived and scoped to one audience? A token issued for one server must not work against another. The MCP specification is blunt about this:
MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server.
Source: Security Best Practices, MCP 2025-06-18. The same page also says how a session must not be used:
MCP Servers MUST NOT use sessions for authentication.
The details of the authorization flow are in MCP authorization with OAuth.
2. Permissions
- Does each agent hold only the tools its step needs? Split broad agents into small ones; see least privilege for agents.
- Are tool arguments constrained, not just tool names? “May call
add_comment” is weak; “may add one comment of at most 2000 characters to tickets matchingSEC-<number>” is a rule you can test. - Do irreversible or external actions require approval? Sending mail, deleting data, moving money and publishing should pause for a person with the right role.
3. Data flow
- Can you draw where data enters and leaves each agent? Sources, sinks and the trust level of each. If you cannot draw it, you cannot reason about leaks.
- Are secrets kept out of prompts and out of tool results? Credentials belong in the tool layer, not in the model’s context. The MCP spec also forbids using elicitation to ask users for sensitive data (see MCP sampling, elicitation and tasks).
- Is there a rule for what may leave the organisation? Outbound domains, attachments and logs sent to third parties need an explicit allowlist.
4. Injection
- Is all external text treated as untrusted? Web pages, tickets, emails and tool results can contain instructions. The model cannot reliably tell data from commands.
- Does any single agent combine private data, untrusted content and an outbound channel? That combination is the pattern described in the lethal trifecta. Remove one leg.
- Does a deterministic gate sit between model output and side effects? The model proposes, policy decides; see policy decides, audit proves. Claude Code’s documentation says the same in plain words:
Servers that fetch external content can expose you to prompt injection risk.
Source: Connect Claude Code to tools via MCP (live documentation, wording as fetched on 2026-10-04).
5. Sandboxing
- Does code execution happen in a sandbox with no ambient credentials? No home directory, no cloud metadata endpoint, no inherited environment variables.
- Is network egress denied by default? Allow named hosts only. See sandboxing agents.
- Are CPU, memory, time and spend limited per run? A looping agent should hit a ceiling before it hits your invoice.
6. Supply chain
- Do you know every MCP server and skill an agent can load, and who owns it? The same docs give the rule in one line:
Verify you trust each server before connecting it.
- Are versions pinned and changes reviewed? A server or skill that changes silently is a new, unreviewed dependency.
- Do you know which model and provider handle which data? Include the region, retention and the fallback model if the primary fails.
7. Audit
- Is every tool call recorded with arguments, decision and acting identity? A plain log is a start; an append-only, tamper-evident record is better.
- Can you stop an agent and reconstruct what it did? A kill switch and a replayable record are the minimum for incident response.
Mapping to the OWASP list
The checklist is not a copy of any standard. If you need to map it, the OWASP Top 10 for Agentic Applications for 2026 is the natural reference; it describes itself like this:
identifies the most critical security risks facing autonomous and agentic AI systems.
A walk-through of the list is in the OWASP agentic top 10. For terminology on attacks and mitigations, NIST’s adversarial machine learning taxonomy is a useful vocabulary:
provides a taxonomy of concepts and defines terminology in the field of adversarial machine learning (AML).
An example of a policy decision
In the current demo, a policy decision shows the requested tool, its arguments, the rule that matched and the outcome, for example “require approval” for a write to an example ticket. That is the shape question 12 asks for: a visible rule, a visible decision, no model discretion.

Screenshot of the current demo (fake data).
Scoring and next steps
Count the answers you can back with evidence. Fewer than fifteen means the agent is a pilot, not a production service. Fix questions 6, 11 and 12 first: they limit the damage of the failures you cannot predict. Then repeat the review when tools, skills or models change; a checklist answered once is a snapshot.
Key takeaways
- Twenty questions in seven areas: identity, permissions, data flow, injection, sandboxing, supply chain, audit.
- Answer per agent and keep evidence; “probably” counts as “no”.
- The most valuable controls are structural: remove a leg of the trifecta, gate side effects, deny egress.
- Re-run the checklist whenever a tool, skill, server or model changes.
Sources
- OWASP Gen AI Security Project, OWASP Top 10 for Agentic Applications for 2026, published 2025-12-09.
- Model Context Protocol, Security Best Practices (2025-06-18).
- NIST, AI 100-2 E2025, Adversarial Machine Learning, published 2025-03-24.
- Anthropic, Connect Claude Code to tools via MCP, live documentation.