AgentShield

Pydantic AI Security for Pydantic AI Agents, MCP Tools, Guardrails and Human Approval

Pydantic AI gives Python teams typed tools, validated arguments and a clean way to pause a run for approval. Validation is not authorization, though, and Pydantic says so in its own docs. With V2 stable since June 23, 2026 and V1 security fixes promised for only six months after that, most production teams are porting agents right now, which is the moment to decide what enforces the rules around every tool call.

OWASP LLM Top 10 Immutable audit trail Never trains on your data

Direct answer

Pydantic AI security is strong on input shape and weak by default on permission. Tools run without approval unless you set requires_approval or raise ApprovalRequired, Pydantic states that approval is not an authorization boundary against an untrusted client, and ten security advisories were published between June and September 2026, several in the web fetch tool and the UI adapters. AgentShield adds the missing layer outside the process: a per-tool policy on every call, holds on high-impact actions for a named approver, inspection of fetched and retrieved content for injected instructions, and a tamper-evident record of each decision.

Try it live

Watch AgentShield block an attack in real time.

Pick a scenario and drive the inspection lane yourself. No signup needed.

Threat Console
Interactive demo · 0 blocked in this session

Run a request

Runs the live engine on your text. Nothing is stored, no account needed.

Inspection lane

INSPECTING
⌖ untrusted input

Policy trace

High-risk action held for approval

Audit trail

The risk

A fintech team ships a Pydantic AI support agent behind a web chat built on the Vercel AI adapter. It has a refund tool, an account lookup tool and the web fetch tool so it can read help articles. Pydantic validation guarantees the refund amount is a positive decimal and the account ID is a string, and the team marks the refund tool requires_approval. What they missed is that the approval arrives from the same browser the customer controls, and Pydantic's docs are explicit that this is not an authorization boundary. A crafted request approves its own refund, and the only record is an OpenTelemetry span.

How AgentShield handles it

Keep Pydantic AI doing what it does well: typed tool schemas, argument validation, deferred tool requests, capabilities and instrumentation. Put AgentShield on the outbound path. Register the MCP servers and HTTP APIs your agents call behind the AgentShield gateway, write a policy per tool (allow, deny, or hold above a threshold), send held calls to a named approver on your side rather than back to the client that asked, inspect fetched pages and retrieved documents for injected instructions before the model reads them, and record the requester, the tool, the validated arguments and the verdict for every action.

The controls

The controls that secure which tools a Pydantic AI agent may call, with what arguments, and who approves it.

What Pydantic AI enforces for you, and what it leaves to your code

Pydantic AI is the agent framework from the team behind Pydantic, the validation library used across most of the Python ecosystem, and its core promise is type safety: tool arguments and model outputs are validated against Python types before your code sees them. That is a real security benefit. A tool that expects an integer account ID will never receive a string of SQL, and a structured output that fails validation is retried or rejected instead of passed along.

Validation answers whether an argument has the right shape. It does not answer whether this agent, on behalf of this user, should call this tool with these values right now. Those are different questions, and the second one is where agent incidents happen. A refund of 480 dollars is a perfectly valid decimal. So is a refund of 48,000.

ControlWhat Pydantic AI providesWhat is left to you
Tool argument shapeSchema generation and validation from type hints, with retries on failureNothing extra. This is the framework's strength
Which tools may runToolsets, prepare_tools, and per-run tool selectionDeciding and enforcing the allowlist per user, tenant and environment
Human approvalrequires_approval, ApprovalRequired and DeferredToolRequestsWhich tools need it, and making sure the approver is not the requester
AuthorizationRunContext and dependency injection to carry identity into toolsEvery permission check inside every tool, as Pydantic recommends
Prompt injectionHarness guardrails for input, output and tool calls, with regex detectorsDetecting instructions in ordinary language, which the detectors cannot do
AuditOpenTelemetry instrumentation, viewable in Logfire or any OTel backendA decision record your auditor accepts, with verdicts, not just spans

The first row is why teams choose Pydantic AI, and it is worth keeping. The remaining rows are all "the framework gives you a hook, you write the rule", which is the honest design for a library and the reason a separate enforcement layer exists.

Pydantic AI human in the loop and why approval is not authorization

Pydantic AI's deferred tools are a clean design. Mark a tool requires_approval=True and the model's call becomes a pending request instead of an execution. The run ends with a DeferredToolRequests output listing each call with its validated arguments, you approve or deny each one, and you resume the run with a DeferredToolResults object. Denials can carry a message back to the model through ToolDenied. For approval that depends on the arguments, a tool can raise ApprovalRequired conditionally, and MCP servers can be wrapped so their tools require approval too.

Two details decide whether this is a control or a formality. First, it is opt-in: by default a Pydantic AI tool runs without any approval. Second, Pydantic's own documentation carries this warning, and it is the most important sentence on the page: "Approval is not an authorization boundary against an untrusted client." When an agent is served through a UI adapter, the client that sent the request is also the one returning approvals, so a malicious client can approve its own calls. Pydantic's guidance is that approval protects against unwanted model actions, while real authorization has to be enforced inside the tool function itself.

That matches what the advisories show. GHSA-jpr8-2v3g-wgf9, published July 2026, described a case where the UI adapters' message sanitizing could be sidestepped so that "A remote client could use this to have a registered server tool run with arguments it supplied rather than arguments the model produced." It was fixed in 1.107.1 and 2.5.0, and the advisory notes that teams relying on model-request hooks for guardrails were the most affected, because forged calls skip the model turn entirely.

Approval designWho approvesHolds against a hostile client
No requires_approval on the toolNobodyNo. The tool runs when the model calls it
requires_approval, approvals returned by the same web clientThe requesterNo, per Pydantic's own warning
requires_approval, approvals collected by your back officeA staff member in your systemYes, if the tool also checks permission
Hold enforced at a gateway on the outbound callA named approver outside the agent processYes, and it covers every agent calling that tool

Our human approval for AI agents page explains how the gateway hold works, including timeouts and what the approver sees.

Pydantic AI V1 to V2 migration security checklist

Pydantic AI 2.0.0 went stable on June 23, 2026, after betas that started May 20. Pydantic's version policy says: "We'll continue to provide security fixes for V1 for at least 6 months after V2's stable release, so you have time to upgrade your applications." Six months from June 23 is roughly December 23, 2026. After that, a V1 agent has no promise of patches, and the V1 line did receive security backports in September (1.107.6 carried the four fixes from 2.44.0).

Most of the V2 changes are renames, but several touch the controls security teams care about. These are the ones to review, not just recompile.

V2 changeWhy it matters for securityWhat to check after porting
Default end_strategy changed from early to gracefulFunction tools now run alongside a successful output tool instead of being skippedTools with side effects that previously never ran at the end of a turn may now run
DeferredToolCalls renamed DeferredToolRequests, DeferredToolset renamed ExternalToolsetYour approval flow is built on these typesEvery approval path still pauses, and none silently executes
Per-transport MCP server classes replaced by MCPToolset, mcp_servers moved to toolsetsMCP wiring is rebuilt, not renamedEach MCP server still has the same tool allowlist and approval wrapping
sequential=True became a per-tool barrierOrdering assumptions between tools can changeAny check that relied on one tool finishing before another
Instrumentation format version 5, approvals no longer recorded as span errorsDashboards and alerts built on span errors go quietAlerts on held or denied calls still fire

The first row deserves a test of its own. Under V1's early strategy, a model that produced its final output and a tool call in the same response often never ran that tool. Under V2's graceful strategy it does. For a read tool that is harmless. For a tool that sends email or moves money, it is a behavior change you want to see in staging, not in production.

Pydantic AI security advisories from June to September 2026

Pydantic publishes advisories promptly and fixes them fast, which is to its credit. Ten appeared on the project's GitHub security page between June and September 2026. None is a flaw in the core agent loop; they cluster in the parts of the framework that touch the outside world: fetching URLs, serving a chat UI and exporting telemetry.

AdvisorySeverityAreaWhat it allowed
GHSA-h4xc-3qfq-jf93 (Aug 2026)HighDevelopment web chat UIA website a developer visited could trigger agent runs and tool calls with the local process's credentials. Fixed in 1.107.4 and 2.28.0
GHSA-jpr8-2v3g-wgf9 (Jul 2026)ModerateUI adaptersA client could get a server tool to run with arguments it supplied. Fixed in 1.107.1 and 2.5.0
GHSA-vmxc-h2x2-jmf3 (Sep 2026)Moderateweb_fetch SSRF protectionCloud-metadata and private-IP blocklist bypass via an IPv6 zone identifier
GHSA-22h6-qm39-v87j (Sep 2026)Lowweb_fetch_tool domain listsBlocked domains bypassed through a hostname the resolver normalizes differently
GHSA-4x9p-g9wm-8q7f and GHSA-3gh4-cghq-f8v4 (Aug and Sep 2026)LowOpenTelemetry instrumentationContent and retry prompts reaching spans despite include_content=False

The earlier SSRF issue, CVE-2026-25580, affected URL downloads before 1.56.0. The pattern across all of them is the same one we see on every framework: the moment an agent can fetch a URL or be reached from a browser, the attack surface is the network, not the model. Upgrading closes each known bug. It does not give you a policy on which domains an agent may fetch from in the first place, or a record of what it fetched and why.

Two notes on the telemetry advisories, because they are easy to dismiss. If you export spans to a third-party backend and set include_content=False to keep customer data out of it, those bugs meant some content went anyway. Treat traces as business data, restrict who can read them, and do not rely on one flag as your data-loss control. Our AI data leak prevention page covers where redaction belongs.

Pydantic AI guardrails compared with a runtime security layer

Pydantic AI Harness, a separate package from core, now ships input, output and tool guardrails as capabilities. In Pydantic's words, guardrails "put a validation layer on the three edges of an agent run". The built-in detectors redact secrets, redact personal data such as card numbers and US Social Security numbers, and block keywords. Pydantic is candid about the limit: "A regex finds a credential because credentials have a shape. It does not find a prompt injection, which is ordinary language." It also notes that during run_stream() an output guardrail runs on the final output only, so a block cannot un-send chunks already streamed.

NeedPydantic AI and Harness aloneWith AgentShield added
Validated, typed tool argumentsYes, nativelyNo change. Keep using it
Redacting secrets and card numbers from promptsYes, Harness detectorsNo change needed for this alone
Detecting injected instructions in fetched pages and documentsCustom guardrail you write and maintainInspection on retrieved content before the model reads it
Per-tool allow, deny and hold policy across many agentsCode in each agent and each toolOne policy at the gateway, per tool and per argument threshold
Approval that a hostile client cannot grant itselfPossible, if you build the back-office flowHeld calls go to a named approver outside the agent process
Consistent controls across Pydantic AI, LangChain, OpenAI and Microsoft agentsPydantic AI onlySame policy for every framework calling the same tools

Two of those six rows say buy nothing, and we mean it. A single internal agent that reads data, with no write tools and no web fetch, is well served by Pydantic AI plus Harness. Agents that answer from a vector store carry their own version of the injection problem, covered on our RAG security page. The case for a runtime layer starts when an agent can change something that matters, when it reads untrusted content from the web or email, or when you run more than one framework and want one set of rules.

How to put AgentShield in front of Pydantic AI tools and MCP servers

You keep writing agents exactly as before. The change is where outbound calls go.

  1. Inventory tools with side effects. List every function tool, MCP server and external API your agents can reach, and mark which ones write, send, pay or delete. Include web_fetch and anything that reads email or documents, because that is where injected instructions arrive.
  2. Route MCP and HTTP through the gateway. Point your MCPToolset servers at the AgentShield MCP gateway and your HTTP tools at the AgentShield proxy. Pure in-process functions that only compute stay as they are.
  3. Write policy per tool. Allow reads broadly, cap writes by value, deny anything a given agent should never touch. Our AI agent permissions management is where those rules live, and they apply no matter which framework made the call.
  4. Keep requires_approval, and move the approver. Pydantic's deferred tools still pause the run cleanly. Pair them with a gateway hold so the final yes comes from a named person in your organization, not the client that made the request.
  5. Record decisions, not just spans. Each verdict lands in a tamper-evident audit trail with requester, tool, validated arguments and outcome, alongside the OpenTelemetry traces you already export, and live agent monitoring flags unusual tool activity as it happens.

Teams running Pydantic AI next to other frameworks should compare notes with our LangChain security and OpenAI agent security pages, since the same policy covers both. For tool servers you build yourself, MCP server security covers what to lock down on the server side.

FAQ

Common questions about pydantic ai security.

Is Pydantic AI secure?

Pydantic AI is well engineered and patches advisories quickly, and its typed validation removes a class of malformed-input bugs. It does not decide what an agent is allowed to do: tools run without approval by default, authorization must be written inside each tool, and approval returned by an untrusted client is not a security boundary.

Does Pydantic AI have guardrails?

Yes. Pydantic AI Harness, a separate package, provides input, output and tool guardrails as capabilities, with detectors that redact secrets and personal data and block keywords. Pydantic notes that regex detectors cannot find prompt injection, which is ordinary language, and that output guardrails cannot recall content already streamed.

Does Pydantic AI support human in the loop?

Yes. Set requires_approval=True on a tool, or raise ApprovalRequired conditionally, and the run ends with DeferredToolRequests listing the pending calls. You approve or deny each and resume with DeferredToolResults. It is opt-in, and Pydantic warns that approval is not an authorization boundary against an untrusted client.

How long will Pydantic AI V1 get security fixes?

Pydantic commits to security fixes for V1 for at least six months after V2's stable release. V2.0.0 went stable on June 23, 2026, so the guaranteed window runs to roughly late December 2026. V1 received a security backport in September 2026 with release 1.107.6.

What changed in Pydantic AI V2 that affects security?

The default end_strategy moved from early to graceful, so tools can now run alongside a final output; deferred tool types were renamed; MCP servers moved to MCPToolset and the toolsets argument; and instrumentation moved to format version 5, where approvals are no longer recorded as span errors. Each needs a test, not just a rename.

Can Pydantic AI agents use MCP servers safely?

They can, with care. Pydantic AI connects to MCP servers through MCPToolset and can wrap them so tools require approval. It does not assess whether a server or its tools are trustworthy, so allowlist the tools each agent may call, cap what writes can change and inspect tool output for injected instructions.

Is the Pydantic AI web fetch tool safe to use?

It has SSRF protection and domain lists, and September 2026 releases 2.44.0 and 1.107.6 fixed bypasses of both, plus a response-processing issue that could block the event loop. Upgrade first. Then limit which domains each agent may fetch and inspect fetched pages for instructions aimed at the model.

Does AgentShield work with Pydantic AI?

Yes, on the outbound path. You route the MCP servers and HTTP APIs your Pydantic AI agents call through AgentShield, which checks each call against your policy, holds high-impact actions for a named approver, inspects fetched and retrieved content for injected instructions and records every decision. Your agent code and types stay as they are.

Secure your pydantic ai security.