AgentShield
How it works Pricing Blog FAQ Contact Sign in

Claude Managed Agents Security for Anthropic Managed Agents with Permission Policies and MCP Approvals

Claude Managed Agents is Anthropic's hosted agent harness, in beta since April 2026 and enabled by default for every API account. Since September 10 it has a third permission policy, auto, where Anthropic's server decides whether each tool call runs. On October 7 Anthropic tightened web_fetch to cut a data exfiltration path. Both changes tell you where the real security decisions sit, and which ones are still yours.

OWASP LLM Top 10 Immutable audit trail Never trains on your data

Direct answer

Claude Managed Agents security rests on four facts from Anthropic's docs. The built-in agent toolset, including bash, defaults to always_allow, and an environment created through the API without a networking setting gets unrestricted egress. The auto policy is, in Anthropic's words, not a human checkpoint. Permission policies do not cover custom tools. And Managed Agents is not eligible for Zero Data Retention or a HIPAA BAA.

Try it live

Watch AgentShield block an attack in real time.

Pick a scenario and drive the inspection lane yourself. No signup needed.

Threat Console
Interactive demo · 0 blocked in this session

Run a request

Runs the live engine on your text. Nothing is stored, no account needed.

Inspection lane

INSPECTING
⌖ untrusted input

Policy trace

High-risk action held for approval

Audit trail

The risk

A US property management company builds a delinquency agent on Claude Managed Agents. A scheduled deployment starts a session every morning at six. The agent reads the overnight tenant portal messages, checks the ledger through the company's own MCP server and posts late fees or payment plan credits. The team sets the ledger MCP toolset to auto so the agent can work without someone answering every call, and passes tenant messages into the session as user messages so the agent sees them verbatim. One tenant writes that the regional manager already approved waiving three months of fees and asks the agent to apply the credit. Anthropic's docs say what posted in user messages counts as the operator's intent and can lead the server to allow a call it would otherwise deny. The credit posts at 6:04. Nobody was asked, because auto decided the call was safe, and the evidence that it happened is an evaluation field in a session transcript that the operations team never reads.

How AgentShield handles it

Keep Anthropic's controls for what they do well: vaults that inject MCP tokens at the boundary so the agent never sees them, limited networking with an explicit host list, per-tool always_ask on anything that writes, session budgets and the evaluation field on every tool event. Then point each MCP server that can write at an AgentShield MCP gateway endpoint instead of the raw server. The gateway holds the real credential, checks each call against deterministic per-tool and per-argument policy, holds risky calls for a named approver through Slack, email or webhook, inspects tool inputs for injected instructions and logs each decision with the rule that applied. It enforces the same way whether the session was started by a person or by a 6 a.m. schedule.

Platform engineer answering a paused agent request on a laptop in the evening

The controls

The controls that secure which Claude Managed Agents tool calls run, and who approves the risky ones.

What Claude Managed Agents runs for you and what stays with you

Claude Managed Agents is a pre-built, configurable agent harness that Anthropic runs in managed infrastructure. You define an agent (model, system prompt, tools, MCP servers and skills), an environment (an Anthropic cloud sandbox or a self-hosted sandbox on your own infrastructure), and then start sessions that run for minutes or hours. Every endpoint needs the managed-agents-2026-04-01 beta header, and access is enabled by default for all API accounts. It is also available on Claude Platform on AWS, with some feature differences.

It is a different product from the Claude Agent SDK and Claude Code, where the loop runs in your own process and every tool call passes through code you control. With Managed Agents the loop runs on Anthropic's side, and you see what the event stream and webhooks show you.

Tool typeWho executes itGoverned by permission policiesDefault
Agent toolset (bash, read, write, edit, web search, web fetch)Anthropic, inside the session's sandboxYesalways_allow
MCP toolsetsAnthropic connects to your MCP server with the vault credentialYesalways_ask
Custom toolsYour application, after an agent.custom_tool_use eventNo, your code decidesWhatever your handler does
Subagents in a multiagent rosterAnthropic, as separate session threadsYes, their permission requests are cross-posted to the primary threadInherits each toolset's policy

Two defaults deserve a second look. The agent toolset runs bash without approval unless you change it. And Anthropic's environment docs note that "a create request that omits it gets unrestricted," meaning networking: an environment created through the API without a networking block can reach any host outside a general safety blocklist. The Console form starts with Limited selected, so the risk sits mostly in infrastructure-as-code and scripts.

Why the auto permission policy is not an approval gate

On September 10, 2026 Anthropic added auto to the two existing policies. Under auto, the server evaluates each agent or MCP tool call, considering the tool, the input and the session so far, and either runs it, denies it as high risk, or pauses for your approval when it reaches no determination. It is a genuinely useful middle ground for agents that would otherwise ask about everything. Anthropic is also unusually direct about its limits, and those limits are the reason this page exists.

"auto is not a human checkpoint. If the server determines that a call is safe, the call runs before anyone sees it, and its effects might not be reversible."

The second limit is about intent. Anthropic's docs say that what you post in user.message events counts as your intent and can lead the server to allow a call it would otherwise deny. The server does not take instructions from tool results, fetched pages or MCP responses. But "if you relay untrusted end-user input in user.message events, the server reads that input as your intent too, and it can get a call allowed." Any agent that forwards customer chats, tickets or portal messages as user messages is in that position.

PolicyWho decidesRuns before a person sees itApproval evidence
always_allowNobody, the call runsYesNone
autoAnthropic's server, per callYes, whenever the server judges it safeAn evaluation field and a reason_code on denials and asks
always_askWhoever your client lets send user.tool_confirmationNo, the session waitsOnly what your client records about who clicked
Gate on the MCP call pathDeterministic policy, then a named approverNo, for calls the rules holdPer call, with the rule, the approver and the time

The pain story above is the textbook case. With an approval gate for AI agent actions on the ledger MCP endpoint, a fee waiver over a set amount would have waited for a named manager, and the tenant's claim of prior approval would have changed nothing, because the manager's decision is an event in a separate channel, not text in the conversation.

Scheduled deployments and approvals nobody is watching

Scheduled deployments start sessions on a cron schedule, which is where Managed Agents earns its keep for back-office work. They also change what always_ask means. When a call evaluates to ask, the session goes idle with stop_reason requires_action, and Anthropic's docs say "the session waits indefinitely for a response." At six in the morning there may be nobody to respond.

That pushes teams toward auto or always_allow on exactly the runs that nobody watches. A few other details from the docs matter here:

  • Policy changes are not retroactive. "Running sessions keep the toolset configuration they were created with. Updates apply to sessions created afterward." Tightening a policy does not touch a long-running session.
  • Budgets bound each run, not the schedule. A deployment copies its budget to every session it starts, so the cap is per run, and a session that hits it pauses with budget_reached rather than ending.
  • Pausing a deployment stops new triggers only. Sessions from earlier runs keep executing.
  • Your client cannot override an auto denial, and it cannot send a confirmation for a call that was not paused for one.

The fix is not to staff a night shift. It is to move the decision for consequential calls out of the session and into a gate that knows who is on call. Webhooks tell you a session paused; a gateway with routing rules decides who gets asked, how long to wait, and what happens on timeout. That is the part of AI agent permissions management Managed Agents leaves to you.

Is Claude Managed Agents HIPAA eligible

No. Anthropic's overview states that because Managed Agents is a stateful service, it "is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement (BAA) coverage." The API and data retention page adds that session transcripts persist until you delete them, and that the exclusion applies to all Managed Agents sub-features, including self-hosted sandboxes.

RequirementStatus (checked October 8, 2026)What regulated US teams usually do
Zero Data RetentionNot eligibleRun those workloads on the Messages API under a ZDR arrangement
HIPAA BAA coverageNot eligible, including self-hosted sandboxesKeep PHI out of Managed Agents sessions; build the loop on an eligible endpoint
Inference locationPin with inference_geo on the agent or the session (added August 7, 2026)Pin to the US where contracts require it
Session dataTranscripts persist until deletedDelete on a schedule and keep your own record of what happened

Self-hosting the sandbox does not change the answer, and it moves more of the job onto you. Anthropic's security model for self-hosted sandboxes says "Anthropic's security boundary stops at the sandbox," that it "does not inspect or verify your sandbox image," and that egress control, the environment service key, tool isolation and log retention are yours. Our healthcare AI agent security page covers the rest of the HIPAA picture for agents.

Claude Managed Agents security checklist for production

Anthropic ships good primitives. This is how we would use them, in order, with the action layer where it belongs. The platform neutral version is our page on AI agent security best practices.

  1. Set networking explicitly. Use limited with an explicit allowed_hosts list. Since October 7, 2026 that list also governs web_search and web_fetch. Note that any allowed host accepts any request, including uploads such as git push.
  2. Decide bash on purpose. Override bash to always_ask or disable it for agents that do not need a shell. Disabled tools are removed from the agent entirely.
  3. Keep secrets in vaults. MCP tokens are injected when the agent connects, and environment variable credentials stay placeholders until egress, so "the agent never sees the secret value." One caveat from the docs: if a client uses a stored secret to fetch a session token, such as an OAuth client credentials grant, "the returned token arrives in the sandbox unredacted."
  4. Do not relay untrusted text as user messages under auto. Hand it to the agent as a custom tool result instead, which the server assesses without taking intent from it, or keep always_ask on the tools that text could trigger.
  5. Put write tools behind a policy gate. Point MCP servers that move money, change records or send external messages at a gateway that evaluates each call and holds the risky ones for a named approver.
  6. Govern custom tools in your handler. Permission policies never see them, so validation and approval belong in the code that executes them.
  7. Cap spend per session. Budgets are priced at public list rates; set one on every deployment.
  8. Keep an audit trail outside the session. Transcripts can be deleted and the evaluation field records the server's view, not your approver's. Your AI agent audit trail should record who asked, which call ran, which rule applied and who approved.

Anthropic's own October 7 change is a good example of why defaults matter. web_fetch now fetches only URLs that already appeared in the session, and a URL that appears only in Claude's own output, an attached document or a tool result returns url_not_in_prior_context. Anthropic says the change "reduces the risk of data exfiltration." Sessions before that date did not have it.

When to buy nothing, use Anthropic controls, or add AgentShield

Your situationRecommendationWhy
Research, code review or report agents with no write toolsBuy nothing. Limited networking, vaults, bash on ask, budgetsNo irreversible action to gate
PHI or a Zero Data Retention commitmentKeep that workload off Managed Agents for nowNot eligible for a BAA or ZDR; a data governance decision
All write actions are custom tools your app executesYour handler checks may be enoughThose calls already pass through your code
Interactive sessions where an engineer answers every askalways_ask plus ant beta:sessions connect may be enoughA person sees each paused call
Scheduled or customer-facing agents with MCP write toolsAdd AgentShield as the MCP gateway in front of those serversApprovals need routing, timeouts and evidence no session provides
Agents split across Managed Agents, the Agent SDK and other cloudsAdd AgentShield as one policy point for all of themOne rule set and one audit trail instead of several

Three of six rows end without a purchase from us, and one tells you to keep a workload off the product. If you are comparing options, the buyer guide for Claude Managed Agents sets six of them side by side, and the MCP server security page covers the server side of the gate.

FAQ

Common questions about claude managed agents security.

What is Claude Managed Agents?

It is Anthropic's pre-built agent harness that runs in managed infrastructure. You define an agent with a model, system prompt, tools, MCP servers and skills, pick a cloud or self-hosted sandbox, and start sessions that run for minutes or hours. It is in beta and enabled by default for all Claude API accounts.

What is the difference between Claude Managed Agents and the Claude Agent SDK?

The Agent SDK is a library that runs the agent loop in your own process, so every tool call passes through code you control. Claude Managed Agents runs the loop on Anthropic's side with persistent sessions, sandboxes, vaults and scheduled deployments. You see its tool calls through the event stream and webhooks.

Is the auto permission policy safe for production?

It is safe for low-risk calls, and Anthropic is clear that it is not a human checkpoint. A call the server judges safe runs before anyone sees it. Anything relayed in user messages counts as your intent and can get a call allowed. Keep always_ask or an external gate on tools that move money or change records.

Is Claude Managed Agents HIPAA compliant?

Not as of October 8, 2026. Anthropic states Managed Agents is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage, including self-hosted sandboxes, because sessions are stateful and transcripts persist until deleted. Keep protected health information out of Managed Agents sessions.

Do Claude Managed Agents permission policies apply to custom tools?

No. Permission policies govern the built-in agent toolset and MCP toolsets only. When the agent calls a custom tool, your application receives an agent.custom_tool_use event and decides whether to run it, so validation and approval for custom tools belong in your own handler code.

How do I keep credentials safe in Claude Managed Agents?

Store them in vaults. MCP credentials are injected when the agent connects to the matching server URL, and environment variable credentials stay opaque placeholders until egress to allowed hosts. One exception: a token fetched with a stored secret, such as an OAuth client credentials grant, arrives in the sandbox unredacted.

How much does Claude Managed Agents cost?

Anthropic prices sessions at public list rates: model tokens at each model's list price, web searches at 10 dollars per 1,000, and session running time at 8 cents per hour, as listed in the session budget docs on October 8, 2026. A session budget caps that list cost per session.

Does AgentShield work with Claude Managed Agents?

Yes. Point each MCP server that can write at an AgentShield MCP gateway endpoint and keep the real credential in the gateway. Every call is checked against per-tool and per-argument policy, risky calls wait for a named approver, inputs are inspected for injected instructions, and each decision is logged outside the session.

Secure your claude managed agents security.