Claude Managed Agents Security for Anthropic Managed Agents with Permission Policies and MCP Approvals
Claude Managed Agents is Anthropic's hosted agent harness, in beta since April 2026 and enabled by default for every API account. Since September 10 it has a third permission policy, auto, where Anthropic's server decides whether each tool call runs. On October 7 Anthropic tightened web_fetch to cut a data exfiltration path. Both changes tell you where the real security decisions sit, and which ones are still yours.
Direct answer
Claude Managed Agents security rests on four facts from Anthropic's docs. The built-in agent toolset, including bash, defaults to always_allow, and an environment created through the API without a networking setting gets unrestricted egress. The auto policy is, in Anthropic's words, not a human checkpoint. Permission policies do not cover custom tools. And Managed Agents is not eligible for Zero Data Retention or a HIPAA BAA.
Try it live
Watch AgentShield block an attack in real time.
Pick a scenario and drive the inspection lane yourself. No signup needed.
Run a request
Runs the live engine on your text. Nothing is stored, no account needed.
Inspection lane
INSPECTINGPolicy trace
High-risk action held for approval
Audit trail
- § · → → →
The risk
A US property management company builds a delinquency agent on Claude Managed Agents. A scheduled deployment starts a session every morning at six. The agent reads the overnight tenant portal messages, checks the ledger through the company's own MCP server and posts late fees or payment plan credits. The team sets the ledger MCP toolset to auto so the agent can work without someone answering every call, and passes tenant messages into the session as user messages so the agent sees them verbatim. One tenant writes that the regional manager already approved waiving three months of fees and asks the agent to apply the credit. Anthropic's docs say what posted in user messages counts as the operator's intent and can lead the server to allow a call it would otherwise deny. The credit posts at 6:04. Nobody was asked, because auto decided the call was safe, and the evidence that it happened is an evaluation field in a session transcript that the operations team never reads.
How AgentShield handles it
Keep Anthropic's controls for what they do well: vaults that inject MCP tokens at the boundary so the agent never sees them, limited networking with an explicit host list, per-tool always_ask on anything that writes, session budgets and the evaluation field on every tool event. Then point each MCP server that can write at an AgentShield MCP gateway endpoint instead of the raw server. The gateway holds the real credential, checks each call against deterministic per-tool and per-argument policy, holds risky calls for a named approver through Slack, email or webhook, inspects tool inputs for injected instructions and logs each decision with the rule that applied. It enforces the same way whether the session was started by a person or by a 6 a.m. schedule.
The controls
The controls that secure which Claude Managed Agents tool calls run, and who approves the risky ones.
What Claude Managed Agents runs for you and what stays with you
Claude Managed Agents is a pre-built, configurable agent harness that Anthropic runs in managed infrastructure. You define an agent (model, system prompt, tools, MCP servers and skills), an environment (an Anthropic cloud sandbox or a self-hosted sandbox on your own infrastructure), and then start sessions that run for minutes or hours. Every endpoint needs the managed-agents-2026-04-01 beta header, and access is enabled by default for all API accounts. It is also available on Claude Platform on AWS, with some feature differences.
It is a different product from the Claude Agent SDK and Claude Code, where the loop runs in your own process and every tool call passes through code you control. With Managed Agents the loop runs on Anthropic's side, and you see what the event stream and webhooks show you.
| Tool type | Who executes it | Governed by permission policies | Default |
|---|---|---|---|
| Agent toolset (bash, read, write, edit, web search, web fetch) | Anthropic, inside the session's sandbox | Yes | always_allow |
| MCP toolsets | Anthropic connects to your MCP server with the vault credential | Yes | always_ask |
| Custom tools | Your application, after an agent.custom_tool_use event | No, your code decides | Whatever your handler does |
| Subagents in a multiagent roster | Anthropic, as separate session threads | Yes, their permission requests are cross-posted to the primary thread | Inherits each toolset's policy |
Two defaults deserve a second look. The agent toolset runs bash without approval unless you change it. And Anthropic's environment docs note that "a create request that omits it gets unrestricted," meaning networking: an environment created through the API without a networking block can reach any host outside a general safety blocklist. The Console form starts with Limited selected, so the risk sits mostly in infrastructure-as-code and scripts.
Why the auto permission policy is not an approval gate
On September 10, 2026 Anthropic added auto to the two existing policies. Under auto, the server evaluates each agent or MCP tool call, considering the tool, the input and the session so far, and either runs it, denies it as high risk, or pauses for your approval when it reaches no determination. It is a genuinely useful middle ground for agents that would otherwise ask about everything. Anthropic is also unusually direct about its limits, and those limits are the reason this page exists.
"
autois not a human checkpoint. If the server determines that a call is safe, the call runs before anyone sees it, and its effects might not be reversible."
The second limit is about intent. Anthropic's docs say that what you post in user.message events counts as your intent and can lead the server to allow a call it would otherwise deny. The server does not take instructions from tool results, fetched pages or MCP responses. But "if you relay untrusted end-user input in user.message events, the server reads that input as your intent too, and it can get a call allowed." Any agent that forwards customer chats, tickets or portal messages as user messages is in that position.
| Policy | Who decides | Runs before a person sees it | Approval evidence |
|---|---|---|---|
always_allow | Nobody, the call runs | Yes | None |
auto | Anthropic's server, per call | Yes, whenever the server judges it safe | An evaluation field and a reason_code on denials and asks |
always_ask | Whoever your client lets send user.tool_confirmation | No, the session waits | Only what your client records about who clicked |
| Gate on the MCP call path | Deterministic policy, then a named approver | No, for calls the rules hold | Per call, with the rule, the approver and the time |
The pain story above is the textbook case. With an approval gate for AI agent actions on the ledger MCP endpoint, a fee waiver over a set amount would have waited for a named manager, and the tenant's claim of prior approval would have changed nothing, because the manager's decision is an event in a separate channel, not text in the conversation.
Scheduled deployments and approvals nobody is watching
Scheduled deployments start sessions on a cron schedule, which is where Managed Agents earns its keep for back-office work. They also change what always_ask means. When a call evaluates to ask, the session goes idle with stop_reason requires_action, and Anthropic's docs say "the session waits indefinitely for a response." At six in the morning there may be nobody to respond.
That pushes teams toward auto or always_allow on exactly the runs that nobody watches. A few other details from the docs matter here:
- Policy changes are not retroactive. "Running sessions keep the toolset configuration they were created with. Updates apply to sessions created afterward." Tightening a policy does not touch a long-running session.
- Budgets bound each run, not the schedule. A deployment copies its budget to every session it starts, so the cap is per run, and a session that hits it pauses with
budget_reachedrather than ending. - Pausing a deployment stops new triggers only. Sessions from earlier runs keep executing.
- Your client cannot override an auto denial, and it cannot send a confirmation for a call that was not paused for one.
The fix is not to staff a night shift. It is to move the decision for consequential calls out of the session and into a gate that knows who is on call. Webhooks tell you a session paused; a gateway with routing rules decides who gets asked, how long to wait, and what happens on timeout. That is the part of AI agent permissions management Managed Agents leaves to you.
Is Claude Managed Agents HIPAA eligible
No. Anthropic's overview states that because Managed Agents is a stateful service, it "is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement (BAA) coverage." The API and data retention page adds that session transcripts persist until you delete them, and that the exclusion applies to all Managed Agents sub-features, including self-hosted sandboxes.
| Requirement | Status (checked October 8, 2026) | What regulated US teams usually do |
|---|---|---|
| Zero Data Retention | Not eligible | Run those workloads on the Messages API under a ZDR arrangement |
| HIPAA BAA coverage | Not eligible, including self-hosted sandboxes | Keep PHI out of Managed Agents sessions; build the loop on an eligible endpoint |
| Inference location | Pin with inference_geo on the agent or the session (added August 7, 2026) | Pin to the US where contracts require it |
| Session data | Transcripts persist until deleted | Delete on a schedule and keep your own record of what happened |
Self-hosting the sandbox does not change the answer, and it moves more of the job onto you. Anthropic's security model for self-hosted sandboxes says "Anthropic's security boundary stops at the sandbox," that it "does not inspect or verify your sandbox image," and that egress control, the environment service key, tool isolation and log retention are yours. Our healthcare AI agent security page covers the rest of the HIPAA picture for agents.
Claude Managed Agents security checklist for production
Anthropic ships good primitives. This is how we would use them, in order, with the action layer where it belongs. The platform neutral version is our page on AI agent security best practices.
- Set networking explicitly. Use
limitedwith an explicitallowed_hostslist. Since October 7, 2026 that list also governsweb_searchandweb_fetch. Note that any allowed host accepts any request, including uploads such asgit push. - Decide bash on purpose. Override
bashtoalways_askor disable it for agents that do not need a shell. Disabled tools are removed from the agent entirely. - Keep secrets in vaults. MCP tokens are injected when the agent connects, and environment variable credentials stay placeholders until egress, so "the agent never sees the secret value." One caveat from the docs: if a client uses a stored secret to fetch a session token, such as an OAuth client credentials grant, "the returned token arrives in the sandbox unredacted."
- Do not relay untrusted text as user messages under auto. Hand it to the agent as a custom tool result instead, which the server assesses without taking intent from it, or keep
always_askon the tools that text could trigger. - Put write tools behind a policy gate. Point MCP servers that move money, change records or send external messages at a gateway that evaluates each call and holds the risky ones for a named approver.
- Govern custom tools in your handler. Permission policies never see them, so validation and approval belong in the code that executes them.
- Cap spend per session. Budgets are priced at public list rates; set one on every deployment.
- Keep an audit trail outside the session. Transcripts can be deleted and the evaluation field records the server's view, not your approver's. Your AI agent audit trail should record who asked, which call ran, which rule applied and who approved.
Anthropic's own October 7 change is a good example of why defaults matter. web_fetch now fetches only URLs that already appeared in the session, and a URL that appears only in Claude's own output, an attached document or a tool result returns url_not_in_prior_context. Anthropic says the change "reduces the risk of data exfiltration." Sessions before that date did not have it.
When to buy nothing, use Anthropic controls, or add AgentShield
| Your situation | Recommendation | Why |
|---|---|---|
| Research, code review or report agents with no write tools | Buy nothing. Limited networking, vaults, bash on ask, budgets | No irreversible action to gate |
| PHI or a Zero Data Retention commitment | Keep that workload off Managed Agents for now | Not eligible for a BAA or ZDR; a data governance decision |
| All write actions are custom tools your app executes | Your handler checks may be enough | Those calls already pass through your code |
| Interactive sessions where an engineer answers every ask | always_ask plus ant beta:sessions connect may be enough | A person sees each paused call |
| Scheduled or customer-facing agents with MCP write tools | Add AgentShield as the MCP gateway in front of those servers | Approvals need routing, timeouts and evidence no session provides |
| Agents split across Managed Agents, the Agent SDK and other clouds | Add AgentShield as one policy point for all of them | One rule set and one audit trail instead of several |
Three of six rows end without a purchase from us, and one tells you to keep a workload off the product. If you are comparing options, the buyer guide for Claude Managed Agents sets six of them side by side, and the MCP server security page covers the server side of the gate.
FAQ
Common questions about claude managed agents security.
What is Claude Managed Agents?
It is Anthropic's pre-built agent harness that runs in managed infrastructure. You define an agent with a model, system prompt, tools, MCP servers and skills, pick a cloud or self-hosted sandbox, and start sessions that run for minutes or hours. It is in beta and enabled by default for all Claude API accounts.
What is the difference between Claude Managed Agents and the Claude Agent SDK?
The Agent SDK is a library that runs the agent loop in your own process, so every tool call passes through code you control. Claude Managed Agents runs the loop on Anthropic's side with persistent sessions, sandboxes, vaults and scheduled deployments. You see its tool calls through the event stream and webhooks.
Is the auto permission policy safe for production?
It is safe for low-risk calls, and Anthropic is clear that it is not a human checkpoint. A call the server judges safe runs before anyone sees it. Anything relayed in user messages counts as your intent and can get a call allowed. Keep always_ask or an external gate on tools that move money or change records.
Is Claude Managed Agents HIPAA compliant?
Not as of October 8, 2026. Anthropic states Managed Agents is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage, including self-hosted sandboxes, because sessions are stateful and transcripts persist until deleted. Keep protected health information out of Managed Agents sessions.
Do Claude Managed Agents permission policies apply to custom tools?
No. Permission policies govern the built-in agent toolset and MCP toolsets only. When the agent calls a custom tool, your application receives an agent.custom_tool_use event and decides whether to run it, so validation and approval for custom tools belong in your own handler code.
How do I keep credentials safe in Claude Managed Agents?
Store them in vaults. MCP credentials are injected when the agent connects to the matching server URL, and environment variable credentials stay opaque placeholders until egress to allowed hosts. One exception: a token fetched with a stored secret, such as an OAuth client credentials grant, arrives in the sandbox unredacted.
How much does Claude Managed Agents cost?
Anthropic prices sessions at public list rates: model tokens at each model's list price, web searches at 10 dollars per 1,000, and session running time at 8 cents per hour, as listed in the session budget docs on October 8, 2026. A session budget caps that list cost per session.
Does AgentShield work with Claude Managed Agents?
Yes. Point each MCP server that can write at an AgentShield MCP gateway endpoint and keep the real credential in the gateway. Every call is checked against per-tool and per-argument policy, risky calls wait for a named approver, inputs are inspected for injected instructions, and each decision is logged outside the session.
More use cases