OpenAI Agents API Security for OpenAI Managed Agents on the Codex Harness
OpenAI released the Agents API in public beta on September 10, 2026, and added computer use on September 29. It puts the Codex harness behind one API: OpenAI runs the session, the orchestration and the sandbox, and your application supplies tools. That changes where security decisions can happen, and three lines in OpenAI's own documentation decide whether your design holds up.
Direct answer
OpenAI Agents API security comes down to three facts from OpenAI's docs. The API retains session state, supports data residency only in the United States and does not support Zero Data Retention, even with a self-hosted sandbox. Remote MCP calls go from OpenAI straight to your server, so your app never sees them. And a confirmation step built as a function tool only runs if the agent decides to call it.
Try it live
Watch AgentShield block an attack in real time.
Pick a scenario and drive the inspection lane yourself. No signup needed.
Run a request
Runs the live engine on your text. Nothing is stored, no account needed.
Inspection lane
INSPECTINGPolicy trace
High-risk action held for approval
Audit trail
- § · → → →
The risk
A US revenue cycle company builds a claims follow-up agent on the Agents API in its first beta week. The agent gets an OpenAI-hosted sandbox, a vault with a bearer token for the company's billing MCP server, and an instruction to "always call request_supervisor_approval before posting an adjustment." In the demo it calls the approval function every time. Three weeks later a payer remittance file arrives with a note in the free text field telling the reader the adjustment was pre-approved by the supervisor. The agent posts the adjustment directly through the MCP tool. Nothing in the application logs shows it happening, because remote MCP calls run from OpenAI's service to the billing server and never pass through the application. The approval function was a request to the model, not a control. Separately, compliance asks whether patient account data is covered by the BAA the company accepted on October 5. The Agents API is not on OpenAI's HIPAA eligible endpoint list.
How AgentShield handles it
Keep OpenAI's controls for what they do well: scoped application keys, vaults that keep secrets out of the sandbox, restricted network mode, an allowed tools list on each MCP connection and origin approvals for the hosted browser. Then point every MCP connection that can write at an AgentShield MCP gateway endpoint instead of the raw server. The gateway holds the real credential, evaluates each call against a deterministic per-tool, per-argument policy, holds risky calls for a named approver in a separate channel, inspects tool inputs for injected instructions and logs every decision with the rule that applied. Because enforcement sits on the call path, it does not depend on the agent choosing to ask.
The controls
The controls that secure what your OpenAI managed agents may call, and who approves it first.
What the OpenAI Agents API manages and what stays with you
The OpenAI Agents API is a beta API that gives your application an OpenAI-managed Codex agent. In OpenAI's words, "OpenAI manages sessions, orchestration, context compaction, and recovery while your application provides tools and chooses its execution environment." An agent can run code, edit files, connect to MCP servers, delegate to subagents and keep working in the same durable session across turns. Requests carry the OpenAI-Beta: agents=v1 header, and the API is separate from the OpenAI Agents SDK, which runs the agent loop inside your own process.
That split matters for security. With the SDK, every tool call passes through code you wrote, so a guardrail or an approval hook can sit in the path. With the Agents API, the loop runs on OpenAI's side and your application sees only what OpenAI routes back to it.
| Component | Who runs it | Does your app see each call | What you control |
|---|---|---|---|
| Session, orchestration, compaction, subagents | OpenAI | Through streamed events and webhooks | Instructions, model, which tools exist |
| Function tools | Your application | Yes, as agent.session.requires_action | The handler, its credentials, its result |
Remote MCP over HTTP (default connection_origin: "service") | OpenAI's service calls your MCP server | No. OpenAI: "Your application does not need to handle each call" | allowed_tools, the vault credential, the server itself |
| Executor MCP (HTTP from the environment, or stdio) | Your session's environment | Only if the server logs it | Network rules and the environment |
| OpenAI-hosted sandbox | OpenAI | Through events and artifacts | Packages, setup commands, network mode |
| Hosted browser (computer use, added September 29, 2026) | OpenAI | Origin approval requests only | Approve, deny or cancel each new website origin |
The third row is the one that surprises teams. When a remote MCP server holds write tools, the call goes from OpenAI to that server with the credential from your vault, and the first place a policy can act is the server itself. If you need a per-tool, per-argument permission check, it has to live there or in a gateway in front of it.
Anthropic's managed equivalent has the same shape with different defaults: MCP tools ask first, the built-in shell does not, and a server-side auto policy can approve calls on its own. Our Claude Managed Agents security page covers it.
Does the OpenAI Agents API support Zero Data Retention
No. OpenAI's Agents API FAQ states: "The Agents API currently supports data residency in the United States only. It does not support Zero Data Retention, including when you use a self-hosted sandbox." The reason is structural. The API retains session state so an agent can continue work across turns, and that retained state is the product. OpenAI lets you delete sessions and published artifacts when you no longer need them, which is deletion on request rather than retention you can turn off.
For a US buyer, the residency half is rarely a problem. The ZDR half and the BAA question are where procurement stalls.
| Requirement | Agents API status (checked October 7, 2026) | What regulated US teams usually do |
|---|---|---|
| Zero Data Retention | Not supported, including with a self-hosted sandbox | Use the Responses API or Chat Completions on a ZDR-approved project for those workloads |
| Data residency | United States only | Fine for US data; a blocker for EU or other regional residency commitments |
| HIPAA BAA coverage | OpenAI's list of HIPAA eligible API endpoints includes /v1/responses and /v1/chat/completions and does not list the Agents API | Keep PHI out of Agents API sessions until OpenAI adds it, or build the loop with the Agents SDK on an eligible endpoint |
| Session data deletion | Sessions and published artifacts can be deleted; save files first | Delete on a schedule and record the deletion |
OpenAI made the BAA much easier to get on October 5, 2026: eligible API organizations can now accept the standard Business Associate Agreement in Organization settings. API eligibility still requires an executed BAA and, unless OpenAI says otherwise, Modified Retention. OpenAI is also direct that accepting a BAA and enabling HIPAA support "do not, by themselves, make your application HIPAA compliant." If you are a covered entity or business associate, the practical rule is simple: an endpoint that is not on the eligible list is not where PHI goes. Our healthcare AI agent security page covers the rest of the HIPAA picture for agents.
None of this is a criticism of the product. It is a beta, and the retention model is what makes durable sessions work. It does mean the security review for an Agents API project starts with data classification, not with prompt injection.
Why an approval function tool is not an approval gate
The most common pattern we see in early Agents API builds is a function tool named something like request_approval plus an instruction telling the agent to call it before anything consequential. It looks like a human in the loop. OpenAI's own computer use guide explains why it is not one: "Asking for confirmation through a function tool relies on the agent calling that function."
The same page adds that origin approval for the hosted browser "does not enforce confirmation before individual actions," and that if you must guarantee confirmation before purchases or destructive changes, you should restrict the browser to resources that cannot perform them or use a browser runtime you control. That is good, honest guidance, and it generalizes to every tool the agent can reach.
| Control | Where it is enforced | Survives a prompt injection | Produces approval evidence |
|---|---|---|---|
| Instruction to "always ask first" | In the model's reasoning | No | No |
| Approval function tool | Only when the agent calls it | No, the agent can skip it | Only for calls that were routed through it |
allowed_tools on an MCP connection | OpenAI, at discovery and call time | Yes, for tools not on the list | No, it is allow or not offered |
| Hosted browser origin approval | OpenAI, before each new website origin | Yes, for new origins | Yes, per origin, not per action |
| Policy gate on the MCP server or a gateway in front of it | On the call path, after the model decides | Yes | Yes, per call, with the approver's identity |
The difference between the last row and the second is the whole argument. A gate on the call path evaluates the actual tool name and arguments every time, whether or not the model remembered its instructions. In the pain story above, an approval gate for AI agent actions on the billing MCP endpoint would have held the adjustment for a named supervisor, and the injected "pre-approved" note would have changed nothing, because the supervisor's decision is an event in a separate channel, not text the agent read.
OpenAI Agents API security checklist for production
OpenAI publishes a solid sandbox security guide. Its opening line sets the tone: "Agent-generated code can access the files, credentials, and network available to its environment." Here is that guidance plus the action layer, in the order we would apply it. The platform neutral version of this list is our page on AI agent security best practices.
- Split keys. Give the application key only
api.agents.read,api.agents.writeandapi.responses.write, plus vault scopes if it manages vaults. Give a self-hosted executor its own environment key. OpenAI notes that agent-generated code can read the environment key, so it can do nothing except connect environments. - Keep third-party secrets out of the sandbox. Use vaults: a
static_bearerormcp_oauthcredential for remote MCP, or anenvironment_variablecredential where sandbox code sees a placeholder and a proxy swaps in the real secret for approved hosts. OpenAI warns that injecting a stored secret into the environment still exposes it to agent-generated code. - Restrict the network. Restricted mode on an OpenAI-hosted sandbox accepts 1 to 100 exact host names, no wildcards, and redirect targets need their own entries. Note that hosted stdio MCP servers currently require network access to be enabled, so put write tools on HTTP MCP behind a gateway instead.
- Trim every MCP connection with
allowed_tools. A tool the agent cannot discover is a tool it cannot be talked into calling. - Put write tools behind a policy gate. Point remote MCP connections that can move money, change records or send external messages at a gateway that evaluates each call and holds the risky ones. This is the step that turns an instruction into a control.
- Verify outcomes, not idle sessions. OpenAI's FAQ says a completed turn "does not mean every tool succeeded" and an idle session "alone does not establish that the task finished successfully." Reconcile against the system of record.
- Isolate tenants. Use separate environments for users or workloads that must not share data, and a dedicated OpenAI project per application.
- Keep an audit trail outside the session. Session history lives with OpenAI and can be deleted. Your AI agent audit trail should record who asked, which call was made, which rule applied and who approved, in a store the agent cannot write.
Steps one through four and six through seven are configuration, not purchases. Step five is the one most teams cannot build in a sprint, which is why it is the part we sell. If your MCP servers are your own, the MCP server security page shows what the server side of that gate needs.
When to buy nothing, use OpenAI controls, or add AgentShield
| Your situation | Recommendation | Why |
|---|---|---|
| Read only agents: research, code review, report drafting with no write tools | Buy nothing. Scoped keys, vaults, restricted network, allowed_tools | No irreversible action to gate |
| Workloads that need ZDR, or PHI under a BAA | Do not use the Agents API for them yet. Build the loop with the Agents SDK on an eligible endpoint | A data governance decision, not a tooling purchase |
| Write tools only through function tools your app already validates | Your own handler checks may be enough | Function calls already pass through your code |
| Remote MCP servers with payment, refund, CRM or email send tools | Add AgentShield as the MCP gateway in front of them | Those calls never pass through your application |
| Hosted browser tasks that can buy, submit or delete | Restrict the browser to resources that cannot, as OpenAI advises, or run your own browser runtime | Origin approval is per site, not per action |
| Agents split across the Agents API, the Agents SDK and other clouds | Add AgentShield as one policy point for all of them | One rule set and one audit trail instead of three |
Three of six rows do not end in a purchase from us, and one tells you to keep a workload off the product entirely for now. If you are also running Agent Builder workflows, note that OpenAI shuts Agent Builder down on November 30, 2026; our Agent Builder migration page covers that move, and the buyer guide for the OpenAI Agents API compares six options side by side.
FAQ
Common questions about openai agents api security.
What is the OpenAI Agents API?
It is a beta API, released September 10, 2026, that gives your application an OpenAI-managed Codex agent. OpenAI runs sessions, orchestration, context compaction and recovery, and optionally the sandbox. Your application supplies the task, tools and configuration, then follows progress through streamed events or webhooks.
What is the difference between the OpenAI Agents API and the Agents SDK?
The Agents SDK is a library that runs the agent loop inside your own process, so every tool call passes through your code. The Agents API runs the loop on OpenAI's side as a managed service with durable sessions. Remote MCP calls in the Agents API go from OpenAI to your server without passing through your application.
Does the OpenAI Agents API support Zero Data Retention?
No. OpenAI states the Agents API supports data residency in the United States only and does not support Zero Data Retention, including when you use a self-hosted sandbox. Session state is retained so work can continue across turns. You can delete sessions and published artifacts when they are no longer needed.
Is the OpenAI Agents API HIPAA eligible?
Not as of October 7, 2026. OpenAI's list of HIPAA eligible API endpoints includes the Responses API and Chat Completions, among others, and does not list the Agents API. Teams handling PHI should keep it out of Agents API sessions and use an eligible endpoint under an executed BAA with Modified Retention.
Can I require human approval before an Agents API agent calls a tool?
Only partly with native features. Function tools pause for your code to return a result, and the hosted browser asks before each new website origin. Remote MCP calls have no per-call approval step, and OpenAI notes that confirmation through a function tool relies on the agent calling that function. A gate on the call path closes that gap.
How do I keep API keys and secrets safe in an Agents API sandbox?
Keep your application key outside the sandbox, give a self-hosted executor its own environment key, and store third-party credentials in vaults. For hosted sandboxes, a vault environment variable gives code a placeholder that a proxy swaps for the real secret on approved hosts. Anything injected into the environment is readable by agent code.
How much does the OpenAI Agents API cost?
There is no separate Agents API fee. Model usage is billed at the selected model's API rates, OpenAI tools at their standard tool rates, and OpenAI-hosted sandboxes at standard container rates. Check OpenAI's pricing page for current numbers before you budget a production workload.
Does AgentShield work with the OpenAI Agents API?
Yes. Point each MCP connection that can write at an AgentShield MCP gateway endpoint instead of the raw server, and keep the real credential in the gateway. Every call is checked against per-tool, per-argument policy, risky calls wait for a named approver, inputs are inspected for injected instructions, and each decision is logged.
More use cases