AgentShield
How it works Pricing Blog FAQ Contact Sign in

Best AI Agent Security Software for the OpenAI Agents API and Managed Codex Agents

AgentShield Security Team·Oct 7, 2026·8 min read

Try it live

Watch AgentShield block an attack in real time.

Pick a scenario and drive the inspection lane yourself. No signup needed.

Threat Console
Interactive demo · 0 blocked in this session

Run a request

Runs the live engine on your text. Nothing is stored, no account needed.

Inspection lane

INSPECTING
⌖ untrusted input

Policy trace

High-risk action held for approval

Audit trail

Two engineers comparing architecture options at a whiteboard in a small office

For most teams building on the OpenAI Agents API, the best AI agent security software starts with what OpenAI already gives you: scoped API keys, vaults that keep secrets out of the sandbox, restricted network mode and allowed_tools on every MCP connection. Add a dedicated runtime gate only when the agent can move money, change records or send messages through remote MCP servers, because those calls go from OpenAI straight to your server and never pass through your application.

We sell that runtime gate, so opening with "OpenAI's controls may be enough" is deliberate. For a read only research or code review agent, it is the correct answer.

Why teams are reviewing Agents API security now

OpenAI released the Agents API in public beta on September 10, 2026. It exposes the Codex harness as a managed service: OpenAI runs sessions, orchestration, context compaction and recovery, and your application supplies tools and picks an execution environment. On September 29 OpenAI added computer use, so agents can now work in an OpenAI-hosted browser too.

That is a different security model from the OpenAI Agents SDK, where the loop runs in your process and every tool call passes through your code. Three facts from OpenAI's documentation shape every buying decision below:

  • Retention. "The Agents API currently supports data residency in the United States only. It does not support Zero Data Retention, including when you use a self-hosted sandbox."
  • Remote MCP. HTTP MCP connections default to running from OpenAI's service, and OpenAI notes that "your application does not need to handle each call." It also does not get the chance to.
  • Confirmation. From the computer use guide: "Asking for confirmation through a function tool relies on the agent calling that function."

There is also a compliance date. On October 5, 2026 OpenAI added a self-serve BAA flow for API organizations. As of October 7, its list of HIPAA eligible API endpoints includes the Responses API and Chat Completions and does not list the Agents API. If your agent touches PHI, that settles the first question before any tool comparison starts.

Six options compared

OptionWhat it coversWhere it stopsBest for
OpenAI native controlsScoped keys, vaults, restricted network (1 to 100 exact hosts), allowed_tools, origin approval for the hosted browserNo per-call policy or approval on remote MCP toolsEvery team, as the baseline
Validation inside your function tool handlersAny check you write, on calls routed to your codeOnly function tools; MCP calls bypass itAgents whose write actions are all function tools
Agents SDK on a ZDR or BAA eligible endpoint insteadFull control of the loop, SDK guardrails, approvals in your processYou run orchestration, compaction and recovery yourselfPHI, ZDR contracts, strict data governance
Prompt and content guardrail productsInjection and sensitive data screening on textAnswers whether text is safe, not whether this call should runTeams with heavy untrusted input
Build your own MCP gatewayExactly the rules you wantApproval routing, audit storage and upkeep are yoursPlatform teams with spare capacity
AgentShield MCP gatewayPer-tool, per-argument policy on every MCP call, holds for a named approver, injection checks on inputs, audit trailDoes not replace OpenAI's data controls or make the API ZDR eligibleAgents with write tools on remote MCP servers

Option 1, OpenAI's native controls

Start here regardless of what else you buy. Give the application key only api.agents.read, api.agents.write and api.responses.write. Keep that key out of the sandbox, and give any self-hosted executor its own environment key, which OpenAI notes agent-generated code can read but which can do nothing except connect environments. Put third-party credentials in vaults, where sandbox code sees a placeholder and a proxy swaps in the secret only for approved hosts. Use restricted network mode and trim every MCP connection with allowed_tools.

What none of this gives you is a decision about a specific call. allowed_tools decides whether the agent can see issue_refund, not whether this refund, to this account, for this amount, should go through.

Option 2, checks in your function tool handlers

Function tools are the part of the Agents API that does pass through your code. The session emits agent.session.requires_action, your handler runs, and you return a result or an error. If every write action your agent can take is a function tool, validating arguments and holding risky calls in that handler is a sound design and costs nothing extra.

The trap is the hybrid build: write actions through MCP, plus a request_approval function the agent is told to call first. That function only runs when the model chooses to call it, which is exactly what a prompt injection in an email, ticket or uploaded file is designed to change.

Option 3, keep regulated workloads on the Agents SDK

This is the honest answer for a lot of US healthcare and financial services teams, and it means buying nothing. If a workload needs Zero Data Retention or falls under a BAA, the Agents API is not the place for it today. Run the loop with the Agents SDK against an eligible endpoint, where input, output and tool guardrails and per-tool approvals live in your own process. You give up the managed sessions and compaction. You keep a clean data governance story, which is what your auditor asks about first. Our healthcare AI agent security page covers the rest of that review.

Option 4, prompt and content guardrails

Guardrail products screen text for injection attempts, secrets and personal data. They are useful when agents read a lot of untrusted input, and several already integrate with OpenAI tooling. They answer one question well: is this text dangerous. They do not answer whether this agent should make this payment, which is a policy question about the call, not the words around it. Treat them as a layer, not the gate.

Option 5, build the gateway yourself

A gateway in front of your MCP servers that checks each call is not exotic engineering. The rules engine is the easy part. The work is in approval routing to the right person, timeouts and escalation, an audit store the agent cannot write, and keeping rules current as tools change. If your platform team has the capacity, build it. If it does not and you still want it in house, a custom software development company can scope it against your MCP inventory. Budget for owning it after launch.

Option 6, AgentShield as the MCP gateway

You point each Agents API MCP connection that can write at an AgentShield gateway endpoint instead of the raw server, and the gateway holds the real credential. Every call is evaluated against per-tool, per-argument permissions, risky calls wait for a named approver through a human approval gate, inputs are checked for injected instructions, and every decision lands in an AI agent audit trail outside the session. Because enforcement sits on the call path, it works whether or not the agent remembers to ask.

What it does not do: change OpenAI's retention, make the Agents API ZDR or BAA eligible, or secure the hosted browser. For the browser, follow OpenAI's own advice and restrict it to resources that cannot buy, submit or delete.

Which option fits your situation

SituationPickBuy anything new
Read only research, code review or report agentsOption 1No
PHI, ZDR contracts or strict retention rulesOption 3No
All write actions are function tools you ownOptions 1 and 2No
Agents read lots of untrusted email, tickets or filesAdd Option 4 to whatever else you chooseMaybe
Write tools on remote MCP servers, platform team has capacityOption 5Engineering time
Write tools on remote MCP servers, auditor asks who approvedOption 6Yes

Three of six situations buy nothing new. That is the right ratio for a beta product used mostly for research and coding today, and it will shift as more teams connect production systems.

Questions buyers ask

Does the OpenAI Agents API support Zero Data Retention?

No. OpenAI states the Agents API supports data residency in the United States only and does not support Zero Data Retention, even with a self-hosted sandbox, because session state is retained so work can continue across turns. You can delete sessions and published artifacts, but that is deletion on request, not zero retention.

Is the OpenAI Agents API covered by a HIPAA BAA?

Not as of October 7, 2026. OpenAI's HIPAA eligible endpoint list covers the Responses API, Chat Completions and other endpoints, and does not list the Agents API. Keep PHI out of Agents API sessions and run those workloads on an eligible endpoint under an executed BAA.

Can I add human approval to OpenAI Agents API tool calls?

For function tools, yes, because they pause for your code. For remote MCP tools there is no native per-call approval, and OpenAI notes that confirmation through a function tool depends on the agent calling it. A gateway on the MCP call path is how you make approval mandatory.

What does the OpenAI Agents API cost?

There is no separate fee. Model usage is billed at the chosen model's API rates, OpenAI tools at their standard rates and hosted sandboxes at container rates.

Next step

If your agents only read, configure OpenAI's controls and stop there. If they write through MCP, read our OpenAI Agents API security page for the full control map and checklist, and compare it with the Agent Builder migration page if you are moving workflows off Agent Builder before its November 30, 2026 shutdown. Then create your workspace and put the gateway in front of one write tool first.

See the firewall block an attack live.

Drive the Threat Console and watch a real prompt injection get stopped, then put AgentShield in front of your own agents.

Open the console