Mastra AI Security for Mastra Agents, MCP Tools and Guardrails After the npm Compromise
Mastra is the TypeScript agent framework many product teams picked in 2025 and 2026, and it ships real controls: tool approval, processors for prompt injection and PII, and a classifier policy gate. Most of them are opt-in, several default to failing open, and the June 17, 2026 npm compromise showed that the machines building Mastra agents hold exactly the keys an attacker wants. This page maps what Mastra enforces, what it leaves to you, and where a runtime policy layer earns its cost.
Direct answer
Mastra AI security is solid on primitives and permissive on defaults. Tools execute without approval unless you set requireApproval or requireToolApproval, model-backed guardrails other than ClassifierProcessor default to warn and continue when their model fails, and since release 1.71.0 tool calls start before the model finishes streaming. Pin packages after the June 2026 npm compromise, turn approval on for every tool that writes, and enforce per-tool policy outside the agent process.
Try it live
Watch AgentShield block an attack in real time.
Pick a scenario and drive the inspection lane yourself. No signup needed.
Run a request
Runs the live engine on your text. Nothing is stored, no account needed.
Inspection lane
INSPECTINGPolicy trace
High-risk action held for approval
Audit trail
- § · → → →
The risk
A US software company runs a Mastra support agent that can look up an account in HubSpot, issue a Stripe refund and send a follow-up email through Gmail, all wired through @mastra/connect. The team set requireApproval on the refund tool and added a PromptInjectionDetector. Then they moved the agent to a durable agent so runs survive redeploys, which means their approval function (approve refunds over 200 dollars only) could no longer be serialized, and every tool call started asking for approval. After a week of clicking approve on account lookups, someone set the flag to false. The detector, meanwhile, was on its default warn setting, so on the afternoon its model provider timed out, it logged the failure and let a pasted email with hidden instructions straight through.
How AgentShield handles it
Keep Mastra doing what it does well: agents, workflows, memory, typed tools, processors and tracing. Put AgentShield on the outbound path. Route the MCP servers and HTTP APIs your Mastra tools call through the AgentShield gateway, write one policy per tool (allow, deny, or hold above a value threshold) that works the same for regular, durable and stored agents, send held calls to a named approver outside the agent process, and record every decision in an audit trail your security team can export.
The controls
The controls that secure which tools a Mastra agent may call, with which arguments, and who signs off first.
What Mastra enforces for you, and what it leaves to your code
Mastra is an open-source TypeScript framework for agents, workflows, memory and RAG, built by Kepler Software, and its core is licensed under Apache 2.0. It gives you more security building blocks than most frameworks its age: tool-level and request-level approval, suspend and resume for human input, a family of input and output processors, and tracing. What it does not give you is a decision about which of those to turn on. Almost every control is a hook you wire, and the defaults favor an agent that keeps working over an agent that stops.
That is a reasonable design for a framework, and it is also why incidents in Mastra deployments rarely come from a bug in Mastra. They come from the gap between the controls that exist and the controls a given team switched on, then kept on after the first week of friction.
| Control | What Mastra provides | Default | What is left to you |
|---|---|---|---|
| Tool approval | requireApproval on a tool, requireToolApproval on a request, or a policy function per call | Off. Tools execute without approval | Choosing which tools pause, and who approves |
| Prompt injection detection | PromptInjectionDetector processor, LLM-backed, with block, rewrite and warn strategies | errorStrategy warn, so a failed check logs and continues | Setting strict where unchecked input must not proceed |
| Policy gate on content | ClassifierProcessor with typed classifier policies on input, output and streams | Fails closed since 1.69.0, with an opt-in warn mode | Writing the classifiers and thresholds |
| PII and moderation | PIIDetector and ModerationProcessor, both LLM-backed | Configurable per processor | Knowing that a model call now sits in your hot path |
| Spend limits | TokenCostControl with block or warn | Approximate, checked from asynchronously stored metrics | A hard cap outside the agent if overspend matters |
| Authorization inside tools | requestContext and workspace passed to tools and approval policies | None | Every permission check in every tool |
Read down the Default column and the pattern is clear. The one control that fails closed out of the box, ClassifierProcessor, only changed to that behavior in the September 23, 2026 release. Everything else needs a deliberate setting, and a setting is only as good as the next person who edits the config.
Mastra is also a common destination for chatflows being rewritten as code after the Flowise archive. While that port is under way, the old instance still holds live credentials; see keeping self-hosted Flowise safe until the migration lands.
The Mastra npm supply chain compromise and what it means for agent credentials
On June 17, 2026, an attacker who had taken over the npm account of a Mastra contributor with publish rights pushed new versions of more than 140 packages across the mastra and @mastra scopes. Microsoft Threat Intelligence published its analysis the same day and attributed the operation with high confidence to Sapphire Sleet, the North Korean group also tracked as BlueNoroff, which Microsoft had already tied to the April 2026 axios package compromise.
The mechanism was simple and effective. Each poisoned Mastra package added a dependency on easy-day-js at version range ^1.11.21, a typosquat of the popular dayjs library. That range resolved to 1.11.22, whose postinstall hook ran during installation whether or not your code ever imported it, disabled TLS certificate verification, contacted attacker infrastructure and pulled a second-stage payload for Windows, macOS or Linux. Reporting describes a publishing window of roughly 45 minutes.
| Item | Detail |
|---|---|
| Date | June 16 to 17, 2026, weaponized version published at 01:01 UTC on June 17 |
| Scope | More than 140 packages in the mastra and @mastra scopes |
| Entry point | A hijacked maintainer account with publish rights across the ecosystem |
| Malicious dependency | easy-day-js 1.11.22, pulled in through the range ^1.11.21 |
| Payload targets | Credentials, tokens and API keys, 166 crypto wallet browser extensions, browser history, host and process inventory |
| Attribution | Sapphire Sleet, per Microsoft, with high confidence |
Microsoft's own mitigation list is the right starting point: review dependency trees for affected @mastra packages, check node_modules and package-lock.json for easy-day-js, pin known-good versions, install with scripts disabled, and rotate any credentials, tokens or API keys that could have been exposed.
The part that is specific to agent frameworks is what "credentials" means on a machine that builds agents. A developer laptop or CI runner that installs Mastra typically also holds the model provider key, the Stripe or HubSpot token behind the agent's tools, database URLs and cloud credentials. An infostealer that lands there does not need to trick the agent at all. It walks off with everything the agent is allowed to do. That is the strongest argument for keeping long-lived write credentials out of the agent process entirely and letting a gateway hold them, so a stolen developer environment yields a token that can reach the gateway and nothing else.
Mastra added a related feature on September 24, 2026: environment() in @mastra/connect materializes Platform connection credentials into sandbox environments such as e2b, Modal, Daytona, Docker or a subprocess, with GitHub first, exported as GH_TOKEN and GITHUB_TOKEN. It is convenient, and it means code the agent runs inside that sandbox can read the token. Scope those tokens as narrowly as you would for an untrusted contractor.
Mastra tool approval, eager tool execution and durable agents
Mastra's approval model is well thought out. Set requireApproval: true on a tool and every call to it pauses. Set requireToolApproval: true on a stream or generate call and every tool call in that request pauses. The two combine with OR logic. The stream emits a tool-call-approval chunk carrying the tool name, call ID and arguments, and you call approveToolCall or declineToolCall. Approvals raised by subagents bubble up to the supervisor, through several levels of delegation.
The strongest option is a policy function: requireToolApproval can take an async function that sees the tool name, arguments, request context and workspace and decides per call. That is what lets you approve refunds under 200 dollars automatically and hold the rest. Three details decide whether any of this holds up in production.
- Durable and stored agents cannot use the function. Mastra's documentation is explicit: "Function-based requireToolApproval is only available on regular stream() / generate() calls. Durable agents and stored agents persist their options, and a function can't be serialized, so they accept only a boolean." Pass a function there and it falls back to requiring approval for every tool call. That is a safe fallback, and it is also the fastest route to approval fatigue, which is how approval gets switched off.
- Approval needs storage. Human in the loop uses snapshots to capture request state, and Mastra warns you will see a "snapshot not found" error without a storage provider. Suspended runs only survive restarts with persistent storage configured.
- Approval should be bound to the exact arguments. Mastra's docs recommend fingerprinting the tool name and arguments shown to the reviewer so a call cannot run under an old approval if its arguments drift, and storing that fingerprint in durable storage scoped to user, run, tool call and policy version. It is good advice, and it is code you write and maintain.
Release 1.71.0, published September 24, 2026, added eager tool execution: "Tool calls now start as soon as their own arguments are complete (instead of waiting for the model to finish streaming the whole step)." It is on by default and can be disabled per run. The release notes list what still waits for the model: tools that need approval, declare a suspend schema, run on the provider or client, or run in the background, plus runs using the called concurrency strategy or output processors that run after the stream.
Read that list as a security note. A write tool without approval now fires while the model is still generating the rest of its step, so any check you imagined happening at the end of the step is too late for it. If a tool can move money, send mail or change records, give it an approval flag or a gateway policy, not a post-hoc review.
| Setup | Per-argument policy (hold only above a threshold) | What happens in practice |
|---|---|---|
| Regular agent, policy function | Yes | Works as designed, if every write tool is covered |
| Durable or stored agent, boolean true | No | Every call pauses, reviewers approve everything, fatigue follows |
| Durable or stored agent, boolean false | No | Nothing pauses, eager execution runs write tools immediately |
| Any agent, policy enforced at a gateway | Yes, per tool and per argument | Same rule for regular, durable and stored agents |
The last row is the reason a separate layer exists. Our human approval for AI agents page shows what the approver sees and how timeouts are handled.
Mastra guardrails compared with a runtime security layer
Mastra calls its guardrails processors. Input processors run before the model sees a message, output processors run on what comes back, and hybrid processors can do both. The built-ins include UnicodeNormalizer, PromptInjectionDetector, LanguageDetector, SystemPromptScrubber, ModerationProcessor, PIIDetector, TokenCostControl and, since September 23, 2026, ClassifierProcessor for typed policy gates.
Two defaults matter more than the list. First, from Mastra's guardrails documentation: "Model-backed guardrail processors other than ClassifierProcessor default to errorStrategy: 'warn', which logs internal model failures and continues with the processor's fallback." In plain terms, if the model behind your injection detector is slow or down, the message goes through. Set strict where that is unacceptable. Second, TokenCostControl is honest about its limits: "Cost checks are approximate, metrics are persisted asynchronously, so fast-running agents may briefly exceed the configured limit before the guard triggers."
There is also a limit no processor can remove. Processors inspect text. A guardrail can tell you a message looks like an injection attempt; it cannot tell you whether this agent, acting for this customer, should issue this 900 dollar refund. That is an authorization decision, and it belongs where the tool call leaves your system.
| Need | Mastra processors and approval alone | With AgentShield added |
|---|---|---|
| Redact PII and system prompt leaks from responses | Yes, PIIDetector and SystemPromptScrubber | No change needed for this alone |
| Block unsafe content with a fail-closed gate | Yes, ClassifierProcessor since 1.69.0 | No change needed for this alone |
| Catch injected instructions in fetched pages, email and documents | PromptInjectionDetector, if set to strict and wired to that content | Inspection on retrieved content before the model reads it |
| Per-argument hold on durable and stored agents | Not available, boolean only | One policy per tool and value threshold at the gateway |
| Write credentials kept out of the build and agent environment | Not addressed | Gateway holds the credential, the agent holds a scoped token |
| Same controls across Mastra, LangChain and OpenAI agents | Mastra only | One policy for every framework calling the same tools |
Two of those six rows say buy nothing, and we mean it. A Mastra agent that answers questions from your docs, with no write tools and no inbound email, is well served by Mastra's processors with strict error handling. The case for a runtime layer starts when the agent can act.
Mastra Enterprise License and what the open-source core leaves out
Mastra is open core, and the line matters to security planning. The repository's license file states that everything under any directory named ee, including @mastra/core/auth/ee, @mastra/core/agent-builder/ee and @mastra/editor/ee, falls under the Mastra Enterprise Edition License, while the rest is Apache 2.0. Version 2.0 of that license took effect on September 22, 2026, and it is direct: "Source availability lets you read, build, modify, and test the code; it does not permit Production Use." Production use of an enterprise feature requires a written agreement with Kepler and a valid license key.
This is a normal open-core arrangement and we do not think it is a trap. It does mean a team planning to fork Mastra, or to run everything self-hosted on the open-source core, should check whether the auth features it is counting on sit inside an ee directory before it designs around them. The Apache 2.0 core is yours to keep; the enterprise auth module is not.
On the hosted side, Mastra's pricing page (checked September 30, 2026) lists a Teams plan at 250 dollars a month that adds multiple teams, SSO and SOC 2 documentation, and an Enterprise plan at custom pricing with RBAC, audit logs, support and uptime SLAs. Those are platform controls for who can operate agents and view traces. They do not decide whether a given tool call should run, which is the gap this page is about. Our tamper-evident audit trail records the decision on each call, and it sits alongside whatever observability you already use.
How to put AgentShield in front of Mastra tools and MCP servers
You keep writing Mastra agents and workflows the way you do now. The change is where outbound calls go and where write credentials live.
- Clean up after June 17 first. Confirm easy-day-js is absent from every lockfile, pin Mastra packages to known-good versions, install with scripts disabled in CI, and rotate the keys any affected machine held.
- Inventory tools that act. List every createTool definition, @mastra/connect provider and MCP server your agents can reach, and mark which ones write, send, pay or delete. Stripe, HubSpot, Gmail and GitHub providers all belong on that list.
- Route MCP and HTTP through the gateway. Point your Mastra MCP client at the AgentShield MCP gateway and your HTTP tools at the AgentShield proxy. Move the long-lived write credentials to the gateway so the agent process holds a scoped token.
- Write policy per tool. Allow reads broadly, cap writes by value, deny what an agent should never touch. Rules live in AI agent permissions management and apply the same way to regular, durable and stored agents.
- Keep requireApproval, and move the final yes. Mastra's pause and resume still work. Pair them with a gateway hold so a named person in your organization approves the high-value calls, and set PromptInjectionDetector to strict on inputs that carry outside content.
- Watch it live. Every verdict streams to agent monitoring next to your Mastra traces, so a burst of denied calls shows up as an alert rather than a line in a log.
Teams running Mastra next to Python agents should compare notes with our Pydantic AI security and LangChain security pages, since the same gateway policy covers all three. If you build MCP servers yourself, MCP server security covers the server side.
FAQ
Common questions about mastra ai security.
Is Mastra secure?
Mastra is a well-maintained framework with real security primitives: tool approval, suspend and resume, injection and PII processors, and a fail-closed classifier gate. Its defaults are permissive, though. Tools run without approval, most model-backed guardrails continue when their model fails, and eager tool execution starts calls before the model finishes. Secure deployments turn those controls on deliberately.
Is Mastra open source?
Mostly. The core framework is Apache 2.0. Anything under a directory named ee, including @mastra/core/auth/ee and the editor enterprise module, is covered by the Mastra Enterprise Edition License, which since September 22, 2026 lets you read, modify and test the code but requires a written agreement and license key for production use.
What happened in the Mastra npm supply chain attack?
On June 17, 2026, an attacker used a hijacked maintainer account to publish poisoned versions of more than 140 mastra and @mastra packages. Each added easy-day-js, a dayjs typosquat whose postinstall hook installed an infostealer. Microsoft attributed it to North Korean group Sapphire Sleet. Check lockfiles, pin versions and rotate exposed keys.
Does Mastra have guardrails?
Yes. Mastra calls them processors. Built-ins include PromptInjectionDetector, PIIDetector, ModerationProcessor, SystemPromptScrubber, TokenCostControl and ClassifierProcessor. Model-backed ones other than ClassifierProcessor default to errorStrategy warn, which logs a failure and continues, so set strict where unchecked content must not proceed. They inspect content, not whether an action is authorized.
Does Mastra support human in the loop?
Yes, in two ways. requireApproval or requireToolApproval pause a tool call until you call approveToolCall or declineToolCall, and a tool can call suspend() to ask for data and resume later. Both rely on snapshots, so configure persistent storage or suspended runs will not survive a restart.
How does Mastra tool approval work?
Mark a tool requireApproval: true, or pass requireToolApproval on a request, and matching calls pause with a tool-call-approval chunk showing the tool name and arguments. On regular agents requireToolApproval can be a function that decides per call. Durable and stored agents accept only a boolean, so per-argument rules need to live elsewhere.
Can Mastra agents use MCP servers safely?
Yes, with care. Mastra connects agents to MCP servers as tool sources, and those tools follow the same approval rules as any other tool. Mastra does not judge whether a server or its tools are trustworthy. Allowlist the tools each agent may call, require approval on the ones that write, and route calls through a gateway that logs every decision.
Does AgentShield work with Mastra?
Yes, on the outbound path. You route the MCP servers and HTTP APIs your Mastra tools call through AgentShield, which checks each call against a per-tool policy, holds high-value actions for a named approver, inspects retrieved content for injected instructions and records every decision. It works the same for regular, durable and stored agents.
More use cases