Block the attack before your agent acts on it
Your agent reads documents, tool results and messages from other agents. SafePrompt checks every one of them and blocks the attacks, in one call, whatever model you run.
Key points
SafePrompt checks any text your agent is about to act on: a tool input, a retrieved document, or a message from another agent. One POST to /api/v1/validate returns JSON with safe true or false, a confidence score, and the threat categories that fired. Most checks come back in under a second. It works with any model provider, because it validates the content rather than the model.
Enter your email and the next screen shows your key. No card.
10,000 free validations a month on the free plan.
SafePrompt checks every input, not just the first one
A chatbot reads one thing: the user's message. An agent reads many things it did not write. It pulls documents from a vector store, calls tools over MCP, browses pages and passes results between sub-agents. Every one of those is a place an instruction can arrive, and that is where modern attacks land. SafePrompt drops in at each checkpoint, so the whole pipeline is covered rather than the front door.
What are the three ways an agent gets injected?
SafePrompt reads the content at all three, so a hidden instruction has to get past a check before it can become an action.
Indirect injection
A retrieved document, web page or email carries hidden instructions that your agent reads as if you typed them. The user prompt looked clean. The data did not.
Tool poisoning
An MCP tool description or a tool result smuggles redirect instructions into the loop, steering the agent toward an action you never approved.
Cross-context bleed
An injection that enters through one tool or sub-agent rides along into the next call, so a single poisoned step can reach the whole chain.
How it scores against the benchmarks
We ran a public agent benchmark against SafePrompt, and no injection reached its goal: 0 of 249 attacks across the banking and slack suites got through, against 40.3% and 60.0% with no defence. Screening tool output has a cost you can measure too: under attack the agent finished 28.5% of its banking tasks and 16.2% of its slack tasks, against 42.4% and 53.3% undefended.
| Agent attacks that never reached their goal (AgentDojo, banking and slack) | 100% |
| Direct hijack attempts caught (TensorTrust) | 91% |
| Jailbreak attempts caught (HackAPrompt) | 93% |
| Ordinary messages passed (TensorTrust benign set, 200 access codes) | 97% |
| Ordinary messages passed (BIPIA, documents and retrieved content) | 96% |
Strict mode, public sets, our own runs, 12 September 2026. AgentDojo: GPT-4o-mini agent, SafePrompt checking tool output.
The suite is in the public repository, with every run and every failed case published.
Security facts last reviewed: 12 September 2026
Install it in one function
SafePrompt is a single HTTP call between your orchestrator and tool execution. Check the tool input before the agent acts on it, and the same call covers retrieved chunks and inter-agent messages.
// Check anything your agent is about to act on
async function isSafe(content, userIp, sessionToken) {
const res = await fetch('https://api.safeprompt.dev/api/v1/validate', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'X-API-Key': process.env.SAFEPROMPT_API_KEY,
'X-User-IP': userIp, // End user's IP address, required
},
// session_token ties the steps of one conversation together, so an
// attack split across turns is read with the earlier turns attached
body: JSON.stringify({
prompt: content,
sensitivity: 'strict',
session_token: sessionToken,
}),
})
const { safe } = await res.json()
return safe
}
// In your MCP tool handler, check the input first
if (!(await isSafe(toolInput, userIp, sessionToken))) {
throw new Error('Prompt injection detected in tool input')
}It travels with your agent, across every provider
SafePrompt validates the content, not the model, so one checkpoint sits in front of every step, whoever serves the tokens. Your pipeline can change providers without changing your security layer.
Use Azure Prompt Shields if your stack lives entirely in Azure. Use Bedrock Guardrails if everything runs in Bedrock. Use a model provider's own filter if one model serves your whole pipeline and the chat box is the only input you take. Agent pipelines usually span more than one of those. Our Azure Prompt Shields comparison sets the first of those choices next to ours.
Questions developers ask first
Does SafePrompt work with any MCP server?
Yes. SafePrompt sits between your orchestrator and tool execution, so it works with any MCP server, custom tool or agent framework. It validates the content rather than the model, which is why one checkpoint covers every step of the pipeline.
What does it cost in latency to check every tool input?
Most checks come back in under a second. The median and the 95th percentile are measured continuously on production traffic and published on this site, so you can size the check against your own agent loop before you add the call.
Does it work with LangChain, LlamaIndex and CrewAI?
Yes. SafePrompt is a single HTTP POST, so it drops into LangChain, LlamaIndex, CrewAI or any other agent framework. You check content at each checkpoint, whether that is a tool input, a retrieved chunk or an inter-agent message, and block when an injection is detected. The install example on this page sends strict sensitivity.
Secure your agents in one call
10,000 free validations a month, on the same detection engine every paid plan runs.
Enter your email and the next screen shows your key. No card.
Read the MCP guide · See what it costs · What SafePrompt covers