explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What a decision model for security does
  • The benchmark numbers, in context
  • Why you must tune the threshold on your own traffic
  • Where it fits in an agent defense
  • How to evaluate it this week
  • Limits to keep in mind
  • A worked example of threshold choice
  • Logging for audit and retuning
  • What this means for what you build or pay
  • Related reading
← Back to blog

explainx / blog

Security-One 27B: An Open Decision Model for Prompt-Injection and Security Triage

AI Security, Prompt Injection, Decision Models, Open Weights, Agents

Superagent released Security-One, a 27B Apache 2.0 decision model that scores prompts, tool calls and alerts. 599 of 600 injections caught, with real caveats.

Oct 6, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Security-One 27B: An Open Decision Model for Prompt-Injection and Security Triage

Prompt injection is the problem that does not go away. As agents read web pages, emails and repository files, any of that text can try to hijack them. The standard defenses are least privilege, sandboxing and human approval, and increasingly a screening model that looks at every input and tool call and flags the suspicious ones. On October 5, 2026, Superagent's Alan Zabihi released one as open weights: Security-One, a 27B decision model under Apache 2.0.

The headline number is striking, 599 of 600 prompt-injection attacks caught on one benchmark, and the less quoted numbers are more instructive. This post covers what the model is, how to use it, what the benchmarks say and do not say, and how to fit it into an agent defense without over-trusting it.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A 27B decision model returning unsafe probabilities for security-relevant events.
What can it screen?Prompts, agent tool calls, code changes and alerts.
License?Apache 2.0, weights on Hugging Face, plus an API.
Best number?599 of 600 attacks on BIPIA, 1 of 200 benign flagged, at threshold 0.70.
Weaker numbers?47 of 60 on Deepset; higher false positives than Jev on NotInject.
Can it enforce?No. The vendor says use it for screening.
Price?Not specified for the API in the coverage we reviewed.

What a decision model for security does

We explained the category in our decision models guide and covered a commercial example in Liquid AI's d1. A decision model answers a closed question with probabilities in one forward pass, with no generated text. For security triage that means you hand it an event and a question, such as "is this input an attempt to override the agent's instructions?", and get back a probability for each answer.

Security-One is built for always-on use. Because it generates no tokens, it is cheap and quick enough to run on every prompt, tool call or alert, not only on a sample. Superagent describes it as a routing layer that identifies suspicious activity and sends those events to human analysts or more powerful models for deeper investigation. It follows the "System One" decision-model idea, popularized by the Jev family, of fast intuitive classification with deliberate review reserved for exceptions.

The benchmark numbers, in context

Superagent reports results against public prompt-injection datasets. Here they are in one place.

table · 3 cols
DatasetResultNotes
BIPIA599 of 600 attacks detected (99.83%); 1 of 200 benign flagged (0.50%)Threshold 0.70
Deepset47 of 60 attacks detected (78.33%); 0 of 56 benign flaggedSame model, different dataset
NotInject0 of 339 false positives, but a 12.39% false-positive rate versus 2.36% for Jev in the vendor's comparisonOver-flagging benign text that looks suspicious

The spread across datasets is the real finding. A 99.83 percent catch rate on one benchmark and 78 percent on another says that performance depends on how attacks are written, and that no single number describes how it will behave on your traffic. Public datasets are also known to be imperfect proxies: attackers adapt, benchmarks age, and models can be tuned to them. The over-flagging result matters too, because a screening model that blocks benign requests teaches users to bypass it.

Why you must tune the threshold on your own traffic

The benchmark threshold is 0.70 unsafe probability. That is a trade-off, not a truth. Lower the threshold and you catch more attacks and flag more benign inputs. Raise it and the reverse happens. The right setting depends on:

  • the cost of a missed attack in your system (a read-only chatbot versus an agent that can send money),
  • the cost of a false alarm (a blocked customer versus a delayed review),
  • the base rate of attacks in your traffic, which is usually far lower than in a benchmark, so even a small false-positive rate can swamp your analysts.

Superagent provides a calibration recipe and says explicitly that thresholds must be tuned on your own data. Do that with a labeled sample that includes your real benign traffic, especially the odd-looking but legitimate inputs, like pasted logs or security discussions.

Where it fits in an agent defense

A screening model is one layer. A sensible stack looks like this:

  1. Least privilege. Give agents the narrowest credentials that work. A hijacked agent that cannot do anything dangerous is a nuisance, not an incident.
  2. Sandboxing. Run tools in isolated environments, as in our guides to agent sandbox isolation and the Codex Auto-review reviewer.
  3. Screening. Run Security-One, or a similar model, on inbound content and on each proposed tool call.
  4. Escalation. Send medium-probability events to a stronger model or a human, and block the highest ones.
  5. Logging. Record the probability and the decision, so you can audit and retune.

The reason to screen tool calls and not only inputs is that injections usually succeed through an action: an agent tries to send data somewhere or run a command. Catching the action is more robust than guessing every phrasing of the attack. For real-world examples of what goes wrong, see our coverage of prompt injection in GitHub agentic workflows and the rise of AI-driven web traffic.

How to evaluate it this week

  1. Collect two sets. Attack examples that resemble your threat model, and a large sample of benign inputs from your logs.
  2. Run Security-One and plot the distribution of unsafe probabilities for each set.
  3. Choose a threshold that meets your false-positive budget, then record the catch rate at that threshold.
  4. Red-team it. Try rephrased attacks, other languages, encoded text and multi-step instructions. Note which pass.
  5. Test tool-call screening by showing it proposed actions, not just prompts.
  6. Measure latency and cost in your serving setup. A 27B model needs a capable GPU, so compare with the API.
  7. Plan for drift. Re-run the evaluation periodically with fresh attacks.

Limits to keep in mind

  • It is a prediction model. Superagent itself says it can be confidently wrong.
  • Dataset-dependent performance. The gap between BIPIA and Deepset shows how uneven it can be.
  • Over-flagging. A higher false-positive rate on benign-but-odd text can hurt usability.
  • Unclear base model. The vendor post does not name the base model in the content we reviewed, so check the model card for lineage and licensing of the training data.
  • Adaptive attackers. Anyone who knows you use a given screener can test against it. Treat it as a speed bump, not a wall.
  • No pricing details for the hosted API in the coverage we reviewed.

A worked example of threshold choice

Suppose your agent processes 100,000 inbound items a day and real attacks are rare, say 20 a day. A screener with a 0.5 percent false-positive rate flags about 500 benign items daily, which means 25 false alarms for each real attack even if it catches every attack. If each review costs a minute, that is more than eight analyst-hours a day. Raise the threshold and the false alarms drop, but so does the catch rate. The point of the exercise is that base rates, not benchmark accuracy, drive operational cost. Measure your own base rate, set a review budget, and choose the threshold that fits it, then monitor drift as attackers adapt.

It also helps to separate tiers. Events above a high threshold can be blocked automatically, events in a middle band can be sent to a stronger model for a second opinion, and events below it can pass with logging. A three-tier policy gets more value from a cheap screener than a single cutoff does.

Logging for audit and retuning

Whatever threshold you choose, log the probability, the input hash, the decision and the eventual human verdict for escalated items. Those records let you measure real precision, find recurring false-alarm patterns and build a labeled set for retuning. They also give you evidence if an incident occurs and you need to show how the screening layer behaved. Mind privacy: store hashes or redacted excerpts when the inputs may contain personal data.

What this means for what you build or pay

If you ship agents that read untrusted content, a cheap open screening model is a reasonable addition, and an Apache 2.0 license means you can run it yourself and inspect the behavior. Budget the work to calibrate it, because an uncalibrated threshold is the most common way these tools disappoint. And keep the main defenses where they belong: scoped credentials, isolation and human review for irreversible actions. A good screener reduces the number of events that reach those layers. It does not replace them.

Related reading

  • What are decision models? AI classifiers guide
  • Liquid AI d1: a decision model with vision
  • Codex Auto-review: a reviewer agent for approvals
  • Prompt injection in GitHub agentic workflows
  • AI traffic overtakes humans and prompt injection
  • Agent sandbox isolation: five things to know
  • Perplexity's decision API

Primary: Superagent, "Introducing Security-One: a decision model for security" (superagent.sh) · HuggingNews and Datastudios coverage (October 5 to 6, 2026)

Details are accurate as of October 6, 2026 and come from vendor materials and press reports. Benchmark results are self-reported, thresholds must be tuned on your own traffic, and the model is a screening aid, not a security guarantee.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 26, 2026

The Provenance Tax: how LLM watermarking can break agent tool calls and refusals

Watermarks exist for provenance, but generation-time marks like SynthID-Text change which tokens get sampled — the same tokens agents use for tools and safety refusals. Lasso's September 2026 study reports sampling drift: up to ~17% paired disagreement on tool calls and higher attack success under a fixed prompt injection when watermark keys shift refusal behavior.

Sep 25, 2026

Meta Muse Filesystem Export: Feature, Not Breach, but Read the Fine Print

On September 24, 2026, The Verge reported that Meta Muse users could export large parts of its virtual machine, including system files and internal docs. Meta says that is intended behavior because each user gets their own Linux VM. Both sides are partly right, and the real security question is narrower than the headline.

Sep 19, 2026

Jev's Actual Security Use Case: Detecting Prompt Injection, Not Getting Hacked

There's no published adversarial research on gaming or poisoning Jev, TypeSafe AI's non-generative "System One Model" — a search for that angle comes up thin. What does exist is the inverse: Jev being positioned as a security tool itself, with a `contains_prompt_injection` classification primitive meant to sit in front of a main LLM and flag jailbreak or injection attempts fast and cheap, before they reach the model actually generating your response.