Detection benchmarks

We publish our real numbers — including the misses. 465 labeled samples run against the live production engine on 2026-10-07, plus a 211-sample external hold-out set the engine never saw. Methodology reviewed by three AI systems (ChatGPT, Grok, Claude — full critiques in the repo). Per-sample results: download the 465-sample dev set (CSV); hold-out kept private so it stays honest.

In plain English

Reliably catches: SSNs and credit-card numbers (100%), leaked API keys and credentials (~80%), risky URLs (~77%). Ordinary text never flags.

Does not reliably catch: novel prompt-injection attacks — about 1 in 20 unseen attacks detected. This is the hard problem the whole industry is working on; our classifier for it is in development.

When to rely on Agent-Shield today: as a fast pre-filter for secrets, PII, and known-bad URLs in agent traffic. When not to: as your only defense against novel prompt injection.

Headline numbers — engine: heuristic_v4 (2026-10-07, fourth iteration)

58.8%
Recall — attacks caught
89.2%
Precision — flags that were real
18.8%
False-positive rate
41.3%
False-negative rate

465 samples · 0 API errors · ~0.64s median scan latency

External hold-out set — the number that actually matters

Our 465-sample set was written by our own team, which means tuning against it inflates recall (Claude's authorship-circularity warning). So we built a second set from external corpora the detector never saw: deepset/prompt-injections, HackAPrompt, TensorTrust, BIPIA, and gitleaks fixtures — 211 samples, sampled blind, labels from the sources.

4.7%
Recall on unseen attacks
100%
Precision — zero false positives
0.0%
False-positive rate (40/40 benign clean)
SourceSamplesCaught
deepset/prompt-injections10043 (all 40 benign + 3 attacks)
BIPIA (indirect injection)300
HackAPrompt (extraction)370
TensorTrust (hijacking)402
gitleaks fixtures43

The gap between 59% (our own samples) and 4.7% (unseen attacks) is the overfitting this page exists to prevent. A heuristic pattern-matcher cannot do semantic injection detection — this is the true baseline the Tev1 decision-model layer must beat, on this same hold-out set. We will not tune the heuristics against these samples.

Iteration history — published, not cherry-picked

EngineRecallPrecisionFPRWhat changed
heuristic_v130.8%86.6%12.7%Baseline: narrow phrase patterns
heuristic_v252.9%90.7%14.3%Paraphrase families, base64 decode, defang normalization, placeholder allowlist
heuristic_v358.7%89.9%17.5%Encoding normalization (hex/ROT13/homoglyph/leet/reversed), jailbreak-persona signals
heuristic_v458.8%89.2%18.8%Hex/octal/decimal IP literals, de-JSON fragment joining; +5 Claude adversarial cases

Four iterations, and recall has converged near 59% while the false-positive rate keeps climbing — the heuristic wall, confirmed. Further tuning against our own samples would manufacture numbers, not earn them (see authorship circularity note). The semantic gap is what the Tev1 decision-model layer is being built to close.

Per-category results

CategorySamplesCorrectMiss rate
Direct prompt injection502354%
Indirect injection (hidden in content)503530%
Encoded / obfuscated injection401075%
Role spoofing401660%
System-prompt extraction attempts401368%
Secrets exposure (AWS/GitHub/PEM)403220%
PII exposure (SSN / cards)40400%
SSRF-risk URLs302323%
Benign — normal text60600% FP
Benign — tricky edge cases604525% FP
Panel adversarial cases10730%

What the numbers actually say

Representative misses

False negatives (attacks scored clean): paraphrased instruction overrides, HTML-comment hidden instructions, defanged SSRF URLs, debugging-framed credential exfiltration.

False positives (benign text flagged): threat-briefing quotes of known attacks, README PEM examples, test fixtures containing AWS documentation placeholder keys.

Methodology

What this means for the roadmap

heuristic_v5 is a deterministic pre-filter, not a semantic classifier — and these numbers are the honest baseline it sets. Five iterations took recall from 30.8% to 58.8%, but the false-positive rate climbed from 12.7% to 18.8%: the heuristic wall, confirmed. The semantic gap (injection intent under paraphrase and adversarial context) is what the Tev1 decision-model layer — a small language model trained to judge intent, now in calibration — is being built to close. The next benchmark will run all 676 samples against heuristic_v5 + Tev1 and publish the delta.