We publish our real numbers — including the misses. 465 labeled samples run against the live production engine on 2026-10-07, plus a 211-sample external hold-out set the engine never saw. Methodology reviewed by three AI systems (ChatGPT, Grok, Claude — full critiques in the repo). Per-sample results: download the 465-sample dev set (CSV); hold-out kept private so it stays honest.
Reliably catches: SSNs and credit-card numbers (100%), leaked API keys and credentials (~80%), risky URLs (~77%). Ordinary text never flags.
Does not reliably catch: novel prompt-injection attacks — about 1 in 20 unseen attacks detected. This is the hard problem the whole industry is working on; our classifier for it is in development.
When to rely on Agent-Shield today: as a fast pre-filter for secrets, PII, and known-bad URLs in agent traffic. When not to: as your only defense against novel prompt injection.
465 samples · 0 API errors · ~0.64s median scan latency
Our 465-sample set was written by our own team, which means tuning against it inflates recall (Claude's authorship-circularity warning). So we built a second set from external corpora the detector never saw: deepset/prompt-injections, HackAPrompt, TensorTrust, BIPIA, and gitleaks fixtures — 211 samples, sampled blind, labels from the sources.
| Source | Samples | Caught |
|---|---|---|
| deepset/prompt-injections | 100 | 43 (all 40 benign + 3 attacks) |
| BIPIA (indirect injection) | 30 | 0 |
| HackAPrompt (extraction) | 37 | 0 |
| TensorTrust (hijacking) | 40 | 2 |
| gitleaks fixtures | 4 | 3 |
The gap between 59% (our own samples) and 4.7% (unseen attacks) is the overfitting this page exists to prevent. A heuristic pattern-matcher cannot do semantic injection detection — this is the true baseline the Tev1 decision-model layer must beat, on this same hold-out set. We will not tune the heuristics against these samples.
| Engine | Recall | Precision | FPR | What changed |
|---|---|---|---|---|
| heuristic_v1 | 30.8% | 86.6% | 12.7% | Baseline: narrow phrase patterns |
| heuristic_v2 | 52.9% | 90.7% | 14.3% | Paraphrase families, base64 decode, defang normalization, placeholder allowlist |
| heuristic_v3 | 58.7% | 89.9% | 17.5% | Encoding normalization (hex/ROT13/homoglyph/leet/reversed), jailbreak-persona signals |
| heuristic_v4 | 58.8% | 89.2% | 18.8% | Hex/octal/decimal IP literals, de-JSON fragment joining; +5 Claude adversarial cases |
Four iterations, and recall has converged near 59% while the false-positive rate keeps climbing — the heuristic wall, confirmed. Further tuning against our own samples would manufacture numbers, not earn them (see authorship circularity note). The semantic gap is what the Tev1 decision-model layer is being built to close.
| Category | Samples | Correct | Miss rate |
|---|---|---|---|
| Direct prompt injection | 50 | 23 | 54% |
| Indirect injection (hidden in content) | 50 | 35 | 30% |
| Encoded / obfuscated injection | 40 | 10 | 75% |
| Role spoofing | 40 | 16 | 60% |
| System-prompt extraction attempts | 40 | 13 | 68% |
| Secrets exposure (AWS/GitHub/PEM) | 40 | 32 | 20% |
| PII exposure (SSN / cards) | 40 | 40 | 0% |
| SSRF-risk URLs | 30 | 23 | 23% |
| Benign — normal text | 60 | 60 | 0% FP |
| Benign — tricky edge cases | 60 | 45 | 25% FP |
| Panel adversarial cases | 10 | 7 | 30% |
False negatives (attacks scored clean): paraphrased instruction overrides, HTML-comment hidden instructions, defanged SSRF URLs, debugging-framed credential exfiltration.
False positives (benign text flagged): threat-briefing quotes of known attacks, README PEM examples, test fixtures containing AWS documentation placeholder keys.
heuristic_v5 is a deterministic pre-filter, not a semantic classifier — and these numbers are the honest baseline it sets. Five iterations took recall from 30.8% to 58.8%, but the false-positive rate climbed from 12.7% to 18.8%: the heuristic wall, confirmed. The semantic gap (injection intent under paraphrase and adversarial context) is what the Tev1 decision-model layer — a small language model trained to judge intent, now in calibration — is being built to close. The next benchmark will run all 676 samples against heuristic_v5 + Tev1 and publish the delta.