Field notes · AI security
← WritingLLM guardrails fail silently — measure them
Grey Ridge Signals Group · August 2026
LLM guardrails fail silently — measure them
Grey Ridge Signals Group · August 2026
A guardrail that blocks 90.9% of attacks while also blocking 2.6% of benign traffic is a liability, not a defense — and the regression went unnoticed until we measured it at scale. This is how we put false positives under permanent, automated measurement, and what the numbers caught.
1. Guardrails fail silently
The failure mode nobody demoes is the false positive. A prompt filter that stops attacks but also kills legitimate requests is worse than no filter at all once you run it at volume: every false positive is a benign request blocked or flagged, and at triage scale a 2.6% false-positive rate is not a rounding error — it is a standing interruption to your own traffic.
The costs compound in ways that don't show up in a demo:
- Desensitization. Operators learn that the guardrail cries wolf. Alerts get ignored, exceptions get rubber-stamped, and the defense becomes background noise.
- Budget erosion. In a red-team or triage pipeline, false positives burn the same review hours as real detections. An FP-heavy guardrail spends your team's attention on itself.
- Silent regressions. Raw "attacks blocked" numbers hide all of this. "Blocked 90% of attacks" is compatible with both a scalpel and a brick wall. If you only measure true positives, a guardrail can regress on false positives and you will ship the regression without noticing.
That last point is not hypothetical. We measured it.
2. The setup: measurement as a permanent process
Our adversarial eval harness already scored attacks. What it did not do, until recently, was score benign traffic — and it only ran when we remembered to run it.
That changed in August 2026. The weekly benchmark — three fleet models (qwen3:8b, qwen2.5:14b-instruct, llama3.1:8b) across seven engines — now runs every Sunday at 02:00 with a control corpus wired in. Each engine's benign control set is scored alongside the attack sets, and true-positive rate, false-positive rate, and F1 are computed from the results.
The important part is the tripwire: a fail-closed FPR gate sits at the end of the run. If any gated engine's false-positive rate exceeds 2.0%, the run exits non-zero and the publish step is skipped — a regression can no longer reach the site silently. If the metrics cannot be computed at all, the gate fails closed rather than guessing. One-off audits are nice; a gate that runs every week is what actually holds the line.
3. What the measurement caught
The first thing the control corpus surfaced was real. At scale on the modern engine (115 benign controls, 2026-08-10), our llm-guard defense showed a false-positive rate of 2.61% — three benign prompts, benign-033/038/043, blocked. A sibling defense (seal-epd-llm) held 0.0% on the same set. That single comparison changed the verdict: llm-guard was not safe to promote as a default at scale until the FP problem was fixed.
The fix landed as a shape pre-filter: a whitelist plus a 300-character cap plus a directive denylist, verified against the attack library (0 of 172 attacks matched the whitelist — the filter could not be used to smuggle attacks through). Acceptance was FPR=0.0 at 115 controls with true-positive rate held at 90.9%.
Then came the re-measurement at scale, on the first gated weekly run (2026-08-12):
- 115 benign controls, zero false positives
- True-positive rate 90.9% (attack detection unchanged)
- F1 0.952, discrimination 90.9
- Gate verdict: PASS — publish proceeded
The corpus did not stay at 115. The at-scale re-measurement — the control corpus grown to 205 file-based prompts plus 15 built-ins, 220 controls, published 2026-08-19 — measured FPR 0.91% (2/220, benign-164/179) with TPR held at 90.9%. Still PASS under the fail-closed 2% ceiling: the zero was a 115-control result, and 0.91% is the number the tripwire now holds the line against.
The regression that would have shipped silently was caught, fixed, and permanently wired into the tripwire. That is the entire argument for measuring false positives automatically: not because we expected a problem, but because the measurement found one we had not seen.
4. The cross-model check
A zero-FP result on one model could be a coincidence of that model's tokenizer or refusal style. So we ran the same 115-control measurement on the other two canonical fleet models on 2026-08-13:
- llama3.1:8b: FPR=0.0, TPR 90.9%, F1 0.952, zero false positives
- qwen2.5:14b-instruct: FPR=0.0, TPR 90.9%, F1 0.952, zero false positives
Identical numbers across all three models. The zero-FP result is a property of the defense on our harness, not an artifact of one model's behavior. And because the sweep is weekly, this is not a one-time photograph: the 2026-08-16 run records the first full-matrix verdict (all models × all gated engines), and every Sunday after that re-verifies it automatically.
5. What we changed because of it
Three permanent changes came out of this:
- FPR is now a first-class metric in every defense verdict we publish. A raw "blocked X% of attacks" without a false-positive figure is not a verdict — it is a number that cannot distinguish a filter from a firewall.
- The gate is permanent and automatic. Defense promotion is conditioned on it. A guardrail earns default status by passing FPR at scale, repeatedly, and loses it the moment a weekly run trips the wire.
- We now give clients the same instrument, not the conclusion. If you run a guardrail and you do not measure false positives on your own traffic, you do not know what you are running. The control-corpus methodology is portable; we hand it over with every defense assessment.
6. Honest limitations
FPR=0.0 was a 115-control measurement on our harness and control corpus; re-measured at 220 controls it came back 0.91% (2/220) — still inside the fail-closed 2% gate. Neither number is an absolute guarantee. Your traffic differs, and your false-positive rate will differ; the point is that you can now measure yours.
The cross-model result covers the modern engine and the llm-guard defense; the first full-matrix verdict lands with the 2026-08-16 sweep. And the gate is only as good as its corpus — the control set has to evolve with the attack library, which is exactly why the measurement runs weekly rather than once.
The defense that holds 90.9% of attacks with a 0.91% false-positive rate at 220 controls is the claim we can defend. The defense we merely believed in — until the numbers said otherwise — is the one this whole exercise exists to prevent.
Grey Ridge Signals Group LLC provides AI security and security architecture advisory. If you found a hole in our methodology, we would rather hear it than publish next quarter's numbers on top of it.