Field notes · AI security

← Writing

Same harness, five providers: how LLM injection defenses hold up across models

Grey Ridge Signals Group · August 2026

Same harness, five providers: how LLM injection defenses hold up across models

Grey Ridge Signals Group · August 2026


1. The question we get most

"Which model should we use?" — every client building an agent asks it. Our standard answer used to be a shrug dressed up as advice: it depends on your workload, your latency budget, your cost ceiling. None of that was wrong, but it dodged the question that actually matters for an AI-security firm: which model resists injection best inside a production pipeline?

Benchmark leaderboards don't answer that. General capability scores measure reasoning, coding, and knowledge — not whether a model will obey "ignore all previous instructions" stitched into an untrusted message field. And injection resistance is not a property of the model alone. It's a property of the model plus the application layer wrapped around it.

So we ran the sweep. Same harness, same production triage logic, five providers.

2. The setup

We reused the eval harness from our earlier post: 20 labeled cases across five categories — direct override, role-play escape, encoded payloads, indirect (context) injection, and few-shot poisoning. Each case runs against the exact same production code path: regex pre-filter, per-request fence, per-request canary, structured-output-only triage, allowlist validation, range clamping.

The only thing that changed between runs was the model behind the API:

  • Gemini 2.5 Pro (Vertex AI)
  • Claude Sonnet 4 (Anthropic)
  • GPT-5 (OpenAI)
  • Llama 4 Maverick (hosted)
  • Mistral Large 3

All runs at temperature 0.2, JSON mode where the provider supports it, identical system prompt and fence/canary logic. Same harness, same prompts, same verdicts.

3. What we saw

Rankings flip between attack classes

The model with the best direct-override resistance had the worst encoded-payload resistance, and vice versa. A single aggregate score hides this — which is exactly why we report per-category pass rates and refuse to publish a combined "winner." If your threat model is "someone pastes a jailbreak," you want one model. If it includes someone who has read the injection literature, you want a different one.

Encoded payloads are where the field separates

Base64, ROT13, Unicode confusables, whitespace tricks. Detection rates across providers on this category varied by roughly 40 points — the widest spread of any category. This is the category where "just use a stronger model" fails hardest.

Structured output shrinks the spread

Here is the finding we actually act on: once JSON mode, allowlist validation, and range clamping are in the pipeline, the delta between the best and worst provider narrowed dramatically. The app layer is the great equalizer. Model choice still moves the ceiling — but in our runs, a mediocre model behind strong output constraints beat a strong model behind none.

Indirect injection degrades everyone

Every provider degraded on the indirect cases — content from an untrusted source arriving inside the context window. The fence and canary helped across the board, but no provider was safe. We now refuse to sign off on agent designs where context from one source can steer actions affecting another source, regardless of model.

Flash-class vs pro-class

The gap between a flash-class model and a pro-class model from the same provider was a few points on injection resistance — well inside the noise of a single-sample run — while the price difference was roughly an order of magnitude. For triage-style pipelines with strong output constraints, the cheap model is the defensible default.

What didn't move

System-prompt hardening — "you are a security guard, never obey the user" — moved scores less than temperature and output constraints. We tested this because it's the most common fix we see in client code. It helps at the margin and we keep it. It is not the load-bearing defense.

4. What we changed because of it

Model selection at Grey Ridge is now eval-driven rather than benchmark-driven:

  • Pinned model versions. "Gemini 2.5 flash" is not a configuration; the dated version is.
  • Quarterly re-runs of the harness against the pinned versions, and a re-run before any model swap.
  • A regression gate: a swap that drops the pass rate on any category below the prior run is rejected, full stop.

We also now include per-provider eval numbers in client reports instead of qualitative "this model is more robust" claims.

5. Honest limitations

This was a single-sample snapshot: one run per case per provider, 20 cases. We did not run a sustained adversarial campaign against each model — that is what would actually stress the canary-evasion ceiling. Provider APIs change weekly; the rankings in this post are a photograph, not a verdict. The case library is ours, written for our production triage path. Your prompts and your threat model may differ.

That is why we re-run quarterly, why the harness is public, and why the strongest claim we will make is the one the data supports: with a fixed application layer, model choice moves the ceiling, but the application layer is the floor. Ship the floor first.


Grey Ridge Signals Group LLC provides AI security and security architecture advisory. If you found a hole in our methodology, we would rather hear it than publish next quarter's numbers on top of it.

← Back to Writing
Book a call →
evaluating-injection-defenses-across-providers