Field notes · AI security

← Writing

How to scope an AI security review

Grey Ridge Signals Group · August 2026

How to scope an AI security review

Grey Ridge Signals Group · August 2026

The scope is the review. Before a single probe runs, someone decides what is in, what is out, and what "pass" means — and that decision, more than any test, determines whether the report tells you something real about your system. Most AI security reviews that disappoint do not fail during testing; they fail at the scoping call, when "review our AI" was accepted without asking what "AI" means here.

This is the checklist we use when a client asks us to look at their system. It is written from the review side, but it is meant to be read from the buyer side: if you can produce a scope that answers these six questions, you already understand your system well enough to make the review worth paying for.

1. Draw the boundary

A scope begins with an in/out list. For an agentic system the in-list is usually the agent itself, the tools it can call, the data it can touch, the memory it reads and writes, and the human-in-the-loop gates between it and consequential actions. The out-list matters just as much: the hosting platform, the vendor APIs, the business processes no agent reaches. An unbounded scope is not thorough, it is shallow — the budget gets spent on the wrong surface and the report proves nothing about the parts that matter.

2. Name the threat model

Who attacks, what do they want, and what can they touch? For agentic systems the interesting answers are rarely "random internet users." They are: a user who wants the agent to exceed its mandate, a compromised tool the agent trusts, poisoned memory that steers later behavior, a delegated sub-agent whose output becomes trusted instruction. If the scope cannot name the attackers and the capabilities they would need, the review will test the wrong things — or test everything superficially. (We wrote the patterns up separately, in our agentic-system review methodology.)

3. Choose the test classes

The tests follow from the threat model. Our own eval harness runs 20 labeled cases across six classes — injection, gaming, exfiltration, false-positive probes, spam, and edge inputs — because a review that only tests "can the attacker override the instructions" misses the gaming case where the input simply claims to be the result, or the exfiltration case where the model is asked to leak a secret. Your scope should list the failure classes on the table. It should also name the ones that are off the table, so nobody pretends they were covered.

4. Write the acceptance bar as numbers, before the work starts

This is the step most scopes skip, and the one that separates a review from a report. Our latest harness run: 95% overall case pass (19 of 20), 89% injection detection, 90% classification integrity, 100% exfiltration stripped, 100% spam accuracy — and a 25% false-positive rate on the one benign probe deliberately built to trip the cheap filter. That last number is the point: the false-positive probe exists to measure the known cost of a fast heuristic, and whether that cost is acceptable is a scoping decision the buyer — not the tester — should make. A guardrail that catches 89% of injections while flagging a quarter of adversarial benign text is a different product from one that trades detection for fewer false alarms. If the acceptance bar is not written down before testing, the review's quality gets negotiated after the fact, which is how "thorough" comes to mean "whatever we found."

5. Write the constraints

A scope without constraints is either reckless or toothless. Every AI system with real capability can take consequential actions; the review must say, up front, what the testers may and may not do: no writes to production data, no real user records, no payments, no bookings, nothing a monitored human would not approve. When we wired our own lead pipeline toward a booking system, the automation ran for weeks with the booking branch env-gated, unable to create a real event until a human flipped a single variable — the constraint was the feature. A review of your system should be equally explicit about what happens if a probe "succeeds": the report names the change that closes it, not a body count.

6. Agree on the deliverable

A scope names the output: findings tied to realistic threats, each paired with the specific design change that closes it — not a list of theoretical risks. It names the format, the evidence standard (a probe that reproduces the issue, not a screenshot of a prompt), and the shape of the engagement: fixed scope, senior-led, time-boxed. And it names the follow-up: which findings are fix-now, which are accepted risk, and who decides.

Red flags

Three scope requests should make a reviewer nervous. "Review everything" — no boundary, so nothing gets depth. "Find all vulnerabilities" — no acceptance bar, so the review can only stop, never finish. And "prove we're safe" — a scope written for a rubber stamp, which usually comes with constraints that forbid the questions that matter.

Why this matters before a call

Our own intake pipeline is a scoping filter. The triage system only qualifies inquiries that are actually in scope for AI security work: a dental group with a $10k budget and a two-week timeline was classified as nurture — outside our specialization — while a CTO asking for an LLM security review at $25k scored qualified. That is the same decision this article is about, applied at the door: the system decides what it will and won't review before a human ever talks to you. If you cannot yet write the boundary, threat model, and acceptance bar for your own system, that is precisely the conversation worth having. Book an intro call and we will scope it together — the scope is the deliverable, and the call is where it gets made.

← Back to Writing
Book a call →
how-to-scope-an-ai-security-review