Writing · Field notes
Field notes
Case studies and notes on AI security, drawn from the systems we build and run. No think-pieces — the work, and what it taught us.
Operations · Measurement
Our first month of leads was 100% test traffic — how to tell a probe from a prospect
We published 27 leads, 13 qualified, 0 booked — then audited the funnel and found every lead was a test address. The probe-contamination audit: how E2E discipline pollutes its own metrics, how to fingerprint test traffic, and what to ask any vendor quoting pipeline numbers.
Read → Aug 2026Agent security · Methodology
How to scope an AI security review
The scope is the review — the boundary, threat model, test classes, and acceptance bar decided before any probe runs. The six-question checklist we use when a client asks us to look at their system, grounded in what our own harness actually measures.
Read → Aug 2026Operations · Measurement
27 leads, 13 qualified, 5 nurtured, 0 booked: the funnel between a website and a call
Our monthly metrics line for August: 27 leads, 13 qualified, 5 nurture sends, 0 bookings. What each funnel stage actually does, why the measurement had to be fought for, and why the last stage is the only one that can't be automated.
Read → Aug 2026Operations · Lead lifecycle
The lead nobody books gets a drip, not a dead end
The lead that never qualifies used to fall into a silent dead end — classified, logged to nothing, never contacted. The nurture drip replaced it with a timed follow-up loop: log, wait 48 hours, send, and mark.
Read → Aug 2026Agent security · Operations
An agent that books meetings needs a human in the loop
Our lead workflow could have created a real Cal.com booking for every qualified lead on a single env flip — no dedupe, no confirmation. The dry run that proved the path, the guardrail that gated it, and the tests that pinned it.
Read → Aug 2026Security engineering · Measurement
LLM guardrails fail silently — measure them
A guardrail that blocks 90.9% of attacks while also blocking 2.6% of benign traffic is a liability, not a defense. How we put false-positive rates under permanent, automated measurement — and what the numbers caught.
Read → Aug 2026Benchmarks · Injection defense
Same harness, five providers: how LLM injection defenses hold up across models
"Which model should we use?" — every client building an agent asks it. We ran the sweep: same production harness, five providers, five different injection resistance profiles.
Read → Aug 2026Cost engineering · Serverless
Serverless cost engineering for AI workloads — what we spend and why
Our AI features run on Cloudflare Pages/Workers with Vertex AI on GCP. The cost model behind that choice: tokens, not requests; system prompt dominance; and why serverless wins at our volume.
Read → Jun 2026Case study · Prompt injection
How we built a prompt-injection-hardened AI receptionist on Cloudflare Pages + Vertex AI
We dogfooded the attack we sell defense against: the threat model, the layered defenses, and an honest account of what we deliberately did not build.
Read → Jun 2026Security engineering · Eval harness
Building a reproducible adversarial evaluation harness for LLM injection defenses
How we built a portable, parameterized eval harness that runs production triage logic against 20 labeled cases and turns "hardened" into a number you can reproduce.
Read → Jun 2026Agent security · Methodology
Patterns and pitfalls in agentic AI security — reviewing systems that act
A practical framework for agentic AI security review: tool-abuse patterns, context poisoning in multi-turn sessions, and why authorization must live outside the model.
Read → Jun 2026Security engineering · Cryptographic provenance
The Seal VPE protocol — cryptographic provenance for AI agent prompts
Seal’s Verifiable Prompt Envelope (VPE) protocol replaces linguistic injection detection with cryptographic prompt signing. How it works, why it matters, and what we’ve learned building it.
Read →More in progress — published as work becomes ready.
Ready to discuss your systems?
Book a call →