Writing · Field notes

Field notes

Case studies and notes on AI security, drawn from the systems we build and run. No think-pieces — the work, and what it taught us.

Aug 2026

Operations · Measurement

Our first month of leads was 100% test traffic — how to tell a probe from a prospect

We published 27 leads, 13 qualified, 0 booked — then audited the funnel and found every lead was a test address. The probe-contamination audit: how E2E discipline pollutes its own metrics, how to fingerprint test traffic, and what to ask any vendor quoting pipeline numbers.

Read →
Aug 2026

Agent security · Methodology

How to scope an AI security review

The scope is the review — the boundary, threat model, test classes, and acceptance bar decided before any probe runs. The six-question checklist we use when a client asks us to look at their system, grounded in what our own harness actually measures.

Read →
Aug 2026

Operations · Measurement

27 leads, 13 qualified, 5 nurtured, 0 booked: the funnel between a website and a call

Our monthly metrics line for August: 27 leads, 13 qualified, 5 nurture sends, 0 bookings. What each funnel stage actually does, why the measurement had to be fought for, and why the last stage is the only one that can't be automated.

Read →
Aug 2026

Operations · Lead lifecycle

The lead nobody books gets a drip, not a dead end

The lead that never qualifies used to fall into a silent dead end — classified, logged to nothing, never contacted. The nurture drip replaced it with a timed follow-up loop: log, wait 48 hours, send, and mark.

Read →
Aug 2026

Agent security · Operations

An agent that books meetings needs a human in the loop

Our lead workflow could have created a real Cal.com booking for every qualified lead on a single env flip — no dedupe, no confirmation. The dry run that proved the path, the guardrail that gated it, and the tests that pinned it.

Read →
Aug 2026

Security engineering · Measurement

LLM guardrails fail silently — measure them

A guardrail that blocks 90.9% of attacks while also blocking 2.6% of benign traffic is a liability, not a defense. How we put false-positive rates under permanent, automated measurement — and what the numbers caught.

Read →
Aug 2026

Benchmarks · Injection defense

Same harness, five providers: how LLM injection defenses hold up across models

"Which model should we use?" — every client building an agent asks it. We ran the sweep: same production harness, five providers, five different injection resistance profiles.

Read →
Aug 2026

Cost engineering · Serverless

Serverless cost engineering for AI workloads — what we spend and why

Our AI features run on Cloudflare Pages/Workers with Vertex AI on GCP. The cost model behind that choice: tokens, not requests; system prompt dominance; and why serverless wins at our volume.

Read →
Jun 2026

Case study · Prompt injection

How we built a prompt-injection-hardened AI receptionist on Cloudflare Pages + Vertex AI

We dogfooded the attack we sell defense against: the threat model, the layered defenses, and an honest account of what we deliberately did not build.

Read →
Jun 2026

Security engineering · Eval harness

Building a reproducible adversarial evaluation harness for LLM injection defenses

How we built a portable, parameterized eval harness that runs production triage logic against 20 labeled cases and turns "hardened" into a number you can reproduce.

Read →
Jun 2026

Agent security · Methodology

Patterns and pitfalls in agentic AI security — reviewing systems that act

A practical framework for agentic AI security review: tool-abuse patterns, context poisoning in multi-turn sessions, and why authorization must live outside the model.

Read →
Jun 2026

Security engineering · Cryptographic provenance

The Seal VPE protocol — cryptographic provenance for AI agent prompts

Seal’s Verifiable Prompt Envelope (VPE) protocol replaces linguistic injection detection with cryptographic prompt signing. How it works, why it matters, and what we’ve learned building it.

Read →

More in progress — published as work becomes ready.

Ready to discuss your systems?

Book a call →