Field notes · AI security

← Writing

Serverless cost engineering for AI workloads — what we spend and why

Grey Ridge Signals Group · August 2026

Serverless cost engineering for AI workloads — what we spend and why

Grey Ridge Signals Group · August 2026


1. Why serverless

Our AI features — the receptionist and the adversarial eval harness — run on Cloudflare Pages/Workers with Vertex AI on GCP. No servers, no idle capacity, no 3 a.m. pager about a warm pool. At our volume, the difference between serverless and a VM is not a trade-off; it's a rounding error in our favor. This post is the cost model behind that choice — the numbers clients ask for every time we say "pennies per month."

2. The cost model that matters: tokens, not requests

The first mistake in AI cost planning is counting requests. Requests are free; tokens are the bill. And the input side is dominated by the system prompt, which you pay for on every request.

Concrete math from the receptionist pipeline:

  • System prompt: ~1,800 tokens
  • Average incoming message: ~300 tokens
  • ~2,100 input tokens per request, before output

At flash-class list pricing (roughly $0.30 per million input tokens and $2.50 per million output tokens at the time of writing), 10,000 requests a month works out to:

  • Input: 21 million tokens → ~$6.30
  • Output: ~4 million tokens (≈400 tokens per reply) → ~$10
  • ≈ $16/month — before caching, tiering, or gating

Our real volume is a fraction of that. The point is not the absolute number; it's that the arithmetic is dominated by a term most teams never look at: the size of the system prompt times the request count.

3. Four levers, in order of impact

Lever one: model tiering

Classify with a flash-class model; escalate to a pro-class model only for the small fraction of requests that need it. The price spread between tiers is an order of magnitude; for triage workloads the quality spread is a few points. Tiering is the single biggest lever we have found, and it costs nothing but a routing rule.

Lever two: caching

Two caches in the pipeline:

  • Exact-match prompt caching. The system prompt is static, so after the first request the input-token cost of the system prompt mostly disappears. This is provider-side; make sure it is enabled.
  • A KV semantic cache for repeated classifications. Normalize the message, hash it, check the cache before calling the model. We see 30-50% hit rates on recurring lead types — the same "we need help with AI security" message in slightly different wording.

Lever three: output discipline

Output tokens cost roughly 8-10× input tokens on most providers. Long outputs are the expensive tail. Structured output mode plus a hard cap on the reply field turns a 1,500-token ramble into a 400-token reply — and the cap is invisible to quality, because the model knows the schema.

Lever four: edge gating

The cheapest inference call is the one you never make. The regex pre-filter, honeypot field, Turnstile bot-gate, and edge rate limit in the receptionist design exist for security reasons and pay for themselves in cost: every bot request that dies at the edge is tokens we do not buy. It is the one lever that is simultaneously a security control and a cost control.

4. The hidden multipliers

  • Retries. A retry storm on a 429 or a timeout multiplies spend. Idempotency keys and exponential backoff are cost controls, not just reliability controls.
  • Storing raw payloads. Every raw message logged or stored is storage plus egress plus review time. Log the classification, not the payload.
  • Secrets in prompts. A credential embedded in a system prompt is a cost and a risk: it bloats every request and it leaks on the first successful injection.
  • Timeouts. A hung inference call still bills the tokens it consumed. Set timeouts; a slow answer is cheaper than a leaked one.

5. Guardrails

We run three:

  • A hard monthly budget in code. When the spend counter for the month exceeds the threshold, inference refuses (fail-closed) and the owner gets a mail. A surprise bill is a process failure, not an act of God.
  • Per-route accounting. Each pipeline (receptionist, harness) reports its own token spend, so a spike is attributable to a feature, not a mystery.
  • Alert thresholds at 50% and 90% of budget.

6. What we deliberately don't do

No provisioned GPUs, no warm pools, no always-on VMs, no autoscaling drama. Serverless is the point: the infrastructure cost of a boutique AI pipeline should be indistinguishable from zero at rest.

The honest limitation: this math is for boutique volume — thousands of requests a month, bursty, latency-tolerant. At millions of requests, batching, async processing, fine-tuned smaller models, and GPU batch inference change the calculus, and we would revisit. We publish the model anyway because it is the same arithmetic at every scale: token counts, cache hit rates, and how many requests never reach the model.


Grey Ridge Signals Group LLC provides AI security and security architecture advisory. We publish what we spend because "how much does AI cost?" deserves a real answer, not a vendor quote.

← Back to Writing
Book a call →
serverless-cost-engineering-for-ai