Technology

Why Usage-Based APIs Should Stop at a Spending Limit by Default

Martin HollowayPublished 21m ago2 min readBased on 1 source
Reading level
Why Usage-Based APIs Should Stop at a Spending Limit by Default
Photo by Paul Downey from Berkhamsted, UK / CC BY 2.0

On 3 October 2026, Simon Willison argued that pay-by-usage services and APIs should ship with hard budget caps switched on by default, stopping service after $X per month. The argument appeared in a blog post titled “We’re going to need default hard budget caps on pretty much everything” Simon Willison.

The post draws a line between enforcement and notification. A hard cap ends or pauses use once the limit is reached. A soft cap sends warning emails while the meter keeps running. Willison’s position is that warning emails alone will not cut it.

Defaults are central to the proposal. Caps would not be an opt-in control. They would be on out of the box, and running without a cap would require explicit action by the customer. Uncapped use would become a deliberate choice rather than an accident of setup.

Looking at what this means for operators, the semantics are fail-closed, meaning cost control wins over keeping the service up. That choice flows into retry logic, or how systems re-try failed work, queue depth, or how much work waits in line, user-facing errors, and on-call response. A service that hits its cap does not wind down gently on its own. It stops.

The broader context here is the spread of metered consumption as a pricing model for infrastructure and platform services. Metered billing ties cost to use. It also turns normal operational wobbles directly into spending swings. Without a circuit breaker, a fast switch that cuts use when spend spikes, a misconfigured client, runaway loop, or unexpected traffic pattern bills at machine speed.

In my view, the default matters more than the mechanism. Quotas, or fixed allowances, rate limits, or caps on request speed, and usage alerts are well understood primitives. The gap is product policy. Opt-in limits avoid interrupting paying workloads. Default-on limits avoid unbounded invoices. Willison is arguing for the second priority as the baseline.

Looking at what this means for API providers, hard caps need support systems that hold up under load. Tenants need per-key and per-project budgets, or separate limits for each key and project, not only account-level totals. Usage accounting must be close to real time, or enforcement will lag behind spend. Platforms must define behavior for in-flight requests, queued jobs, and retained state when the breaker trips. Restarts must be deterministic, happening the same way each time.

Worth flagging as a design tension, availability and cost control pull in opposite directions. Fail-closed protects the budget and breaks the workload. Fail-open protects uptime and breaks the budget. There is no neutral setting. A default hard cap chooses a side, and forces an explicit override if uptime should win. That explicitness is the point.

In my view, the optimistic read is that constraints improve systems. Strict budgets reward caching, or reusing past results, idempotency, or making retries safe, backoff, or waiting longer between tries, sampling, and local evaluation before paid calls. They turn cost into an architectural input rather than a month-end surprise. If defaults change, engineering practice around metered dependencies will adapt quickly.