Hard Budget Caps for AI Agents Stop Different Things on Every Platform
Every major cloud and model API now sells a hard spending limit. Before an agent deploys anything, find out what yours actually shuts off.
A budget alert is a polite email that arrives after the money is gone. That has been true of cloud billing for years, and nobody minded much while a human had to click through a console to spin up anything expensive.
Coding agents broke that arrangement. An agent can stand up a database, a serverless function, and a model endpoint in the time it takes you to refill your coffee, and every one of those things bills by usage. On October 3, Simon Willison put the consequence plainly in a post titled "We're going to need default hard budget caps on pretty much everything". Nobody, he wrote, wants "to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage."
His fix is simple. Ship a cap that cuts the service off and returns errors, make it the default, and make removing it an explicit checkbox where you accept responsibility for the overage.
The good news is that the hard cap has finally arrived. AWS, Google Cloud, OpenAI and Anthropic all sell one now. The less comfortable news is that the four caps stop four different things, and if you pick one without reading the fine print, you trade a surprise bill for a surprise outage, or in one case a surprise deletion.
Why "hard" matters more once agents deploy
A soft limit depends on a human reading a notification and acting on it. That loop worked when deployment was slow and deliberate. It fails when the thing spending money is a process that runs at 3 a.m. and does exactly what you told it to do, which was keep going until the tests pass.
The failure is rarely a single giant charge. It is a retry loop that hammers a model API, a function that scales to meet traffic from a crawler, or a forgotten preview environment that an agent created and never tore down. Each looks reasonable in isolation. Together they add up while you sleep.
So the question for a builder in October 2026 is not whether to cap spend. It is what happens to your running system at the moment the cap trips. That answer differs on every platform.
What each cap actually stops
AWS: the whole project pauses, and the clock starts
AWS shipped spend limits in its new builder experience on September 16. The announcement says you can "set a monthly spend limit for your project," and that "if a project's usage reaches its spend limit, your project is paused for that month." The same launch connects your coding agent to AWS "with a single prompt" that installs the AWS CLI and the Agent Toolkit for AWS, so the agent and the cap arrive together.
The AWS documentation spells out the mechanism: "AWS pauses your project and stops all resources. Your data is preserved." Then come three sentences worth reading twice. The experience is reaching "a limited number of customers," so you may not have it yet. "After reactivation, some resources may need to be manually restarted." And: "If you take no action within 90 days of your project being paused, AWS permanently deletes your project data."
AWS also describes optional early controls that step in before the limit: new resources stop launching, then idle resources pause, then the highest-cost active resources pause. And it scopes the whole feature honestly. Spend limits "are designed for experimentation, learning, and sandbox workloads," usable in production only "when it is acceptable to have a brief pause of your resources."
That is the widest blast radius of the four. A runaway Lambda takes your database down with it.
Google Cloud: one service stops, commitments keep charging
Google's Spend Caps, in public preview since July 28, work the other way. When spend reaches the cap, Google "automatically restricts further cost-incurring usage for that specific service within that project." It is "non-destructive," with data and resources kept and "services outside the scope of the budget are entirely unaffected."
The covered services matter for agent builders: the Gemini API, Agent Platform, Cloud Run and Cloud Run Functions. Google says caps on AI services "trigger within minutes" of the threshold. The preview limits you to a single project and service for a fixed monthly window, and one exception carries real money: "any underlying fixed commitment fees (such as Committed Use Discounts or Provisioned Throughput) will continue to bill at their flat contractual rate."
Narrow blast radius, but narrow coverage too. A cap on the Gemini API does nothing about the Compute Engine VM your agent also launched.
OpenAI API: a 429, slightly late
OpenAI's spend limits guide separates the two kinds cleanly. A spend alert "sends a notification; API traffic continues." A hard spend limit makes affected requests "return a 429 error," with the code organization_spend_limit_exceeded or project_spend_limit_exceeded depending on where you set it.
The guide is also candid about timing: "Enforcement is not instantaneous," and "recorded spend can slightly exceed the configured amount." The limit resets with the next monthly cycle unless you raise or remove it.
Claude API: two different errors, and one workspace you cannot cap
Anthropic's rate limits page describes two layers. Every organization has a tier cap (the docs list $500 a month for Start, $1,000 for Build, $200,000 for Scale, and none for Custom). Hit it and usage pauses until 00:00 UTC on the first of the next month, with an HTTP 429 carrying error_code: "enforced_spend_limit_reached" and no retry-after header, so automatic SDK retries keep failing until access returns.
Below that, you can set your own lower limit in the Console. When that custom limit trips, requests return HTTP 400 with invalid_request_error. You can also set limits per workspace, with one catch for anyone who never bothered to organize their keys: "You cannot set limits on the default workspace."
The two traps hiding in those details
The first trap is error handling. Look at the status codes again: 429 from OpenAI, 429 or 400 from Anthropic depending on which limit you hit. An agent harness that treats every 429 as "back off and retry" will sit in a polite loop against a cap that does not reset until next month. One that treats a 400 as a bug in its own request may start rewriting perfectly good code. Your agent needs to recognize a spend cap as its own condition and stop.
The second trap is scope. The cap protects the thing it is attached to and nothing else. Agents rarely stay inside one service. A typical prototype touches a model API, a compute platform, a database and storage, often on two vendors. A Gemini cap and an OpenAI cap together still leave the Postgres instance uncapped.
Put this into practice
You can do most of this in an afternoon.
- List every account an agent can spend from. Model API keys, cloud projects, and any SaaS the agent signs up for. If an agent holds the credential, it belongs on the list.
- Move agent keys out of default buckets. On the Claude API, create a dedicated workspace for agent traffic so a workspace limit is even possible. On OpenAI, give agent work its own project with a project-level hard limit.
- Set the hard limit, not just the alert. Pick a number you would shrug at losing. Keep the alert too, at half that number, so you hear about trouble before the wall.
- Teach the harness the error codes. Match on
project_spend_limit_exceeded,organization_spend_limit_exceededandenforced_spend_limit_reached, and have the agent stop and report instead of retrying. - Trip each cap on purpose once. Set a tiny limit in a sandbox, run a loop, and watch what breaks. You want to learn what a paused project does to your public endpoint on a quiet Tuesday, not during a launch.
- Put a calendar reminder on any AWS pause. Ninety days is long enough to forget a side project and short enough to lose it.
Honest limitations
None of these caps is the default Willison asked for. Every one is opt-in, which means the accounts most likely to get burned, the forgotten side projects, are the least likely to have one set.
Coverage is patchy. AWS's version is still rolling out to a limited set of customers. Google's is a preview that covers four services, one at a time. The model APIs cap model spend and nothing downstream of it.
Enforcement lags in at least one documented case. OpenAI says outright that recorded spend can slightly exceed the limit, and AWS says nothing in its docs about how fresh the billing data behind a pause is. A hard cap bounds your loss. It does not make it zero.
And the hardest cap moves the failure instead of removing it. A paused project is an outage. A capped API is a broken feature. For anything with real users, a cap is a backstop for a monitoring setup you still need, not a replacement for it.
The control belongs to you, for now
Willison's argument is that vendors should ship caps on by default and make you opt out. That would be the right design, and the industry is partway there. Until it gets the rest of the way, the default is whatever you set this week.
So pick your caps by blast radius. Use the narrow ones where you can, accept the wide ones where you must, and test every one before you trust it. An agent will find the edge of your budget eventually. The only choice you get is whether that edge is a wall you built or an email you read too late.
Sources: Simon Willison, AWS What's New, AWS spend limit docs, Google Cloud Spend Caps, OpenAI spend limits, Claude API rate limits.
Medium metadata
- Title: Hard Budget Caps for AI Agents Stop Different Things on Every Platform
- Subtitle: Every major cloud and model API now sells a hard spending limit. Before an agent deploys anything, find out what yours actually shuts off.
- Tags: AI Agents, Cloud Computing, AWS, Software Development, Generative AI
- Reading time: about 8 minutes
- Canonical: import from the fervorai.dev URL