Turning an LLM bill into a policy: budgets per team and per feature

An invoice tells you what happened. A budget is a rule about what is allowed to happen, with an owner and a defined behaviour at the limit, enforced where a client cannot bypass it. Most teams have the first and believe they have the second. An LLM budget per team is the usual starting point, because it maps to how money is approved; a budget per feature is what engineers need, because it maps to what they can change. This post turns one bill into both.

A bill is a report, a budget is a policy

A budget has four parts, and a dashboard number has one of them: a ceiling in currency or tokens, a period after which it resets, an owner who receives the alerts and can raise the ceiling, and a behaviour at the limit, decided before launch rather than during the incident.

The unit of enforcement is the key, the only thing the gateway can see on every request before it leaves. A header or tag a client attaches voluntarily is useful for reporting and useless for enforcement, since the one client that forgets it is the one running the loop. One key can drain a shared quota pool covers a ceiling that exists only in aggregate.

Per-key quotas, rate limits and allowed models

Three controls on a key do three jobs, and a budget uses all of them.

Issue one key per team, feature and environment, named so that all three are readable from the id. The evaluation harness, the staging deployment and the internal assistant each get their own; they are the consumers nobody remembers until the production key is empty. A key should be revocable without a redeploy: the kill switch for a runaway loop is revocation, not a hotfix.

LLM budget per team versus per feature

A team budget matches how money is approved: a team owns a ceiling, decides what to spend it on, and answers for it. A feature budget matches how the product is built: a feature has a unit cost, a quality bar, and a model choice an engineer can change this afternoon.

They do not nest: a feature is served by several teams and a team runs many features, so a budget enforced on one dimension can only be reported on the other. Enforce on keys issued per team and feature, so that the key id carries both, and report along whichever dimension the reader needs. Where that produces too many keys, enforce per team and require a feature tag, knowing the tag is reporting, not enforcement.

Set a feature ceiling from its unit cost and a team ceiling from the sum of its features plus headroom; when the two disagree, the disagreement is the finding, and a bill alone could not have shown it.

Hard stops, alerts, and what happens at the limit

An alert tells an owner; a hard stop acts on its own. A budget needs both: an alert during a quiet weekend reaches nobody, and a hard stop with no earlier alert is an outage of your own making. Alert on runway, remaining ceiling divided by current spend, when it is shorter than the days left in the period; a fixed fraction of the ceiling is a weaker second alert, since the same balance means different things at different rates. Stop at the ceiling.

What stopping means is decided per class of traffic, and there are three honest answers:

Write the chosen behaviour into the policy, next to the ceiling it belongs to:

key: TEAM/FEATURE/ENV
period: monthly
ceiling: AMOUNT_USD
alert_when: runway_days below DAYS_LEFT_IN_PERIOD
allowed_models: [MODEL_ID, MODEL_ID]
at_limit: fail_closed     # or degrade | queue
counts: attempts          # retries and failover included
owner: TEAM_ALIAS

Retries and failover draw from the same budget

The line most budgets get wrong is counts: attempts. A retried request re-sends the prompt at full price; a request that failed over was billed by the first provider if it had already processed the prompt; a stream that died after producing output was billed for that output. To the provider each is a separate request; to the budget they are the same intent, spent more than once.

A budget that counts only successes under-counts in exactly the hour it exists for, when a provider is degrading and every client is retrying into it. Count attempts, and make the retry budget and the spend budget read from the same meter, so that retries cannot spend what the ceiling has refused.

This is why the budget belongs in the gateway: client-side checks cannot see SDK retries, the second provider, or the other clients sharing a pool, and the hot path can. Treat it like load shedding in any other service: the SRE book's chapter on handling overload is about CPU and queries, and the reasoning transfers to tokens.

Reporting that closes the loop

A budget is only maintained if the report reaches the owner named in it, on a schedule, in a form that leads to a decision. Per key and per period: spend against ceiling, runway, attempts against successes, the model mix, and the rate applied at call time, so that a line reconciles against an itemised statement. Cost attribution that survives audit covers recording those fields; the report is the same records rolled up.

Then review ceilings every period. A ceiling nobody has revised in a year is last year's traffic shape written as policy, and a budget that is always exactly spent was set by habit rather than cost. The review is where team and feature meet: a team raises a ceiling, and a feature justifies it with a unit cost.

Runix Router, in early access, enforces a quota, rate limits, allowed models and routing per key, revocable without a redeploy, and bills usage in USD on itemised statements. It is designed to be the single place a per-key ceiling is enforced and reported from; what a key's limit should mean for its traffic is a decision for the team that owns it, made before the key is issued.

Questions this raises

Should an LLM budget be per team or per feature?

Enforce on the dimension the gateway can see on every request, which is the key. Issue keys per team and feature, or per team with a mandatory feature tag for reporting, and report on both dimensions.

What should happen when a key reaches its limit?

Decide per class of traffic before launch: fail closed with a distinguishable error for batch work, degrade visibly or fail closed with a clear message for interactive traffic, and never queue interactive requests without a bound.

Do retries count against the budget?

Yes. Every attempt is billed by the provider, so the budget must count attempts, including failover to another provider and streams that died after producing output.

Related to this post: Runix Router. Tell us what you are building and we reply within one business day.