How to read an LLM price list

A price list is a table of rates with a unit hidden in each column header, and most comparison mistakes come from reading the rate without the unit. LLM pricing per million tokens is the common currency of these lists, but the lists also carry cached rates, batch rates, tiers that depend on prompt length, and image and audio units that are not tokens at all. This guide is about reading one vendor's page correctly, recording it with its date, and knowing which number is the baseline.

LLM pricing per million tokens: the unit and the two meters

Every text model has two meters. Input tokens are everything you send: system prompt, tool schemas, history and the user's message. Output tokens are everything the model generates, including reasoning tokens where a vendor bills them as output. The two are priced separately, and output is usually the dearer meter, so a model's price is at least two numbers.

The unit is one million tokens, so a rate reads as currency per million input tokens and per million output tokens. Older pages quote per thousand; convert before you compare, and keep the original beside the converted figure.

A token is not a word. Each model family counts text with its own tokenizer, so the same document is a different number of tokens on two models, and tokenizers change between generations without the price list changing. Two lines with the same rate can bill differently for the same text; the count that matters is the one in the response's usage object.

Cached input has its own column

Most vendors now price a cached prefix separately. A cache read is billed at a fraction of the input rate when the start of a request matches a recent one; a cache write, where the vendor charges one, is billed at a premium on the request that creates the entry. The read rate is usually in the table; the write rate and the lifetime are often in a footnote or a separate page.

Read the conditions, not only the rate: caching is prefix-based and expires after a window that varies by vendor. When prompt caching pays and when it bites covers the mechanics. For a price list the rule is simpler: record the read rate, the write rate and the lifetime as three fields, with a blank where the vendor publishes none; a blank must never be read as a zero.

Batch, tiered and context-band prices

Three further kinds of rate change the arithmetic, and each belongs in its own field.

The baseline for comparing models is the standard, real-time, pay-as-you-go rate for the base tier, with every other rate written beside it as a note that names its condition. Comparing one vendor's batch rate with another's real-time rate compares two services, not two models.

Images, audio and other units

Beyond text, the unit in the column header changes, and the safe habit is to read it every time.

None of these is comparable across vendors without the unit. Record the unit as a field next to the rate, and let a comparison fail when two units differ.

Why the list price is the baseline, and why the bill still differs

The public list price is the baseline because it is the one number every buyer can see, dated and unconditional. Contract rates are private and conditional; an intermediary's price is that intermediary's policy. Compare at list, then negotiate, keeping the list in the comparison so that a discount reads as one.

The bill will still differ from list times tokens, for mechanical reasons: cached reads and writes at their own rates, context injected by tools and agents, tokenizer changes between generations, tiers crossed without noticing, retries that billed the prompt twice, and streams that produced output before they failed. Why your LLM bill does not match the price list walks through the first four and recomputes a single request by hand; the retry budget covers the last two. The list price is not wrong; it is the first term of a longer sum.

Reading the vendor's page and recording the date

Read the vendor's own page, not a screenshot or a third-party table, because the page is the only thing the vendor stands behind; OpenAI's pricing page and Anthropic's pricing page are the kind of page meant. Then record, for every model you care about:

  1. the exact model id, since a dated snapshot and an alias can carry different prices and different retirement dates;
  2. the unit, the currency, and the platform or region, since some vendors publish a different list per region;
  3. the input, output, cached-read and cached-write rates, each in its own field, with a blank where none is published;
  4. any batch, off-peak, volume or context-band condition, as a note that names the condition;
  5. the URL of the page and the date you read it.

The date is not decoration. Vendors change prices without notice and retire models on a schedule, so a rate without a date cannot be reconciled against a given month's invoice. Re-read the page on a schedule and whenever a model is announced or retired, and keep the old rows: a ledger that stores the rate applied at call time is only auditable if that rate still exists somewhere.

Runix applies the same rule to its own catalogue. Every model on the Runix Router catalogue is listed at its vendor's public list price, with the source page and the date it was read; contract rates are quoted in writing rather than shown on the page. It is a dated snapshot of public lists, not a rate card of its own, and Runix Router, in early access, meters usage at those list prices on itemised statements in USD.

Questions this raises

Is a price per million tokens comparable across vendors?

Only after you check the unit, the tier and the tokenizer. Two models can list the same rate and bill differently for the same text because they count tokens differently.

Which price should I use when comparing models?

The standard real-time pay-as-you-go rate for the base tier. Record batch, cached, off-peak and volume rates as separate notes that name their conditions.

How often should a recorded price list be rechecked?

On a fixed schedule and whenever a vendor announces a new model or a retirement. Keep the URL and the date read with every number, and keep old rows rather than overwriting them.

Related to this post: The Router models catalogue, at vendors' public list prices. Tell us what you are building and we reply within one business day.