How we work out these prices

Every figure on this site comes from a provider's own published pricing, entered by hand and checked by a person. This page explains how that happens, what we do to the numbers before showing them, and where each one stops being reliable.

Models listed
170
Providers
10
Current prices
393
Comparisons published
9

Where the prices come from

Every price is taken from the provider's own pricing page and entered by hand. We run a crawler over those pages to spot when something changes, but it does not write to this site: it raises a change for someone to look at, and a person decides what the new figure is and enters it. That is slower than scraping straight into a database, and it is the reason a price here is a price somebody has read rather than one a parser guessed at.

We list what a provider publishes. Negotiated rates, committed-spend discounts and enterprise agreements are not here, because we have no way to see them — if you have a contract, your real cost is lower than anything this site shows.

What we do to a price before showing it

Providers quote in whatever unit suits them: per million tokens, per thousand, per image, per minute of audio. We store the price exactly as it was published, in the provider's own unit, and convert to a per-million-token figure only when we display it. Storing the original means a change in how we present prices can never quietly alter what a provider actually charges.

A price of zero is a real price. A model that costs nothing shows as "Free" and counts as the cheapest option. That is different from a price we do not have, which is left blank. The two are never mixed up, on any page or in any total.

Long prompts and context bands

Some models charge more once a prompt passes a certain length. The higher rate applies to the whole request, not only to the tokens past the line — a 300,000-token prompt to a model that reprices at 272,000 is charged at the higher rate from the first token. We price each band separately and show them as separate rows, and where the cheaper model changes at a band, the comparison says so rather than quoting one rate and leaving you to find out.

Service tiers

Alongside the standard pay-as-you-go rate, many providers sell the same model on other terms: Batch for work that can wait, Flex for lower priority, Priority for guaranteed capacity. We list every tier a provider offers, including tiers only one of two compared models has.

When we say a tier saves 50%, that is against that model's own standard rate, not against the other model in the comparison. Each column answers "what does this tier save me on this model", which is the question a tier can actually answer.

The four worked costs

Rates per million tokens are hard to feel. Every comparison prices the same four jobs on both models, chosen because each one exposes a different way two models can differ:

  • Chat turn — 1,000 tokens in, 500 out. The base rate at human scale.
  • RAG answer — 10,000 in, 800 out. Input-heavy, the shape most products have.
  • Agent loop — twenty calls, each 30,000 cached tokens plus 2,000 fresh in and 700 out. This is where cached-input pricing decides the winner.
  • Long document — 400,000 in, 3,000 out. Long enough to cross a context band.

Each is shown as a single run and as a thousand runs. One chat turn costs a fraction of a cent on almost any model, which tells you nothing; a thousand of them is a number you can compare. The thousand-run column is the same arithmetic, multiplied — it assumes nothing about your volume.

A workload is left out rather than estimated when a model cannot take it: if a model's context window will not hold 400,000 tokens, we do not quote a price for a request the provider would refuse. Where a model has no cached-input price, its cached tokens are charged at the full input rate and the page says so.

The blended rate

Two models with four rates each cannot be ranked by eye, so we also quote one blended number: input and output mixed three to one. That ratio is a simplification and we say so wherever the number appears. If your own traffic is output-heavy, the blended rate will flatter the wrong model — the four worked costs above are the better guide, and the calculator takes your own token counts.

Benchmark scores

We do not run benchmarks. The scores on this site are published by Artificial Analysis and by Design Arena, and we show them with the name of the evaluation beside the number so you can go and check.

Artificial Analysis scores a model once per reasoning effort — a model run at maximum effort scores differently from the same model run cheaply. We take each model's best result and name the effort level next to it, so a comparison puts both models at their strongest rather than mixing a careful run against a hurried one.

Benchmark coverage is uneven. Where both models in a comparison report the same figure we show it; where they do not, the comparison says so instead of filling the gap. We do not publish speed or latency figures at all, because we have no way to measure them ourselves and would be repeating someone else's number without being able to stand behind it.

How current any of this is

Pages carry two dates, and they answer different questions. Prices as of is when the rates shown took effect at the provider. Last updated is when anything the page shows last changed here — a price, a specification, a new benchmark score.

A provider can change a price at any time, and there will be a gap between them publishing it and us entering it. If a figure here disagrees with a provider's own page, theirs is right and ours is stale — please tell us and we will fix it.