How we verify every price
No scraping, no aggregator feeds, no estimates. Here's exactly what goes into a row on this site, and what we deliberately leave out.
What "input" and "output" mean
Input tokens are what you send the model — your prompt, retrieved context, documents, conversation history. Output tokens are what it generates back. Output is almost always priced higher than input, often 3–5x, because generating text is more computationally expensive than reading it.
A "token" is roughly ¾ of an English word, though this varies by tokenizer and language. Providers bill per 1 million tokens; we display every rate that way for consistency, even when a provider's own page uses a different unit (per 1K tokens, for example).
What we standardize on
Most providers publish more than one rate per model — a short-context tier and a long-context tier, a standard rate and a cached-prompt rate, on-demand versus batch or provisioned throughput. To keep the comparison table apples-to-apples, we show the standard, short-context, on-demand rate for every model, and note any other published tier in that model's notes or on its dedicated page.
Some rates are explicitly promotional or time-boxed (a launch discount that steps up on a stated date). Where that's disclosed on the provider's page, we note the step-up date rather than silently treating the promotional rate as permanent.
What we don't do
- We don't track benchmark scores or quality rankings — this is a price comparison, not a capability leaderboard.
- We don't fabricate trend lines. Because this tracker only recently started, there's no historical price data yet; any "change since last check" figure you eventually see here will be real, dated, and sourced the same way as the base price.
- We don't list every pricing tier a provider offers (batch, fine-tuning, caching, provisioned throughput) — only the standard API rate, to keep the table comparable across providers.