Frequently asked questions
About the data, the methodology, and the site itself.
Mistral Small 4 (Mistral) is the cheapest model we track, at $0.15 per million input tokens and $0.60 per million output tokens. Several other budget-tier models sit close behind it — see the full table above.
Wide. Claude Fable 5.1 (Anthropic) charges 66.7x more per input token than Mistral Small 4 (tied with GPT-6 Astra at the top end). Flagship pricing generally reflects larger models with deeper reasoning, not a linear improvement in output quality — worth benchmarking against your own task before assuming the expensive model is the right one.
Input tokens are what you send the model — your prompt, context, documents, conversation history. Output tokens are what it generates back. Output is almost always priced higher, often 3–5x the input rate, because generating text is more computationally expensive than reading it.
Meta doesn't sell first-party API access to Llama — it publishes the model weights and lets other companies host it. The prices we show for Llama 4 are from Together AI, one of several inference providers serving it; other hosts (Groq, Fireworks, and others) may price it differently.
Every price on this page was checked directly against the provider's own pricing page on September 10, 2026. We re-verify every 48 hours going forward — this was the first pass, so there's no history yet, but every future check will be dated and any change logged, so the record builds from here.
No. PerTokens is independent and unsponsored. Every provider is listed the same way, using the same source (their own public pricing page) and the same verification date.
Not yet in the main table — we show each model's standard, short-context, on-demand rate to keep the comparison apples-to-apples across providers. Where a provider publishes a notably different batch or caching rate, we mention it in that model's notes or on its dedicated page.
Email contact@pertokens.com with the model or provider and a link to its official pricing page. The tracked list is hand-curated to major providers with a public API pricing page — additions get the same treatment as everything already here: sourced directly from the provider's own page, dated, and never estimated.
This page answers specific questions; the methodology page explains the full sourcing and verification process in one place, including what we standardize on and what we deliberately leave out.