Best LLM API for long-context tasks
For feeding entire codebases, long documents, or large transcripts into a single call, context window size matters as much as price. Ranked by published context size, largest first.
Long-context use cases — a full codebase, a legal contract, a stack of research papers — live or die on the published maximum context window, and on how the price scales once you're actually filling it. A bigger window doesn't cost more per token by itself; what changes your bill is sending far more tokens per call. Below are our tracked models with a published context window of 500K tokens or more, sorted largest first.
| Model | Provider | Input /1M | Output /1M | Context |
|---|---|---|---|---|
| OpenAI | $10.00 | $50.00 | ~1.05M tokens | |
| OpenAI | $4.00 | $20.00 | ~1.05M tokens | |
| OpenAI | $2.00 | $12.00 | ~1.05M tokens | |
| OpenAI | $0.20 | $1.20 | ~1.05M tokens | |
| Meta (via Together AI) | $0.27 | $0.85 | ~1.05M tokens | |
| Anthropic | $10.00 | $50.00 | 1M tokens | |
| Anthropic | $5.00 | $25.00 | 1M tokens | |
| Anthropic | $2.00 | $10.00 | 1M tokens | |
| DeepSeek | $0.30 | $1.20 | 1M tokens | |
| xAI | $1.25 | $2.50 | 1M tokens | |
| Alibaba | $2.00 | $6.00 | 1M tokens | |
| Alibaba | $0.40 | $1.60 | 1M tokens | |
| Meta (via Together AI) | $0.18 | $0.59 | 1M tokens | |
| Amazon | $0.30 | $2.50 | 1M tokens | |
| xAI | $2.00 | $6.00 | 500K tokens |
We don't benchmark task-specific quality — this shortlist is built from verified price and published context window only. Use it to narrow candidates by cost, then evaluate output quality yourself.
Frequently asked questions
We set the bar at a published maximum context window of 500,000 tokens or more — roughly 375,000 words. Smaller windows (100K–400K) can still be plenty for many use cases; check each model's page for its exact figure.
Not inherently — the per-token price is set independently of window size. What drives up the bill is usage: filling a 1M-token window costs roughly 10x what filling a 100K window does, at the same rate, simply because you're sending 10x the tokens.
Not automatically. Very long contexts can be slower to process and, on some models, degrade in retrieval accuracy toward the middle of a very long input. If your documents fit comfortably in 100–200K tokens, a smaller-window model may serve you just as well for less.