USE-CASE GUIDES

Best LLM API for long-context tasks

For feeding entire codebases, long documents, or large transcripts into a single call, context window size matters as much as price. Ranked by published context size, largest first.

Long-context use cases — a full codebase, a legal contract, a stack of research papers — live or die on the published maximum context window, and on how the price scales once you're actually filling it. A bigger window doesn't cost more per token by itself; what changes your bill is sending far more tokens per call. Below are our tracked models with a published context window of 500K tokens or more, sorted largest first.

ℹ️How this list is built: Filtered to models with a clearly published numeric context window ≥500K tokens, sorted by context size descending.
15 models
Model Provider Input /1M Output /1M Context
OpenAI $10.00 $50.00 ~1.05M tokens
OpenAI $4.00 $20.00 ~1.05M tokens
OpenAI $2.00 $12.00 ~1.05M tokens
OpenAI $0.20 $1.20 ~1.05M tokens
Meta (via Together AI) $0.27 $0.85 ~1.05M tokens
Anthropic $10.00 $50.00 1M tokens
Anthropic $5.00 $25.00 1M tokens
Anthropic $2.00 $10.00 1M tokens
DeepSeek $0.30 $1.20 1M tokens
xAI $1.25 $2.50 1M tokens
Alibaba $2.00 $6.00 1M tokens
Alibaba $0.40 $1.60 1M tokens
Meta (via Together AI) $0.18 $0.59 1M tokens
Amazon $0.30 $2.50 1M tokens
xAI $2.00 $6.00 500K tokens

We don't benchmark task-specific quality — this shortlist is built from verified price and published context window only. Use it to narrow candidates by cost, then evaluate output quality yourself.

See all use cases →

Frequently asked questions

We set the bar at a published maximum context window of 500,000 tokens or more — roughly 375,000 words. Smaller windows (100K–400K) can still be plenty for many use cases; check each model's page for its exact figure.

Not inherently — the per-token price is set independently of window size. What drives up the bill is usage: filling a 1M-token window costs roughly 10x what filling a 100K window does, at the same rate, simply because you're sending 10x the tokens.

Not automatically. Very long contexts can be slower to process and, on some models, degrade in retrieval accuracy toward the middle of a very long input. If your documents fit comfortably in 100–200K tokens, a smaller-window model may serve you just as well for less.