Model Pricing Comparison

AI Model Pricing Comparison: GPT, Claude, Gemini & DeepSeek

A manually reviewed reference comparison of current LLM API pricing — input and output token cost, context scope, intended workload, and official source. Verify provider terms before choosing a production model.

Pricing is stored in a local catalog and updated only after human review. Each entry links to the provider's official pricing page.

Pricing data last checked: Aug 22, 2026 · Data sources: official provider pricing pages

ProviderModelInput / 1M tokOutput / 1M tokContext windowBest use casesLast checkedSource
OpenAI
GPT-5.6 Sol
$5$30Short-context tierComplex professional reasoning and coding
Last checked: Aug 22, 2026
Official
OpenAI
GPT-5.6 Terra
$2.5$15Short-context tierBalanced intelligence and cost
Last checked: Aug 22, 2026
Official
OpenAI
GPT-5.6 Luna
$1$6Short-context tierCost-sensitive, high-volume workloads
Last checked: Aug 22, 2026
Official
Anthropic
Claude Opus 5
$5$25See provider model limitsHigh-capability analysis and agentic work
Last checked: Aug 22, 2026
Official
Anthropic
Claude Sonnet 5
$2$10See provider model limitsGeneral production, coding and agents
Last checked: Aug 22, 2026
Official
Anthropic
Claude Haiku 4.5
$1$5See provider model limitsFast, lower-cost Claude workloads
Last checked: Aug 22, 2026
Official
Google
Gemini 3.7 Flash
$0.75$3.75See provider model limitsAgentic workflows and multimodal reasoning
Last checked: Aug 22, 2026
Official
Google
Gemini 3.5 Flash
$1.5$9See provider model limitsFast grounded and search-oriented workloads
Last checked: Aug 22, 2026
Official
Google
Gemini 3.5 Flash-Lite
$0.3$2.5See provider model limitsHigh-volume translation and simple processing
Last checked: Aug 22, 2026
Official
DeepSeek
DeepSeek V4 Flash
$0.44$1.321MLow-cost general and reasoning workloads
Last checked: Aug 22, 2026
Official
DeepSeek
DeepSeek V4 Pro
$1.32$3.961MHigher-capability DeepSeek workloads
Last checked: Aug 22, 2026
Official

Disclaimer: Prices are manually reviewed reference values in USD per one million tokens and are not real-time. Always verify current pricing on each provider's official page before making a purchasing decision.

Why price isn't the whole story

The cheapest model is not always the best choice

A low per-token price looks attractive on a spreadsheet, but production AI costs are shaped by more than the rate card. Weigh these three tradeoffs before locking in a model for your application.

Accuracy vs. cost

Cheaper models often need more retries, longer prompts, or extra verification steps to reach the same output quality. A model with a higher token price but fewer failed generations can end up cheaper per successful task — and it protects the user experience.

Latency vs. cost

Smaller, cheaper models are usually faster, which matters for real-time chat or agent loops. Larger flagship models can be slower even though they reason better — so the "best" model depends on whether your product needs instant responses or can tolerate a few extra seconds for higher-quality output.

Production requirements

Rate limits, context window size, function-calling support, uptime, and data-handling policies all differ by provider and tier. A model that's cheap in isolated testing can become the expensive choice once you factor in outages, throttling, or the engineering time needed to work around its limitations at scale.

Rule of thumb: use a fast, cheap model for high-volume, low-risk tasks (classification, extraction, simple chat), and reserve a higher-accuracy model for the steps where a mistake is costly — then measure real cost per successful outcome, not just cost per token.

See what this pricing means for your workload

Plug your token counts into the calculator to see real cost per request and per month for any model above.

Open the calculator