AI3Radar
AI3Radar OfficialAI3Radar update

LLM API token costs compared: GPT-5.5, Claude Sonnet 5, Opus 5, and Gemini 3.6 Flash

Models & Real TestsReal testVersion 14 evidence sources

Official per-million-token prices for current flagship and workhorse models - GPT-5.5, GPT-5.4, GPT-5.4 mini, Claude Sonnet 5, Opus 5, and Gemini 3.6 Flash - plus a method for estimating a real monthly API bill instead of trusting benchmark comparisons.

Per-million-token prices look small until you multiply. A price tag of 30 dollars per million output tokens sounds abstract, but one long agentic session can consume tens of thousands of tokens of output, and a monthly integration can run millions. This article lists the official list prices that providers publish today, then shows how to estimate what a real workload actually costs, because the number on the pricing page is only the start of the calculation.

Official list prices per million tokens as of 2026-08-05: OpenAI GPT-5.5 is 5.00 dollars input and 30.00 dollars output with 0.50 dollars cached input; GPT-5.4 is 2.50 and 15.00 dollars with 0.25 dollars cached input; GPT-5.4 mini is 0.75 and 4.50 dollars with 0.075 dollars cached input. Anthropic Claude Sonnet 5 is 2.00 and 10.00 dollars through August 31 2026, then 3.00 and 15.00 dollars; Claude Opus 5 is 5.00 and 25.00 dollars. Google Gemini 3.6 Flash is 1.50 and 7.50 dollars. The table below keeps the full set in one place.

The real bill is three numbers, not two. Input tokens are usually the largest count but the smallest price; output tokens are fewer but far more expensive; cached input can cut input cost dramatically for repeated prefixes. Batch APIs reduce both input and output cost by 50 percent in exchange for delayed completion. A workload that reads a long prompt once and generates a short reply is cheap; a workload that regenerates long outputs in a loop is expensive, regardless of the model's headline price.

A practical estimation method: count a typical day's requests, the average input size in tokens, the average output size, and how much of the input repeats across calls; then run the arithmetic against the list prices above and multiply by your monthly volume. The limitation is that list prices change, introductory rates expire, and enterprise or usage-based discounts are not reflected, so treat the result as an upper-bound estimate, not a quote.

Official API list prices per 1M tokens (2026-08-05)
ProviderModelInputOutputNote
OpenAIGPT-5.5$5.00$30.00Cached input $0.50
OpenAIGPT-5.4$2.50$15.00Cached input $0.25
OpenAIGPT-5.4 mini$0.75$4.50Cached input $0.075
AnthropicClaude Sonnet 5$2.00$10.00Introductory to 2026-08-31, then $3.00/$15.00
AnthropicClaude Opus 5$5.00$25.00Batch API saves 50%
GoogleGemini 3.6 Flash$1.50$7.50Batch API saves 50%

Official API list prices collected on 2026-08-05. Introductory, regional, and enterprise pricing may differ; check provider pricing pages before budgeting.