AI3Radar
AI3Radar OfficialAI3Radar update

How to pick an LLM for your workload: capability tiers, token costs, and a three-step method

Choosing a model by benchmark ranking fails in practice, because your workload is defined by three numbers the rankings never show: how much input you read, how much output you generate, and how often the input repeats. This guide maps the current flagship and workhorse models to those three numbers using official list prices, then gives a selection method that survives model updates.

On official per-million-token list prices as of 2026-08-05: OpenAI GPT-5.5 is $5.00 input and $30.00 output with $0.50 cached input; GPT-5.4 mini is $0.75 and $4.50 with $0.075 cached input. Anthropic Claude Sonnet 5 is $2.00 and $10.00 through August 31 2026 then $3.00 and $15.00; Claude Opus 5 is $5.00 and $25.00. Google Gemini 3.6 Flash is $1.50 and $7.50. Cached input is the cheapest lever: a repeated system prompt that never changes is billed at the cached rate.

The three-step method: first, separate your tasks into quick generation (chat, drafts, extraction) and deep reasoning (coding, analysis, multi-step work); second, run the quick tier on the workhorse model and keep the flagship for the reasoning tier only; third, measure a typical day in input tokens, output tokens, and repeated prefix size, then compute both tiers against the list prices before scaling.

The limitation is that list prices and capability rankings shift quickly, and enterprise or usage-based discounts are not reflected. Re-measure monthly and treat the result as an upper-bound estimate for budgeting, not a permanent answer.

Flagship and workhorse model API prices per 1M tokens (2026-08-05)
ProviderModelInputOutputCached input
OpenAIGPT-5.5$5.00$30.00$0.50
OpenAIGPT-5.4 mini$0.75$4.50$0.075
AnthropicClaude Sonnet 5$2.00$10.00n/a
AnthropicClaude Opus 5$5.00$25.00n/a
GoogleGemini 3.6 Flash$1.50$7.50n/a

Official list prices collected on 2026-08-05. Claude Sonnet 5 intro pricing ($2/$10) runs through 2026-08-31, then $3/$15. Cached-input rates apply to repeated prefixes where the provider supports them.