Model Economics / 7 min read
Flagship vs. Flash: What AI Models Cost and How They Compare
Compare token prices, benchmark scores and task times for 12 AI models from six providers. A snapshot of measurements from 7 September 2026.
Choosing a model is a tradeoff between capability, cost, and the time it takes to finish the job. The most capable model on a benchmark is not automatically the best fit for every step in a business workflow.
The charts below pair a flagship model with a lower-cost counterpart from each of six providers. They compare token prices, benchmark scores, price against capability, and task completion time.
The data is a snapshot from September 7, 2026. Check current prices and benchmark measurements before using it to plan a budget.
Comparison 01
Price per million tokens.
USD per 1M tokens, standard list or peak rates. Input and output are priced separately. Bar lengths use a logarithmic scale; exact prices appear alongside.
Uncached input. No batch discounts or temporary promotions. “Efficient” groups the lower-cost counterpart selected for this comparison; providers use different tier names.
Comparison 02
Benchmark scores.
Artificial Analysis Intelligence Index v4.2. Higher is better. These benchmark scores help shortlist models; they do not predict accuracy on your workflow.
Scores are the September 7 snapshot, not a live leaderboard. Fable’s score uses max effort; the timing below uses high effort because a matched max-effort time was unavailable.
Comparison 03
Price and benchmark score.
Blended price assumes 75% input and 25% output tokens: (3 × input + output) ÷ 4. Further left costs less; higher up scores better.
- 01GPT-6 Astra$20.000 blended / 54.7 index
- 02GPT-5.6 Luna$0.450 blended / 43.4 index
- 03Claude Fable 5.1$20.000 blended / 56.8 index
- 04Claude Haiku 4.5$2.000 blended / 17.4 index
- 05Qwen3.8-Max$3.000 blended / 46.7 index
- 06Qwen3.8-Flash-Next$0.230 blended / 45.6 index
- 07DeepSeek V4 Pro 0813$1.980 blended / 42.1 index
- 08DeepSeek V4 Flash 0731$0.660 blended / 40.8 index
- 09Kimi K3$6.000 blended / 50.2 index
- 10Kimi K2.6$1.712 blended / 35.8 index
- 11GLM-5.3$2.150 blended / 48.6 index
- 12GLM-5.3-Flash$0.237 blended / 46.2 index
Comparison 04
Time and cost per task.
Median minutes per Intelligence Index task, output tokens per second, and average API cost per task. Lower task time and cost are better; higher output speed is faster.
- Output
- 120.6 tokens/s
- Cost / task
- $0.005
GPT-5.6 Luna (max)
- Output
- 127.8 tokens/s
- Cost / task
- $0.004
DeepSeek V4 Flash 0731 (max)
- Output
- 56.8 tokens/s
- Cost / task
- $0.256
Claude Fable 5.1 (high with fallback)
- Output
- 59.2 tokens/s
- Cost / task
- $0.215
GPT-6 Astra (max)
- Output
- 79.9 tokens/s
- Cost / task
- $0.013
DeepSeek V4 Pro 0813 (max)
- Output
- 61.8 tokens/s
- Cost / task
- $0.003
GLM-5.3-Flash
- Output
- 70.4 tokens/s
- Cost / task
- $0.013
GLM-5.3 (max)
- Output
- 41.5 tokens/s
- Cost / task
- $0.066
Kimi K3 (max)
- Output
- 80.6 tokens/s
- Cost / task
- Not published
Claude 4.5 Haiku (Non-reasoning)
- Output
- Not published
- Cost / task
- $0.039
Qwen3.8 Max
- Output
- 73.4 tokens/s
- Cost / task
- Not published
Qwen3.8-Flash-Next
- Output
- 57.7 tokens/s
- Cost / task
- Not published
Kimi K2.6
Missing measurements remain “Not published.” Qwen Max’s estimated completion time from the original chart is omitted because the saved metrics do not include its output speed. Fable uses high-effort timing and the corresponding 56.8 tokens/s measurement. Other values retain their recorded configuration.
Comparison 05
Sources and reading notes.
This is a dated comparison, built from the source chart and its saved Artificial Analysis measurements. Source pages can change after the snapshot date.
The token prices and v4.2 scores come from the September 7, 2026 snapshot. Later index versions can use different scores and cost-per-task methods. Check current measurements and rates before budgeting.
Prices exclude cache-hit discounts. DeepSeek is shown at peak rates; GLM-5.3-Flash uses its list rate, excluding the launch promotion. Kimi K3 uses cache-miss input pricing. Regional endpoints, context length, reasoning settings, retries, and tools can change actual spend.
Data labels use the specific saved variants: DeepSeek Pro 0813, DeepSeek Flash 0731, and Qwen3.8-Flash-Next. “Flash” in the article title is shorthand for the selected efficient tier, not a uniform provider product name.
Pricing references: OpenAI, Anthropic, DeepSeek, Moonshot, Z.AI, Alibaba Cloud. Benchmark source: Artificial Analysis.
Start with the workflow, then choose the model
A benchmark can help you choose models to test. For document extraction, check whether the model finds the right fields and handles missing information. For a support assistant, check whether its answers agree with approved company documents.
An efficient model may be enough for routine classification, formatting, or extraction. A more capable model may be useful for ambiguous cases or steps that need deeper reasoning. Measure both on the same representative examples before deciding where each belongs.
Track the cost per completed result your team can accept, including retries, long outputs, tool calls and human corrections.

