Skip to content
AI for engineering / Singapore
Back to Insights

Model Economics / 7 min read

Flagship vs. Flash: What AI Models Cost and How They Compare

Compare token prices, benchmark scores and task times for 12 AI models from six providers. A snapshot of measurements from 7 September 2026.

By Four LabsPublished Updated

Choosing a model is a tradeoff between capability, cost, and the time it takes to finish the job. The most capable model on a benchmark is not automatically the best fit for every step in a business workflow.

The charts below pair a flagship model with a lower-cost counterpart from each of six providers. They compare token prices, benchmark scores, price against capability, and task completion time.

The data is a snapshot from September 7, 2026. Check current prices and benchmark measurements before using it to plan a budget.

12 models / 6 providersSnapshot / 07 Sep 2026AA Index / v4.2

Comparison 01

Price per million tokens.

USD per 1M tokens, standard list or peak rates. Input and output are priced separately. Bar lengths use a logarithmic scale; exact prices appear alongside.

InputOutput
GPT-6 AstraOpenAI / Flagship
Input
$10.00
Output
$50.00
GPT-5.6 LunaOpenAI / Efficient
Input
$0.20
Output
$1.20
Claude Fable 5.1Anthropic / Flagship
Input
$10.00
Output
$50.00
Claude Haiku 4.5Anthropic / Efficient
Input
$1.00
Output
$5.00
Qwen3.8-MaxQwen / Alibaba / Flagship
Input
$2.00
Output
$6.00
Qwen3.8-Flash-NextQwen / Alibaba / Efficient
Input
$0.15
Output
$0.47
DeepSeek V4 Pro 0813DeepSeek / Flagship
Input
$1.32
Output
$3.96
DeepSeek V4 Flash 0731DeepSeek / Efficient
Input
$0.44
Output
$1.32
Kimi K3Kimi / Moonshot / Flagship
Input
$3.00
Output
$15.00
Kimi K2.6Kimi / Moonshot / Efficient
Input
$0.95
Output
$4.00
GLM-5.3Z.AI / Flagship
Input
$1.40
Output
$4.40
GLM-5.3-FlashZ.AI / Efficient
Input
$0.15
Output
$0.50

Uncached input. No batch discounts or temporary promotions. “Efficient” groups the lower-cost counterpart selected for this comparison; providers use different tier names.

Comparison 02

Benchmark scores.

Artificial Analysis Intelligence Index v4.2. Higher is better. These benchmark scores help shortlist models; they do not predict accuracy on your workflow.

Claude Fable 5.1Anthropic / Flagship
56.8
GPT-6 AstraOpenAI / Flagship
54.7
Kimi K3Kimi / Moonshot / Flagship
50.2
GLM-5.3Z.AI / Flagship
48.6
Qwen3.8-MaxQwen / Alibaba / Flagship
46.7
GLM-5.3-FlashZ.AI / Efficient
46.2
Qwen3.8-Flash-NextQwen / Alibaba / Efficient
45.6
GPT-5.6 LunaOpenAI / Efficient
43.4
DeepSeek V4 Pro 0813DeepSeek / Flagship
42.1
DeepSeek V4 Flash 0731DeepSeek / Efficient
40.8
Kimi K2.6Kimi / Moonshot / Efficient
35.8
Claude Haiku 4.5Anthropic / Efficient
17.4

Scores are the September 7 snapshot, not a live leaderboard. Fable’s score uses max effort; the timing below uses high effort because a matched max-effort time was unavailable.

Comparison 03

Price and benchmark score.

Blended price assumes 75% input and 25% output tokens: (3 × input + output) ÷ 4. Further left costs less; higher up scores better.

FlagshipEfficient
Blended token price versus intelligenceTwelve models plotted by blended price and Intelligence Index score. Numbers correspond to the complete values below. The horizontal price scale is logarithmic.$0.25$0.5$1$2$5$10$202030405060GPT-6 Astra: $20.000 blended, 54.7 intelligence01GPT-5.6 Luna: $0.450 blended, 43.4 intelligence02Claude Fable 5.1: $20.000 blended, 56.8 intelligence03Claude Haiku 4.5: $2.000 blended, 17.4 intelligence04Qwen3.8-Max: $3.000 blended, 46.7 intelligence05Qwen3.8-Flash-Next: $0.230 blended, 45.6 intelligence06DeepSeek V4 Pro 0813: $1.980 blended, 42.1 intelligence07DeepSeek V4 Flash 0731: $0.660 blended, 40.8 intelligence08Kimi K3: $6.000 blended, 50.2 intelligence09Kimi K2.6: $1.712 blended, 35.8 intelligence10GLM-5.3: $2.150 blended, 48.6 intelligence11GLM-5.3-Flash: $0.237 blended, 46.2 intelligence12Blended USD / 1M tokens (log scale)
Exact model values below. Points with identical prices share the same horizontal position.
  1. 01
    GPT-6 Astra$20.000 blended / 54.7 index
  2. 02
    GPT-5.6 Luna$0.450 blended / 43.4 index
  3. 03
    Claude Fable 5.1$20.000 blended / 56.8 index
  4. 04
    Claude Haiku 4.5$2.000 blended / 17.4 index
  5. 05
    Qwen3.8-Max$3.000 blended / 46.7 index
  6. 06
    Qwen3.8-Flash-Next$0.230 blended / 45.6 index
  7. 07
    DeepSeek V4 Pro 0813$1.980 blended / 42.1 index
  8. 08
    DeepSeek V4 Flash 0731$0.660 blended / 40.8 index
  9. 09
    Kimi K3$6.000 blended / 50.2 index
  10. 10
    Kimi K2.6$1.712 blended / 35.8 index
  11. 11
    GLM-5.3$2.150 blended / 48.6 index
  12. 12
    GLM-5.3-Flash$0.237 blended / 46.2 index

Comparison 04

Time and cost per task.

Median minutes per Intelligence Index task, output tokens per second, and average API cost per task. Lower task time and cost are better; higher output speed is faster.

GPT-5.6 LunaOpenAI / Efficient
5.6 min
Output
120.6 tokens/s
Cost / task
$0.005

GPT-5.6 Luna (max)

DeepSeek V4 Flash 0731DeepSeek / Efficient
6.0 min
Output
127.8 tokens/s
Cost / task
$0.004

DeepSeek V4 Flash 0731 (max)

Claude Fable 5.1Anthropic / Flagship
7.0 min
Output
56.8 tokens/s
Cost / task
$0.256

Claude Fable 5.1 (high with fallback)

GPT-6 AstraOpenAI / Flagship
7.8 min
Output
59.2 tokens/s
Cost / task
$0.215

GPT-6 Astra (max)

DeepSeek V4 Pro 0813DeepSeek / Flagship
8.7 min
Output
79.9 tokens/s
Cost / task
$0.013

DeepSeek V4 Pro 0813 (max)

GLM-5.3-FlashZ.AI / Efficient
14.7 min
Output
61.8 tokens/s
Cost / task
$0.003

GLM-5.3-Flash

GLM-5.3Z.AI / Flagship
15.4 min
Output
70.4 tokens/s
Cost / task
$0.013

GLM-5.3 (max)

Kimi K3Kimi / Moonshot / Flagship
18.7 min
Output
41.5 tokens/s
Cost / task
$0.066

Kimi K3 (max)

Claude Haiku 4.5Anthropic / Efficient
Not published
Output
80.6 tokens/s
Cost / task
Not published

Claude 4.5 Haiku (Non-reasoning)

Qwen3.8-MaxQwen / Alibaba / Flagship
Not published
Output
Not published
Cost / task
$0.039

Qwen3.8 Max

Qwen3.8-Flash-NextQwen / Alibaba / Efficient
Not published
Output
73.4 tokens/s
Cost / task
Not published

Qwen3.8-Flash-Next

Kimi K2.6Kimi / Moonshot / Efficient
Not published
Output
57.7 tokens/s
Cost / task
Not published

Kimi K2.6

Missing measurements remain “Not published.” Qwen Max’s estimated completion time from the original chart is omitted because the saved metrics do not include its output speed. Fable uses high-effort timing and the corresponding 56.8 tokens/s measurement. Other values retain their recorded configuration.

Comparison 05

Sources and reading notes.

This is a dated comparison, built from the source chart and its saved Artificial Analysis measurements. Source pages can change after the snapshot date.

The token prices and v4.2 scores come from the September 7, 2026 snapshot. Later index versions can use different scores and cost-per-task methods. Check current measurements and rates before budgeting.

Prices exclude cache-hit discounts. DeepSeek is shown at peak rates; GLM-5.3-Flash uses its list rate, excluding the launch promotion. Kimi K3 uses cache-miss input pricing. Regional endpoints, context length, reasoning settings, retries, and tools can change actual spend.

Data labels use the specific saved variants: DeepSeek Pro 0813, DeepSeek Flash 0731, and Qwen3.8-Flash-Next. “Flash” in the article title is shorthand for the selected efficient tier, not a uniform provider product name.

Pricing references: OpenAI, Anthropic, DeepSeek, Moonshot, Z.AI, Alibaba Cloud. Benchmark source: Artificial Analysis.

Start with the workflow, then choose the model

A benchmark can help you choose models to test. For document extraction, check whether the model finds the right fields and handles missing information. For a support assistant, check whether its answers agree with approved company documents.

An efficient model may be enough for routine classification, formatting, or extraction. A more capable model may be useful for ambiguous cases or steps that need deeper reasoning. Measure both on the same representative examples before deciding where each belongs.

Track the cost per completed result your team can accept, including retries, long outputs, tool calls and human corrections.

Talk to The Four Labs

We can help you test AI on a task your team handles regularly. Tell us about the work.