Figure 1
The same unit of work now costs anywhere from two cents to $2.75
DeepSeek V4 Flash
$0.02MiniO-V2.5-Pro
$0.03DeepSeek V4 Pro
$0.04gpt-oss-120b
$0.06MiniMax-M3
$0.12Grok 4.3
$0.14Nova 2.0 Pro Preview
$0.17Claude 4.5 Haiku
$0.24Nemotron 3 Ultra
$0.24Gemini 3.1 Pro Preview
$0.29Grok 4.5
$0.31Qwen3.5 397B A17B
$0.33Kimi K2.6
$0.35GLM-5.2
$0.37GPT-5.4 mini
$0.48Gemini 3.5 Flash
$0.59Qwen3.7 Max
$1.06Mistral Medium 3.5
$1.08Claude Sonnet 5
$1.53Claude Opus 4.8
$1.80Claude Fable 5
$2.75
$0.02 to $2.75 for the same unit of work — a 137× spread. Several models at the cheap end sit inside the top ten on capability.
Weighted average cost per Intelligence Index task, logarithmic. The models at the cheap end are not toys — several sit inside the top ten on capability. Because the axis is logarithmic, the visible gap understates the real spread.
One row excluded. The published chart lists GPT-5.5 at $0.06, but its plotted bar corresponds to roughly $0.90. Rather than guess which is correct, that model is left out of this figure and its table.
Source: Artificial Analysis — weighted average cost per Intelligence Index task.
Table view — every value in this figure
| Model | Cost per index task |
|---|---|
| DeepSeek V4 Flash | $0.02 |
| MiniO-V2.5-Pro | $0.03 |
| DeepSeek V4 Pro | $0.04 |
| gpt-oss-120b | $0.06 |
| MiniMax-M3 | $0.12 |
| Grok 4.3 | $0.14 |
| Nova 2.0 Pro Preview | $0.17 |
| Claude 4.5 Haiku | $0.24 |
| Nemotron 3 Ultra | $0.24 |
| Gemini 3.1 Pro Preview | $0.29 |
| Grok 4.5 | $0.31 |
| Qwen3.5 397B A17B | $0.33 |
| Kimi K2.6 | $0.35 |
| GLM-5.2 | $0.37 |
| GPT-5.4 mini | $0.48 |
| Gemini 3.5 Flash | $0.59 |
| Qwen3.7 Max | $1.06 |
| Mistral Medium 3.5 | $1.08 |
| Claude Sonnet 5 | $1.53 |
| Claude Opus 4.8 | $1.80 |
| Claude Fable 5 | $2.75 |
The leaderboard and the daily driver have decoupled
Anthropic’s Opus 5 launched at or near the top of the benchmarks at roughly half the price of its own flagship. Within two weeks, experienced users were rolling back to the prior version, and “benchmaxxed” became the recurring accusation — strong on demos and evals, worse on sustained work inside a real codebase.
The figure below is the structural version of the same point. Both columns list the same eleven models in the same order, sorted by intelligence. Read across any row and the two ranks rarely agree: the fastest model on the board scores 24 on intelligence, and the third-most-intelligent is the slowest thing measured.
Figure 2
Capability is a cluster. Speed is a different ranking entirely.
Highlighted: the models whose rank moves furthest. gpt-oss-120b is last on intelligence and first on speed; Kimi K3 is third on intelligence and last on speed.
The same eleven models ranked twice. Nine index points separate first from fifth, and reading across any row the two ranks rarely agree — there is no model that is simply best.
Source: Artificial Analysis — Intelligence Index and output tokens per second.
Table view — every value in this figure
| Model | Intelligence Index | Output tok/sec |
|---|---|---|
| Claude Fable 5 | 60 | 74 |
| GPT-5.6 Sol | 59 | 64 |
| Kimi K3 | 57 | 33 |
| Grok 4.5 | 54 | 67 |
| GLM-5.2 | 51 | 176 |
| Muse Spark 1.1 | 51 | 118 |
| Gemini 3.6 Flash | 50 | 251 |
| MiniMax-M3 | 44 | 91 |
| DeepSeek V4 Pro | 44 | 69 |
| Nemotron 3 Ultra | 38 | 186 |
| gpt-oss-120b | 24 | 267 |
So whatStop treating index position as a procurement signal. Pilot on your own workloads before you standardize, and assume the model you pick today is not the one you will be running in six months.
The frontier moved left, not up
Plot capability against cost and the story of the month is a horizontal one. DeepSeek’s open-weight V4 family put a model at index 50 for two cents a task — a price band that six months ago contained only small models. OpenAI cut its lowest tier by about 80% in the same week.
The efficient frontier now has only three points on it, and the steps between them are brutal: +4 index points costs 15× more per task; +10 costs 137× more. Everything else on the chart is paying for something other than capability.
Is it sustainable? It looks structural rather than subsidized. Mixture-of-experts sparsity, self-optimized kernels and inference-specific silicon are landing at once — and Western hosts serving the same open weights are matching the prices, which is hard to square with a dumping theory.
Figure 3
Up and to the left is the whole story
The frontier has three points on it. Everything else is beaten by something cheaper. The steps are brutal: +4 index points costs 15× more per task, +10 costs 137×.
Eight models for which both capability and cost are published, on a logarithmic cost axis. The frontier line connects the only three that are not beaten by something cheaper.
Rather than trace the source’s 150-point scatter, the two Artificial Analysis charts that could be read exactly were joined, and only the eight models appearing in both are plotted. DeepSeek V4 Flash is placed at the index level published for its 0731 (max) configuration.
Source: Artificial Analysis — Intelligence Index joined to cost per index task.
Table view — every value in this figure
| Model | Cost per task | Intelligence Index |
|---|---|---|
| Claude Fable 5 | $2.75 | 60 |
| Grok 4.5 | $0.31 | 54 |
| GLM-5.2 | $0.37 | 51 |
| DeepSeek V4 Flash | $0.02 | 50 |
| MiniMax-M3 | $0.12 | 44 |
| DeepSeek V4 Pro | $0.04 | 44 |
| Nemotron 3 Ultra | $0.24 | 38 |
| gpt-oss-120b | $0.06 | 24 |
So whatRe-price your roadmap. Workloads you shelved on unit economics — document-scale extraction, per-transaction review, always-on monitoring — may have crossed into viability without you noticing.
More reasoning effort is not better. It is a cost multiplier.
This is the least intuitive finding of the month and the most operationally useful. Frontier software-engineering benchmarks have started sweeping all reasoning-effort levels rather than reporting the maximum — and the curves do not all point the same way.
One frontier model scores 89 at low effort and 85 at high. Another peaks at medium and gives the gain straight back. Only one of the four below is worth its maximum setting. The models are over-thinking: burning tokens on self-review, re-planning and elaborate scaffolding instead of executing the task.
Figure 4
Only one of these four is worth running at maximum effort
GPT-5.6 Sol
+9low → high
Grok 4.5
+8low → high
DeepSeek V4-Flash
0low → high
Claude Fable 5
-4low → high
Shared vertical scale across all four panels. Only GPT-5.6 Sol is worth its maximum setting. Claude Fable 5 scores highest at its cheapest setting and loses four points on the way up.
pass@1 on 23 frontier-hard software-engineering tasks, swept across reasoning-effort levels on a shared vertical scale. The badge on each panel is the signed change from low to high.
Coverage is not uniform. The model showing the four-point drop had 2–4 of the 23 tasks refused by its own safety filters, and the count differs by effort level — so its low and high runs are not measured over an identical task set. A second model is omitted on 5 of 23. The finding holds directionally: DeepSeek V4-Flash shows the same peak-and-fall with full coverage, and Grok 4.5 gains nothing above medium.
Source: VulcanBench Eval Suite 3 — 23 frontier-hard tasks from merged open-source pull requests, 1 August 2026.
Table view — every value in this figure
| Model | Low | Medium | High | Low → High |
|---|---|---|---|---|
| GPT-5.6 Sol | 78 | 83 | 87 | +9 |
| Grok 4.5 | 83 | 91 | 91 | +8 |
| DeepSeek V4-Flash | 87 | 91 | 87 | 0 |
| Claude Fable 5 | 89 | 81 | 85 | -4 |
So whatReasoning effort is a routing parameter, not a quality dial. If your evaluation harness only tests each model at maximum effort, you are almost certainly paying a large premium for worse results on some share of your traffic.
Open weights crossed the enterprise threshold
The month’s open releases were not the usual second tier. Moonshot shipped full weights for a 2.8-trillion-parameter mixture-of-experts model — third or fourth overall on the Intelligence Index and first among open-weight models — along with the surrounding stack: kernels, MoE communication layer, agent infrastructure. DeepSeek V4, GLM-5.2 and Qwen3.8 fill in immediately behind it.
For a regulated enterprise that removes the last real objection. Data residency no longer costs you a capability discount, and weights you hold cannot be deprecated, re-priced or rate-limited out from under a production system. The practical constraint is now hardware footprint, not availability — at roughly 1.5 TB, the largest of these is self-hostable but not casually so.
- Parameters
- 2.8T MoE
- Active per token
- 104B
- Context window
- 1M tokens
- Weights on disk
- ~1.5 TB
So whatStand up an open-weight evaluation track alongside your API vendors this quarter — not to migrate, but to hold a credible alternative. It is the cheapest negotiating position you will ever buy, and the only configuration that survives a vendor deprecating a model you depend on.
Access is not the constraint. Adoption is.
The labour data moved from anecdote to payroll this month: roughly 101,700 AI-attributed US job cuts year to date — about 23% of all 2026 layoffs, and the single most-cited reason for four consecutive months. Stanford finds early-career roles in AI-exposed occupations down around 13% since 2022.
The counter-signal deserves equal weight. In a survey of 350 executives, Gartner found firms that cut headcount for AI reported the same financial gains as firms that did not. Replacement is not where the return is showing up — which makes the figure below the more interesting one.
Figure 5
Eighty-four in every hundred people have never used an AI model
The true proportion
84in every 100 people have never used an AI model
The three paying tiers are the sliver on the right. On a true scale they are almost nothing — which is the finding, so they are not inflated to be visible.
The same four numbers, by depth of use
Never used AI
6.8BFree chatbot user
1.3BPays ~$20/mo
20MUses a coding scaffold
3.5M
Logarithmic. Each step down is a fall of roughly an order of magnitude — 6.8 billion to 3.5 million across the whole range.
The same four tiers on two scales: the true linear proportion, and a logarithmic view by depth of use. From 6.8 billion non-users down to a few million people running agentic coding tools is a fall of more than three orders of magnitude.
Source: population penetration estimate, February 2026. Order-of-magnitude only.
Table view — every value in this figure
| Tier | People | Share of 8.1B |
|---|---|---|
| Never used AI | 6.8B | 83.95% |
| Free chatbot user | 1.3B | 16.05% |
| Pays ~$20/mo | 20M | 0.25% |
| Uses a coding scaffold | 3.5M | 0.04% |
So whatNobody has won this yet. Model access is a commodity your competitors have too; what they mostly do not have is an organization that can absorb it. Treat capability diffusion inside your own company as the scarce input, because it is.