Thought leadershipTechnical BriefAugust 2026

Frontier capability stopped being the bottleneck.

Over the last thirty days the price of frontier-class intelligence fell by roughly two orders of magnitude, while the five leading models converged to within nine index points of each other. That changes the question from “which model is best” to “how fast can we stop being locked to one.”

9 min read

The month in four figures

  1. 137×

    Spread in cost per task between the cheapest and dearest model on the board

    Artificial Analysis

  2. 9 pts

    Separating the top five models on the Intelligence Index — 60 down to 51

    Artificial Analysis

  3. −4 pts

    Change in pass@1 when one frontier model is moved from low to high reasoning effort

    VulcanBench Suite 3

  4. 84%

    Of the world’s population has never used an AI model at all

    Feb 2026 penetration estimate

Figure 1

The same unit of work now costs anywhere from two cents to $2.75

  1. DeepSeek V4 Flash

    $0.02
  2. MiniO-V2.5-Pro

    $0.03
  3. DeepSeek V4 Pro

    $0.04
  4. gpt-oss-120b

    $0.06
  5. MiniMax-M3

    $0.12
  6. Grok 4.3

    $0.14
  7. Nova 2.0 Pro Preview

    $0.17
  8. Claude 4.5 Haiku

    $0.24
  9. Nemotron 3 Ultra

    $0.24
  10. Gemini 3.1 Pro Preview

    $0.29
  11. Grok 4.5

    $0.31
  12. Qwen3.5 397B A17B

    $0.33
  13. Kimi K2.6

    $0.35
  14. GLM-5.2

    $0.37
  15. GPT-5.4 mini

    $0.48
  16. Gemini 3.5 Flash

    $0.59
  17. Qwen3.7 Max

    $1.06
  18. Mistral Medium 3.5

    $1.08
  19. Claude Sonnet 5

    $1.53
  20. Claude Opus 4.8

    $1.80
  21. Claude Fable 5

    $2.75

$0.02 to $2.75 for the same unit of work — a 137× spread. Several models at the cheap end sit inside the top ten on capability.

Weighted average cost per Intelligence Index task, logarithmic. The models at the cheap end are not toys — several sit inside the top ten on capability. Because the axis is logarithmic, the visible gap understates the real spread.

One row excluded. The published chart lists GPT-5.5 at $0.06, but its plotted bar corresponds to roughly $0.90. Rather than guess which is correct, that model is left out of this figure and its table.

Source: Artificial Analysis — weighted average cost per Intelligence Index task.

Table view — every value in this figure
ModelCost per index task
DeepSeek V4 Flash$0.02
MiniO-V2.5-Pro$0.03
DeepSeek V4 Pro$0.04
gpt-oss-120b$0.06
MiniMax-M3$0.12
Grok 4.3$0.14
Nova 2.0 Pro Preview$0.17
Claude 4.5 Haiku$0.24
Nemotron 3 Ultra$0.24
Gemini 3.1 Pro Preview$0.29
Grok 4.5$0.31
Qwen3.5 397B A17B$0.33
Kimi K2.6$0.35
GLM-5.2$0.37
GPT-5.4 mini$0.48
Gemini 3.5 Flash$0.59
Qwen3.7 Max$1.06
Mistral Medium 3.5$1.08
Claude Sonnet 5$1.53
Claude Opus 4.8$1.80
Claude Fable 5$2.75

The leaderboard and the daily driver have decoupled

Anthropic’s Opus 5 launched at or near the top of the benchmarks at roughly half the price of its own flagship. Within two weeks, experienced users were rolling back to the prior version, and “benchmaxxed” became the recurring accusation — strong on demos and evals, worse on sustained work inside a real codebase.

The figure below is the structural version of the same point. Both columns list the same eleven models in the same order, sorted by intelligence. Read across any row and the two ranks rarely agree: the fastest model on the board scores 24 on intelligence, and the third-most-intelligent is the slowest thing measured.

Figure 2

Capability is a cluster. Speed is a different ranking entirely.

Highlighted: the models whose rank moves furthest. gpt-oss-120b is last on intelligence and first on speed; Kimi K3 is third on intelligence and last on speed.

The same eleven models ranked twice. Nine index points separate first from fifth, and reading across any row the two ranks rarely agree — there is no model that is simply best.

Source: Artificial Analysis — Intelligence Index and output tokens per second.

Table view — every value in this figure
ModelIntelligence IndexOutput tok/sec
Claude Fable 56074
GPT-5.6 Sol5964
Kimi K35733
Grok 4.55467
GLM-5.251176
Muse Spark 1.151118
Gemini 3.6 Flash50251
MiniMax-M34491
DeepSeek V4 Pro4469
Nemotron 3 Ultra38186
gpt-oss-120b24267

So whatStop treating index position as a procurement signal. Pilot on your own workloads before you standardize, and assume the model you pick today is not the one you will be running in six months.

The frontier moved left, not up

Plot capability against cost and the story of the month is a horizontal one. DeepSeek’s open-weight V4 family put a model at index 50 for two cents a task — a price band that six months ago contained only small models. OpenAI cut its lowest tier by about 80% in the same week.

The efficient frontier now has only three points on it, and the steps between them are brutal: +4 index points costs 15× more per task; +10 costs 137× more. Everything else on the chart is paying for something other than capability.

Is it sustainable? It looks structural rather than subsidized. Mixture-of-experts sparsity, self-optimized kernels and inference-specific silicon are landing at once — and Western hosts serving the same open weights are matching the prices, which is hard to square with a dumping theory.

Figure 3

Up and to the left is the whole story

The frontier has three points on it. Everything else is beaten by something cheaper. The steps are brutal: +4 index points costs 15× more per task, +10 costs 137×.

Eight models for which both capability and cost are published, on a logarithmic cost axis. The frontier line connects the only three that are not beaten by something cheaper.

Rather than trace the source’s 150-point scatter, the two Artificial Analysis charts that could be read exactly were joined, and only the eight models appearing in both are plotted. DeepSeek V4 Flash is placed at the index level published for its 0731 (max) configuration.

Source: Artificial Analysis — Intelligence Index joined to cost per index task.

Table view — every value in this figure
ModelCost per taskIntelligence Index
Claude Fable 5$2.7560
Grok 4.5$0.3154
GLM-5.2$0.3751
DeepSeek V4 Flash$0.0250
MiniMax-M3$0.1244
DeepSeek V4 Pro$0.0444
Nemotron 3 Ultra$0.2438
gpt-oss-120b$0.0624

So whatRe-price your roadmap. Workloads you shelved on unit economics — document-scale extraction, per-transaction review, always-on monitoring — may have crossed into viability without you noticing.

More reasoning effort is not better. It is a cost multiplier.

This is the least intuitive finding of the month and the most operationally useful. Frontier software-engineering benchmarks have started sweeping all reasoning-effort levels rather than reporting the maximum — and the curves do not all point the same way.

One frontier model scores 89 at low effort and 85 at high. Another peaks at medium and gives the gain straight back. Only one of the four below is worth its maximum setting. The models are over-thinking: burning tokens on self-review, re-planning and elaborate scaffolding instead of executing the task.

Figure 4

Only one of these four is worth running at maximum effort

  1. GPT-5.6 Sol

    +9low → high

  2. Grok 4.5

    +8low → high

  3. DeepSeek V4-Flash

    0low → high

  4. Claude Fable 5

    -4low → high

Shared vertical scale across all four panels. Only GPT-5.6 Sol is worth its maximum setting. Claude Fable 5 scores highest at its cheapest setting and loses four points on the way up.

pass@1 on 23 frontier-hard software-engineering tasks, swept across reasoning-effort levels on a shared vertical scale. The badge on each panel is the signed change from low to high.

Coverage is not uniform. The model showing the four-point drop had 2–4 of the 23 tasks refused by its own safety filters, and the count differs by effort level — so its low and high runs are not measured over an identical task set. A second model is omitted on 5 of 23. The finding holds directionally: DeepSeek V4-Flash shows the same peak-and-fall with full coverage, and Grok 4.5 gains nothing above medium.

Source: VulcanBench Eval Suite 3 — 23 frontier-hard tasks from merged open-source pull requests, 1 August 2026.

Table view — every value in this figure
ModelLowMediumHighLow → High
GPT-5.6 Sol788387+9
Grok 4.5839191+8
DeepSeek V4-Flash8791870
Claude Fable 5898185-4

So whatReasoning effort is a routing parameter, not a quality dial. If your evaluation harness only tests each model at maximum effort, you are almost certainly paying a large premium for worse results on some share of your traffic.

Open weights crossed the enterprise threshold

The month’s open releases were not the usual second tier. Moonshot shipped full weights for a 2.8-trillion-parameter mixture-of-experts model — third or fourth overall on the Intelligence Index and first among open-weight models — along with the surrounding stack: kernels, MoE communication layer, agent infrastructure. DeepSeek V4, GLM-5.2 and Qwen3.8 fill in immediately behind it.

For a regulated enterprise that removes the last real objection. Data residency no longer costs you a capability discount, and weights you hold cannot be deprecated, re-priced or rate-limited out from under a production system. The practical constraint is now hardware footprint, not availability — at roughly 1.5 TB, the largest of these is self-hostable but not casually so.

Parameters
2.8T MoE
Active per token
104B
Context window
1M tokens
Weights on disk
~1.5 TB

So whatStand up an open-weight evaluation track alongside your API vendors this quarter — not to migrate, but to hold a credible alternative. It is the cheapest negotiating position you will ever buy, and the only configuration that survives a vendor deprecating a model you depend on.

Access is not the constraint. Adoption is.

The labour data moved from anecdote to payroll this month: roughly 101,700 AI-attributed US job cuts year to date — about 23% of all 2026 layoffs, and the single most-cited reason for four consecutive months. Stanford finds early-career roles in AI-exposed occupations down around 13% since 2022.

The counter-signal deserves equal weight. In a survey of 350 executives, Gartner found firms that cut headcount for AI reported the same financial gains as firms that did not. Replacement is not where the return is showing up — which makes the figure below the more interesting one.

Figure 5

Eighty-four in every hundred people have never used an AI model

The true proportion

84in every 100 people have never used an AI model

The three paying tiers are the sliver on the right. On a true scale they are almost nothing — which is the finding, so they are not inflated to be visible.

The same four numbers, by depth of use

  1. Never used AI

    6.8B
  2. Free chatbot user

    1.3B
  3. Pays ~$20/mo

    20M
  4. Uses a coding scaffold

    3.5M

Logarithmic. Each step down is a fall of roughly an order of magnitude — 6.8 billion to 3.5 million across the whole range.

The same four tiers on two scales: the true linear proportion, and a logarithmic view by depth of use. From 6.8 billion non-users down to a few million people running agentic coding tools is a fall of more than three orders of magnitude.

Source: population penetration estimate, February 2026. Order-of-magnitude only.

Table view — every value in this figure
TierPeopleShare of 8.1B
Never used AI6.8B83.95%
Free chatbot user1.3B16.05%
Pays ~$20/mo20M0.25%
Uses a coding scaffold3.5M0.04%

So whatNobody has won this yet. Model access is a commodity your competitors have too; what they mostly do not have is an organization that can absorb it. Treat capability diffusion inside your own company as the scarce input, because it is.

Sources & method

Artificial Analysis
Intelligence Index, output speed, and cost per index task. Figures 1, 2 and 3.
VulcanBench Eval Suite 3
23 frontier-hard software-engineering tasks drawn from merged open-source pull requests, run in Docker-sandboxed agent harnesses, 1 August 2026. Figure 4.
Challenger, Gray & Christmas
US job-cut attribution, year to date through June 2026.
Stanford
Early-career employment in AI-exposed occupations.
Gartner
Survey of 350 executives on AI-driven headcount reduction.

Every chart here is redrawn from the underlying values rather than reproduced as an image, so each carries a table view with the full data. Where a source figure was internally inconsistent, the affected row is excluded and the exclusion is stated on the chart.

Compiled by AWALI from published benchmark data covering July–August 2026. Benchmark scores reflect specific task suites and harnesses and should not be read as general capability claims. Figures marked as estimates are order-of-magnitude only.

With thanks to VJAL Institute — Aditya Berlia.