AWALI Thought leadershipQ4 2026
Sources

AWALI · Q4 2026 briefing for CEOs and executive teams

  1. AI got ~12× more capable.

    Longer autonomous task horizon

    The length of software task frontier models complete at 50% success grew from about 1 hour to about 12 hours in a year.

    BenchmarkMETR
  2. Equivalent performance got ~13× cheaper.

    Cheaper each year to reach the same AI performance

    The cost of reaching a fixed benchmark score falls about 13× a year when users switch to the cheapest capable model.

    ResearchEpoch AI
  3. Nearly 9 in 10 organizations now use AI somewhere in the business.

    Organizations already use AI somewhere in the business

    Share of organizations using AI in at least one business function, not necessarily at scale.

    SurveyMcKinsey & Company
  4. Yet only 37% report any EBIT impact.

    Report any EBIT impact from AI

    Slightly fewer than a year earlier, and most who report it attribute less than 5% of EBIT to AI.

    SurveyMcKinsey & Company

So the bottleneck is shifting away from the AI.The constraint is moving inside the enterprise.

Access to intelligence is becoming abundant. Turning it into enterprise value is not.

Q4 2026

The enterprise becomes programmable.

8 shifts. 6 decisions. 90 days.

One operating system for AI. Each of the eight shifts changes one layer.

Governance surrounds it, enforced while it runs. Identity · Authority · Permissions · Observability · Reversibility

Intelligence is becoming abundant. Execution is not.

8 shifts are changing this system. 6 require leadership decisions now.

Explore the briefing

Composed and engineered by Scott Thomas, AWALI

The programmable enterprise: models, context, decisions, loops, workflows, voice, software and creative interfaces, and execution, surrounded by governance: identity, authority, permissions, observability and reversibility.

For the CEO and executive team · the briefing on one page

What you can do over the next 90 days

Eight shifts are changing how companies get value from AI. Six call for a leadership decision this quarter: a commitment of budget, policy, ownership or workflow change. Two are watch areas: set a baseline and a trigger now, and decide later. Each one is collected in your 90-day agenda.

  1. 01Intelligence becomes infrastructureThe best AI model depends on the task, not on the leaderboard. 11×difference in cost between the lowest and highest effort settings of a single frontier modelBenchmarkArtificial Analysis Decision 1Set up task-level model routing
  2. 02From loops to graphsPrompts improve an interaction. Loops automate a job. Graphs redesign a workflow. 19→44%accuracy after a test-and-retry step was added, with no change of modelResearchRidnik et al. · Jan 2024 Decision 2Redesign one cross-functional workflow
  3. 03Routine decisions can run on AIAI can now make routine business decisions, like matching an invoice or routing a support ticket, for about two cents per thousand. $0.02to process 1,000 routine decisions with a purpose-built decision model (TypeSafe's Jev), at list priceAWALI calculation · TypeSafe, OpenAI & Anthropic list prices Decision 3Automate one high-volume decision process
  4. 04The interface recedesVoice is becoming a practical way to get work done in enterprise systems. 125,000+businesses use Wispr Flow for voice dictation, including most of the Fortune 500PressPulse 2.0 · Aug 2026 Decision 4Choose one transaction to run by voice
  5. 05Custom software becomes rational againFor differentiated workflows, custom software is economically viable again, and it increasingly deserves to be tested against SaaS. 32%of organizations have dropped at least one software purchase because they built the capability themselvesSurveyMcKinsey & Company · Aug 2026 Decision 5Re-evaluate one differentiated workflow for build versus buy
  6. 06Creative becomes a production systemCreative AI is becoming an orchestrated production system, and Higgsfield is one of the clearest examples. ~$500Mreported annual revenue run rate at Higgsfield, which licenses most of its models from othersReported, unverifiedTechTimes · Jun 2026 WatchBaseline creative testing speed
  7. 07AI is changing faster than annual plans can respondAWALI tracked 44 significant AI developments in the first nine months of 2026, including 25 major model releases. That is more change than a once-a-year plan is designed to absorb. ~4.3monthsfor the length of task AI can complete to double, based on independent measurementBenchmarkMETR · Jan 2026 WatchInstrument your AI planning cadence
  8. 08AI agents need limits built inEnterprise exposure is determined less by an agent's raw capability than by the authority, access and controls surrounding it. 34%of organizations apply the same security controls to AI agents as to employeesSurveyOkta · May 2026 Decision 6Set real-time controls for AI agents

01Intelligence becomes infrastructureLayer: ModelsDecision

The best AI model depends on the task, not on the leaderboard.

In this illustrative routing scenario, sending every task to the top model costs up to 299× more than sending each task to the least expensive model that meets its quality requirement. Choose a task to see where it goes and what it costs.

Figure 01

Five tasks, four models, three providers

An AWALI illustrative routing policy applied to Artificial Analysis benchmark data from 2 October 2026. Each task goes to the least expensive model that meets its quality requirement.

AWALI scenario · routing policy on Artificial Analysis dataBenchmarkArtificial Analysis
AWALI illustrative routing policy applied to the Q4 Artificial Analysis snapshot. Top option: Claude Opus 5.5 (max), $5.98 per task.
TaskComplexity / riskLatency sensitivityCost sensitivityQuality barRouted toIndexCost per taskCheaper than top option
Classify support tickets1 / 2highhigh22.5GPT-6 Luna · medium30$0.020299×
Real-time voice agent2 / 4highmed34.5MiMo-V2.6-Pro · default46$0.1346×
Review supplier contracts4 / 4lowhigh48.5GPT-6.1 Sol · high50$0.3219×
Draft a board memo4 / 5lowlow53.0Claude Opus 5.5 · high54$1.823×
Fix a production bug5 / 5lowlow60.0Claude Opus 5.5 · max (best available)58$5.98—
Artificial Analysis Intelligence Index v4.3.2 · Cost per index task (USD) · as of 2 Oct 2026.
ModelProviderOpen weightsEffortIndexCost per task
GPT-6 LunaOpenAINolow22$0.0045
GPT-6 LunaOpenAINonone18$0.010
GPT-6 LunaOpenAINomedium30$0.020
GPT-6 LunaOpenAINohigh33$0.030
GPT-6 LunaOpenAINoxhigh35$0.040
GPT-6 LunaOpenAINomax38$0.070
GPT-6.1 SolOpenAINolow42$0.13
MiMo-V2.6-ProXiaomiYesdefault46$0.13
DeepSeek V4.1 FlashDeepSeekYesnone25$0.15
GPT-6.1 SolOpenAINomedium48$0.21
GLM-5.3-FlashZ AIYesdefault42$0.25
DeepSeek V4.1 FlashDeepSeekYesmax39$0.27
GPT-6.1 SolOpenAINohigh50$0.32
GPT-6.1 SolOpenAINoxhigh51$0.39
Claude Opus 5.5AnthropicNolow42$0.55
Claude Sonnet 5.5AnthropicNomedium41$0.59
GPT-6.1 SolOpenAINomax52$0.72
GPT-6 AstraOpenAINolow46$0.82
GLM-5.3Z AIYeslow34$0.85
Gemini 3.8 FlashGoogleNomedium40$0.93
Claude Sonnet 5.5AnthropicNohigh47$1.08
Kimi K3KimiYeslow30$1.15
Gemini 3.8 FlashGoogleNohigh41$1.24
Claude Opus 5.5AnthropicNomedium51$1.34
GPT-6 AstraOpenAINomedium50$1.54
GPT-6 AstraOpenAINohigh51$1.73
Claude Opus 5.5AnthropicNohigh54$1.82
Gemini 4 ArgonGoogleNohigh53$1.99
Kimi K3KimiYesmax44$2.00
GLM-5.3Z AIYesmax45$2.01
GPT-6 AstraOpenAINoxhigh52$2.31
Claude Sonnet 5.5AnthropicNoxhigh52$2.74
GPT-6 AstraOpenAINomax53$3.26
Claude Opus 5.5AnthropicNoxhigh56$3.46
Claude Opus 5.5AnthropicNomax58$5.98
Claude Sonnet 5.5AnthropicNomax56$7.62
Gemini 3.8 FlashGoogleNolow33not published
Evidence
11×

difference in cost between the lowest and highest effort settings of a single frontier model

BenchmarkArtificial Analysis
81%

of large enterprises already use models from three or more providers

SurveyAndreessen Horowitz · Jan 2026
What this means

Model strategy is becoming a routing and economics discipline. A single standard model tends to overpay for routine work and underperform on complex work, and at enterprise volumes that gap can be a large share of the AI budget.

Decision 1 of 6 · 90-day agenda

Set up task-level model routing

For your three highest-volume AI workflows, define the quality each one requires. Then test at least two providers, three effort settings and one open-weight model, using your own data. Review the routing every month.

Suggested owner: CTO
Go deeperWhat's new since Q3, and the research behind this chapter

Two trends are running at the same time. Epoch AI estimates that reaching a given level of AI performance gets about 13 times cheaper each year. A separate academic study finds that running the very newest models keeps getting more expensive. Routing is how a company captures the first trend without paying for the second.

A single model's cost can vary by up to 15 times depending on how much reasoning effort it is allowed to use. Often the cheapest option that meets the quality requirement is a frontier model at a low setting, or a smaller model at a high one.

Open-weight models now score about 12 points below the top of the benchmark at a small fraction of the price. MiMo-V2.6-Pro, for example, scores 46 at $0.13 per task, compared with 58 at $5.98 for the top setting. Because these models can run on your own infrastructure, they also address most data-residency concerns and make costs more predictable at high volume. On Vercel's gateway in June, they handled 29% of traffic for less than 4% of spending.

SourcesResearchEpoch AI · Sep 2026ResearchGundlach · Mar 2026BenchmarkArtificial AnalysisMarket dataConstellation · Vercel data · Aug 2026

02From loops to graphsLayer: Loops · WorkflowsDecision

Prompts improve an interaction. Loops automate a job. Graphs redesign a workflow.

A large share of the reliability gain comes from adding a validation step: the system checks its own output against something it knows is correct and tries again if the check fails. In one study, that step alone raised accuracy from 19% to 44% using the same model.

Figure 02

As capability increases, so does complexity

The same disputed supplier invoice handled three ways: a single prompt, a loop and a full workflow. Keep scrolling to add each layer, or drag and use the control below.

Vendor-reportedAnthropic · Dec 2024ResearchRidnik et al. · Jan 2024
Diagram facts for the illustrative invoice-dispute workflow.
PromptLoopGraph
What it doesImproves one interactionAutomates a jobRedesigns a workflow
Model calls11 + retries1 + retries per loop
Checks01, with a stop condition2 loops
Systems and records003
People001
Permission gates001
Evidence
19→44%

accuracy after a test-and-retry step was added, with no change of model

ResearchRidnik et al. · Jan 2024
3 in 4

of the companies getting the most value from AI have redesigned their workflows, compared with 1 in 4 of other companies

SurveyMcKinsey & Company · Aug 2026
40%

of large organizations are now scaling AI agents, up from 27% a year ago

SurveyMcKinsey & Company · Aug 2026
What this means

Much of the reliability gain comes from the loop, not from the workflow around it. If loops are connected into a larger workflow before each one works reliably, their errors compound.

Decision 2 of 6 · 90-day agenda

Redesign one cross-functional workflow

Choose one workflow that spans at least two teams. Make sure each automated step checks its own output against a written definition of a correct result, and only then connect the steps into a single workflow.

Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter

A prompt produces an answer. A loop checks that answer against a reliable reference, such as the purchase order, the contract or a test, and tries again until it passes or reaches a set limit. A workflow, sometimes called a graph, connects several loops with systems of record, approvals and people.

The evidence supports building in that order. In one coding study, adding a test-and-retry loop raised accuracy from 19% to 44% with the same model. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027 because of cost, unclear value or weak risk controls.

Nearly three quarters of the organizations getting the most value from AI have redesigned their workflows. Among all other organizations, about a quarter have.

SourcesResearchRidnik et al. · Jan 2024Vendor-reportedAnthropic · Dec 2024ForecastGartner · Jun 2025SurveyMcKinsey & Company · Aug 2026

03Routine decisions can run on AILayer: Context · DecisionsDecision

AI can now make routine business decisions, like matching an invoice or routing a support ticket, for about two cents per thousand.

For routine decisions, model cost is rarely the binding constraint now. What leaders need to decide is how sure the AI must be before it acts on its own: a higher bar sends more cases to people, and a lower bar accepts more mistakes.

Figure 03

What happens to 10,000 decisions a day

The model rates its confidence in each decision. Above the threshold, the decision is automated; below it, the case goes to a person. Move the threshold to see the effect on workload, cost and errors.

AWALI calculation · TypeSafe, OpenAI & Anthropic list pricesVendor-reportedTypeSafe · Sep 2026Price listOpenAIPrice listAnthropic
Per day, 10,000 decisions. All manual: 667 review hours, $40,000 labor. Illustrative: confidence profiles, 4 minutes per review, $60 per hour loaded.
DecisionThresholdAutomatedTo a personReview hoursReview laborLikely wrong
Invoice match97%7,9102,090139$8,36018
Invoice match90%8,7451,25584$5,02167
Invoice match80%9,26573549$2,941143
Ticket routing97%6,2573,743250$14,97227
Ticket routing90%7,6472,353157$9,411109
Ticket routing80%8,5841,41694$5,665245
Churn risk97%3,9156,085406$24,34129
Churn risk90%5,8484,152277$16,608146
Churn risk80%7,3682,632175$10,528368
AWALI calculation at list prices: 500 input and 20 output tokens per decision.
ModelList price per 1M tokens (in / out)Per 1,000 decisionsPer daySource
TypeSafe Jev$0.042 / $0.00$0.021$0.21TypeSafe · Sep 2026
GPT-6 Luna$0.100 / $0.50$0.060$0.60OpenAI
Claude Opus 5.5$4.000 / $20.00$2.400$24.00Anthropic
Evidence
$0.02

to process 1,000 routine decisions with a purpose-built decision model (TypeSafe's Jev), at list price

AWALI calculation · TypeSafe, OpenAI & Anthropic list prices
114×

how much more the same 1,000 decisions cost on a frontier model

AWALI calculation · TypeSafe, OpenAI & Anthropic list prices
63→95%

accuracy on one task when a single question was broken into five smaller ones

ResearchBeri · Sep 2026
What this means

In every scenario in this simulation, the labor saved is hundreds to hundreds of thousands of times larger than the model cost. The real limits are how many errors the business can tolerate and how much review capacity it keeps.

Decision 3 of 6 · 90-day agenda

Automate one high-volume decision process

List your 20 most frequent recurring decisions and choose one. Test the model on 1,000 to 2,000 past decisions, set the threshold based on the error rate you measure, and assign someone to own the queue of cases that go to people.

Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter

At list prices, 1,000 short decisions (about 500 tokens of input each) cost $0.02 with Jev, $0.06 with GPT-6 Luna, $0.60 with Claude Haiku 4.5 and $2.40 with Claude Opus 5.5. These prices say nothing about accuracy, which has to be measured on your own cases.

Accuracy depends largely on how the decision is set up. In an independent test, Jev was 62.6% accurate when asked a single question and 95.0% accurate when the same decision was broken into five questions weighted using labeled data. The same analysis recommends testing on 1,000 to 2,000 past decisions and setting a separate threshold for each question.

Where to set the threshold is a business decision about acceptable error. It should be set, documented and reviewed in the same way as a credit limit.

SourcesAWALI calculation · TypeSafe, OpenAI & Anthropic list pricesVendor-reportedTypeSafe · Sep 2026ResearchBeri · Sep 2026Price listOpenAIPrice listAnthropic

04The interface recedesLayer: VoiceDecision

Voice is becoming a practical way to get work done in enterprise systems.

Speech models now respond in about a second, close to the pace of normal conversation. The most valuable use is a single spoken request that gets a task done across company systems.

How people use voice today
  1. DictationEmails, documents and AI prompts, spoken. About three times faster than typing.
  2. The foot pedalHold to talk, release to send, with hands kept on the keyboard.
  3. Thinking out loudVoice mode in ChatGPT, Claude or Gemini to prepare and draft on the move.
  4. Hands-busy workTechnicians, clinicians and drivers update records without a screen.
  5. Customer voice agentsService lines that resolve simple requests, with a person on hand.
  6. Requests that get doneOne instruction that updates several systems. Shown below.
PressPulse 2.0 · Aug 2026PractitionerElgato · Jul 2026BenchmarkArtificial Analysis
Figure 04

From a spoken request to a completed transaction

One spoken request, traced from speech to updated company systems. A simulation; no microphone needed.

BenchmarkArtificial AnalysisPrice listOpenAI · Sep 2026
Simulated request: "Move tomorrow's client meeting to the afternoon and send them the updated prep package." Timings are representative, not measured.
StepStageStart (ms)End (ms)What happens
ListenTranscription03500Audio streams in while the person speaks.
Transcribe (streaming)Transcription2003600Speech becomes text as it arrives, so the system can start before the sentence ends.
Understand intentConversation36003780The sentence is turned into two specific actions, each with the details needed to carry it out.
Find the right recordsConversation37803990The calendar, the CRM and the document library are searched, so ‘the client’ and ‘the prep package’ are matched to real records.
Check permissionsTransaction39904060The assistant is allowed to reschedule meetings the company organizes and to share documents already approved for this client. Anything outside those rules would stop here and go to a person.
Update calendarTransaction40604330The meeting moves to the first afternoon slot that works for every attendee.
Retrieve documentsTransaction43304470The latest approved version of the prep package is selected, not an old draft.
Prepare communicationTransaction44704760A short note to the client is drafted from an approved template, with the new time and the package attached.
Complete and confirmTransaction47605200The message is sent, every step is logged, and the person hears a summary of what was done.
Evidence
125,000+

businesses use Wispr Flow for voice dictation, including most of the Fortune 500

PressPulse 2.0 · Aug 2026
~1.2s

for the leading speech models to start responding

BenchmarkArtificial Analysis
87%

of customers say companies should still give them access to a person

SurveyGartner · Aug 2026
What this means

Speaking takes 3.6 of the 5.2 seconds. The value is in the 1.6 seconds after, when the system finds the records, checks permissions and makes the updates. Voice pays off mainly where those connections already exist.

Decision 4 of 6 · 90-day agenda

Choose one transaction to run by voice

Pick one transaction where typing or navigating screens slows people down, in field service, sales, operations or customer service. Build it end to end, including the permission check and a handoff to a person.

Suggested owner: Chief Customer Officer
Go deeperWhat's new since Q3, and the research behind this chapter

The leading speech models now start responding within 1.1 to 1.4 seconds and score above 90% on a spoken-reasoning benchmark. That is close to the pace of normal conversation, which makes voice practical for customers, field technicians and sales teams.

Most of the work happens after the person stops talking. The request has to be turned into a structured instruction, matched to the right records, checked against policy, recorded in the system of record and confirmed back to the user. The trace above shows each step and how long it takes.

Customer service leaders increased AI spending by 38% while their overall budgets grew 2%. Customers have one condition: 87% want to be able to reach a person.

SourcesBenchmarkArtificial AnalysisSurveyGartner · Aug 2026SurveyGartner · Aug 2026

05Custom software becomes rational againLayer: SoftwareDecision

For differentiated workflows, custom software is economically viable again, and it increasingly deserves to be tested against SaaS.

AI has significantly reduced the cost of writing software. It has not reduced the cost of testing it, maintaining it and owning it over time.

Figure 05

Where should we stop standardizing?

An AWALI illustrative framework. Move a workflow from standard to differentiating and compare the relative five-year cost of buying it with building it. The bars show relative cost, not measured cost.

SurveyMcKinsey & Company · Aug 2026ExperimentMETR · Jul 2025SurveyGoogle Cloud · Sep 2025
AWALI framework. Cost bars in the figure are relative, not measured.
Workflow typeExamplesStrategyWhy
Systems of recordGeneral ledger, payroll, ERP coreBuyEvery company needs the same thing here, and the vendor carries the compliance burden. Buy it and manage it through procurement.
Standard workflowsHR service desk, expense approvalConfigureThe process is much the same across companies. Buy a platform, configure it to fit, and make sure your data stays portable.
Competitive workflowsPricing approval, field dispatch, underwritingBespokeHow you run this process is part of how you compete, and off-the-shelf software forces you to work like everyone else. With AI, building it yourself now makes financial sense.
Core judgmentYour models of risk, demand and customer valueOwnThis is where your advantage comes from. Own the whole system, keep it out of vendors' training data, and invest most heavily in testing.
Evidence
32%

of organizations have dropped at least one software purchase because they built the capability themselves

SurveyMcKinsey & Company · Aug 2026
−19%

measured change in speed for experienced developers using AI, even though they believed they were 20% faster

ExperimentMETR · Jul 2025
What this means

The real difference is what the software touches. Tools that sit on top of existing data are low risk to build in-house. Anything that writes to systems of record, touches customers or money, or has to behave the same way across the business needs a proper data and governance foundation first.

Decision 5 of 6 · 90-day agenda

Re-evaluate one differentiated workflow for build versus buy

Pick one strategically differentiating workflow that a SaaS product constrains, and compare building and buying on full cost, including testing, maintenance and ownership. Decide which parts your team builds and which foundation layers, such as data and governance, come from a specialist partner or platform. Apply the same review to software renewals above a set value.

Suggested owner: CIO
Go deeperWhat's new since Q3, and the research behind this chapter

The research on developer productivity is mixed, and the differences are informative. On new, well-defined tasks, developers were 21% to 55% faster. In METR's 2025 trial, experienced developers working in their own mature codebases were 19% slower, even though they believed they were 20% faster. Google's DORA research links AI use to faster delivery but less stable systems.

The practical conclusion is to build where the logic is specific to your business and the scope is clear, and to fund the review, testing and ownership needed to keep it reliable.

SourcesExperimentGitHub · Sep 2022ExperimentGoogle · Oct 2024ExperimentMETR · Jul 2025ExperimentMETR · Feb 2026SurveyGoogle Cloud · Sep 2025

Figure 05b

The productivity gain depends on the task, and developers tend to overestimate it

Reported effect of AI coding tools in controlled trials. Positive values mean faster work or more output. Each study measures something slightly different.

ExperimentGitHub · Sep 2022ExperimentCui · Feb 2025ExperimentGoogle · Oct 2024ExperimentMETR · Jul 2025ExperimentMETR · Feb 2026
Positive values mean faster or more output. METR 2025 participants believed they had been 20% faster.
StudyYearTaskMeasureEffect95% interval (pts)
GitHub Copilot RCT2022One new task (HTTP server)Faster task completion+55%+21 to +89Source
Microsoft, Accenture, Fortune 100 field trials2025Day-to-day workMore completed tasks+26%+6 to +46Source
Google enterprise RCT2024Bounded enterprise taskLess time on task+21%not reportedSource
METR, experienced developers2025Own mature repositoriesChange in speed (time +19%)-19%-39 to -2Source
METR follow-up, returning developers2026Own repositoriesLess time on task+18%-9 to +38Source
METR follow-up, new developers2026Own repositoriesLess time on task+4%-9 to +15Source

06Creative becomes a production systemLayer: CreativeWatch area

Creative AI is becoming an orchestrated production system, and Higgsfield is one of the clearest examples.

Higgsfield puts the leading video and image models, camera control, consistent characters and ad templates into one workflow. A small team can now produce, and test, far more creative than a traditional production schedule allows.

Figure 06

Higgsfield at a glance

AWALI example · made with Claude, Blender and Higgsfield
Scene
The Teton Range from Driggs, Idaho, in winter, at 5,000 ft
Terrain
Built from USGS elevation data, so the mountains match the real topography
Tools
Claude and Blender built the 3D scene; Higgsfield generated the finished video
Build
A quick build that combined several models in one workflow

Workflow based on a guide by @kyleslyf on YouTube

How teams use it
  1. Write a short shot list and set the look with one key frame from an image model.
  2. Use saved camera presets instead of describing the move in words.
  3. Send wide shots and B-roll to a lower-cost model; keep the best model for the one or two hero shots.
  4. Train a Soul ID if the same presenter or character appears across the campaign.
  5. Generate versions by format and message, then cut, grade and caption in your usual editor.
  6. Test the versions and track the cost of each ad that beats the current one.

Clips run up to 15 seconds per generation, final assembly still happens in an editor, and the capability descriptions are Higgsfield's own.

What it does well
  • Many models, one placeVeo 3.1, Kling 3.0, Seedance 2.0, WAN 2.6 and others sit alongside Higgsfield's own models, so each shot can use the model that suits it.
  • Camera controlDolly, orbit, tracking and crane moves are set when the clip is generated, rather than hoped for in a text prompt.
  • Consistent peopleSoul ID is trained once from 20 or more photos and keeps the same face across every clip and model.
  • From product page to adMarketing Studio reads a product image or link and fills templates for product shots, Meta and Google ads, short UGC-style video and motion graphics.
  • Editing in placeInpainting, relighting, upscaling and layers change a scene without regenerating it from scratch.
Vendor-reportedHiggsfield · Jul 2026Vendor-reportedHiggsfield · Aug 2026PractitionerPractitioner guide · Jun 2026Vendor-reportedHiggsfield · website
Evidence
~$500M

reported annual revenue run rate at Higgsfield, which licenses most of its models from others

Reported, unverifiedTechTimes · Jun 2026
1,500+

ad templates in Higgsfield's Marketing Studio, which builds ads from a product image or product-page link

Vendor-reportedHiggsfield · Aug 2026
~$3–12

base model-generation cost for 30 seconds of video, before retries, editing, review and production overhead

AWALI calculation · Google & Runway list prices
What this means

When producing a campaign's worth of versions takes days rather than months, the useful measures change: the cost of each ad that beats the current one, and how quickly results feed the next brief.

Watch area · baseline and trigger, no major commitment yet

Baseline creative testing speed

Measure how many ad versions you test each month and how many days it takes to act on results. Then run one bounded test in a single channel with a production platform such as Higgsfield. No enterprise commitment is needed yet.

Suggested owner: CMO
Go deeperWhat's new since Q3, and the research behind this chapter

Generated video lists at $0.10 to $0.40 a second, so 30 seconds costs $3 to $12 in model fees before retries, editing and review. Production cost is no longer the main constraint.

Audience reaction still has to be tested rather than assumed. In IAB's survey, 82% of advertising executives believed young consumers feel positive about AI ads; 45% of those consumers do.

SourcesPrice listGoogle CloudPrice listRunwaySurveyIAB · Jan 2026

07AI is changing faster than annual plans can respondLayer: OperationsWatch area

AWALI tracked 44 significant AI developments in the first nine months of 2026, including 25 major model releases. That is more change than a once-a-year plan is designed to absorb.

Under an annual plan, each of those developments waited an average of 204 days before anyone could decide how to respond. With a monthly review, the average wait drops to 13 days.

Figure 07

How long each AI development waits for a decision, by planning cycle

Significant AI developments in 2026, by date, compared with three planning cycles. Choose a cycle to see how long each change waits before anyone can act on it.

Market dataScriptByAI · Sep 2026BenchmarkMETR · May 2026Vendor-reportedOpenAIVendor-reportedOpenAI Help CenterSurveyMcKinsey & Company · Aug 2026
Material AI changes in 2026 to 2 October. Kling 3.0 is dated to February 2026 by Kuaishou; the day shown is approximate. Anthropic's >80% figure is as of May 2026, published June 2026. Model releases are the flagship releases of OpenAI, Anthropic, Google, xAI and Meta plus the leading open-weight releases; dates follow public release trackers, using the first public availability where trackers differ.
DateAreaChange
2026-01-27Model releasesKimi K2.5 (Moonshot AI, open-weight)Source
2026-01-29Coding capabilityMETR: task length doubling every ~4.3 monthsSource
2026-02-05Model releasesClaude Opus 4.6Source
2026-02-15Creative modelsKling 3.0Source
2026-02-17Enterprise tooling & rulesNIST agent standards initiativeSource
2026-02-19Model releasesGemini 3.1 ProSource
2026-02-23Enterprise tooling & rulesgpt-realtime-1.5Source
2026-03-05Model releasesGPT-5.4Source
2026-03-31Creative modelsVeo 3.1 LiteSource
2026-04-07Model releasesClaude Mythos PreviewSource
2026-04-08Model releasesMeta Muse SparkSource
2026-04-15AgentsGartner: 17% have deployed agentsSource
2026-04-16Model releasesClaude Opus 4.7Source
2026-04-22Coding capabilityGoogle: 75% of new code AI-generatedSource
2026-04-23Model releasesGPT-5.5Source
2026-04-24Model releasesDeepSeek V4 (open-weight)Source
2026-04-26Creative modelsSora app discontinuedSource
2026-05-07Enterprise tooling & rulesGPT-Realtime-2Source
2026-05-08Coding capabilityMETR: frontier beyond 16 h, limit of measurementSource
2026-05-19Creative modelsGemini Omni FlashSource
2026-05-28Model releasesClaude Opus 4.8Source
2026-06-09Model releasesClaude Mythos 5 preview and Fable 5Source
2026-06-15Coding capabilityAnthropic: >80% of merged code by ClaudeSource
2026-06-26Model releasesGPT-5.6 SolSource
2026-06-30Model releasesClaude Sonnet 5Source
2026-07-06Enterprise tooling & rulesgpt-realtime-2.1Source
2026-07-08Model releasesGrok 4.5Source
2026-07-16Model releasesKimi K3 (Moonshot AI, open-weight)Source
2026-07-24Model releasesClaude Opus 5Source
2026-07-27Enterprise tooling & rulesEU AI omnibus in forceSource
2026-08-02Model releasesQwen3.8 Max (Alibaba)Source
2026-08-25AgentsMcKinsey: 40% of large firms scaling agentsSource
2026-09-03Model releasesGPT-6 AstraSource
2026-09-06AgentsOpenAI: automated research intern goal metSource
2026-09-07Enterprise tooling & rulesBenchmark index v4.3 resets comparisonsSource
2026-09-10Enterprise tooling & rulesGPT-Live-1 voice APISource
2026-09-21Model releasesGrok 4.7Source
2026-09-22Model releasesClaude Opus 5.5Source
2026-09-22Model releasesGPT-6 Sol and GPT-6 LunaSource
2026-09-22Model releasesMiMo-V2.6-Pro (Xiaomi, open-weight)Source
2026-09-24Creative modelsSora API discontinuedSource
2026-09-28Model releasesClaude Sonnet 5.5Source
2026-09-29Model releasesGPT-6.1 SolSource
2026-09-30Model releasesGemini 4 ArgonSource
Evidence
~4.3months

for the length of task AI can complete to double, based on independent measurement

BenchmarkMETR · Jan 2026
11days

median time between major model releases in 2026 (reported)

Market dataOfficeChai · Apr 2026
What this means

Annual planning still works for setting strategy, and quarterly cycles still work for allocating capital. AI operations increasingly need a decision point every month.

Watch area · baseline and trigger, no major commitment yet

Instrument your AI planning cadence

Log material AI developments and how long each waits for a decision. Set a trigger: if more than one a month affects your priorities, add a 30-day AI operating review alongside annual strategy and quarterly capital allocation.

Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter

The length of software task AI can complete went from about an hour in early 2025 to about 12 hours by February 2026. Since 2023 it has doubled roughly every 4.3 months. METR notes that its tests can barely measure the leading models, and that the task length at 80% reliability is much shorter.

AI labs report that AI now writes most of their own code: more than 80% of merged code at Anthropic and 75% of new code at Google. Both figures are self-reported and measured differently.

Something that was out of reach when the budget was set can become routine before the end of the year. A monthly review needs a clear trigger for revisiting work when that happens.

SourcesBenchmarkMETR · Jan 2026BenchmarkMETR · May 2026Vendor-reportedAnthropic · Jun 2026Vendor-reportedGoogle · Apr 2026ResearchMETR · Jul 2026AWALI · Apr 2026

Figure 07b

The length of task AI can complete doubles every few months

METR's measure of the length of software task AI can complete with a 50% success rate, by model release date, on a log scale with 95% confidence intervals.

BenchmarkMETR · Jan 2026BenchmarkMETR · May 2026
METR 50% time horizon. Method versions differ between rows; see notes.
ModelReleased50% horizon (min)95% intervalNote
GPT-42023-03-143.51.6–6.9TH1.1METR
GPT-4 11062023-11-063.61.6–7.5TH1.1METR
Claude 3.7 Sonnet2025-02-246032–106TH1.1METR
o32025-04-1612174–201TH1.1METR
Claude Opus 42025-05-2210158–170TH1.1METR
GPT-52025-08-07214117–480TH1.1METR
Claude Sonnet 4.52025-09-2911350–235TH1.0 (Oct 2025 post)METR
Gemini 3 Pro2025-11-18240130–440TH1.1, before 3 Mar 2026 correctionMETR
GPT-5.1-Codex-Max2025-11-1916275–350TH1.0 (Nov 2025 post)METR
Claude Opus 4.52025-11-24320170–729TH1.1METR
Claude Opus 4.62026-02-05719wideCorrected value (~12 h); interval very wideMETR
Gemini 3.1 Pro2026-02-19384240–720Apr 2026 postMETR
GPT-5.4 (xhigh)2026-03-05342180–810Standard method, Apr 2026 postMETR
Claude Mythos Preview2026-04-07≥ 960510–3300Lower bound (≥ 16 h); at the limit of the task suiteMETR
GPT-5.6 Sol2026-06-26678300–2400Jun 2026; cheating attempts counted as failuresMETR

08AI agents need limits built inLayer: GovernanceDecision

Enterprise exposure is determined less by an agent's raw capability than by the authority, access and controls surrounding it.

Agents can now send email, move money and change records on their own. A written policy is unlikely to stop one in the middle of a task; controls built into the systems it uses can.

Figure 08

What one AI agent can do on its own, with and without controls

One operations agent with access to seven systems. The rings show how serious each action is, from reading data to actions that can't be undone. This is an illustration; nothing is executed.

StandardOWASP GenAI Security Project · Dec 2025StandardNIST · Feb 2026
Illustrative scenario: one operations agent with access to Email, CRM, Finance, Documents, Calendar, External tools, Customer comms. Nothing is executed.
ControlWhen onWhen removed
IdentityEvery action is signed by the agent's own identity, with a named owner.Actions run on a shared login, so nobody can tell which agent did what.
Permission scopeThe agent can reach only the systems and actions this job requires.The agent has broad access, so deleting records, wiring money and calling any API all become possible.
Approval thresholdsExternal emails, file shares and customer messages need a person's approval.Actions that can't be undone go ahead without anyone approving them.
Spending limitsPayments are capped at $1,000, and the agent can't make purchases or issue refunds.There is no limit on how much money the agent can move.
Rate limitsThe agent can take at most 20 actions a minute, so mistakes stay small.The agent can take thousands of actions a minute, so one bad instruction can repeat many times before anyone notices.
Audit loggingEvery action is recorded with who did it, what went in and what happened.Nothing is recorded, so incidents can't be reconstructed.
Human escalationUnclear or out-of-policy cases go to a named person.The agent decides for itself how to handle exceptions.
Kill switchAn operator can stop the agent and revoke its access in the middle of a task.Once a task is running, it can't be stopped.
Evidence
34%

of organizations apply the same security controls to AI agents as to employees

SurveyOkta · May 2026
>40%

of agentic AI projects are expected to be canceled by 2027, partly because of weak risk controls

ForecastGartner · Jun 2025
What this means

Only 34% of organizations apply the same controls to AI agents as to employees. The gap between a typical pilot and a governed agent is four controls, and each is far easier to put in place before launch than after an incident.

Decision 6 of 6 · 90-day agenda

Set real-time controls for AI agents

Make an inventory of every AI agent and automated account. Give each one a named owner, its own identity, spending and activity limits, approval thresholds, logging, an escalation path and a tested way to shut it down.

Suggested owner: CISO
Go deeperWhat's new since Q3, and the research behind this chapter

The necessary controls are well understood: a separate identity for each agent, only the access it needs, approval thresholds, logging, ongoing testing and the ability to revoke access immediately. Recent incidents show why they need to be enforced while systems run rather than written into policy. In 2025 a coding agent deleted a live production database during a code freeze, and AI carried out 80% to 90% of the work in a state-linked hacking campaign. In 2026, evaluation agents reached a third party's production systems.

Practice is behind. Only 34% of organizations apply the same controls to AI agents as to employees. OWASP's first list of top risks for agentic applications includes identity abuse, misuse of tools and agents acting outside their instructions, and NIST has started standards work on agent identity and authorization.

SourcesIncidentAI Incident Database · Jul 2025IncidentAnthropic · Nov 2025IncidentOpenAI; investigation by METR · Jul 2026SurveyOkta · May 2026StandardOWASP GenAI Security Project · Dec 2025StandardNIST · Feb 2026

Next steps

Turn these six decisions into your 90-day AI operating agenda.

Each of the six decisions has a suggested owner. In an AWALI workshop, we turn them into named owners, dates and a first operating review within 30 days.

Request a Workshop

Composed and engineered by Scott Thomas, AWALI

    Sources and method · 63 sources

    Where every figure comes from

    • Figures are reported as published by each source as of the date shown. This edition was compiled on 2 October 2026, and model and pricing data reflect that date.
    • Where a shorter summary would overstate a finding, we use the source's original, narrower wording. AWALI has not independently reproduced these studies.
    • Self-reported figures, press reports, forecasts and AWALI calculations are labeled wherever they appear.
    • Every chart is drawn from the values shown in its table view. Figures we could not confirm were left out rather than estimated.
    • The interactive scenarios (the routing policy, the decision simulator, the voice trace, the build-versus-buy comparison, the ad-testing example and the agent controls map) illustrate how these systems behave. They are not measurements.

    Independent benchmark 6

    Independent research 5

    RCT / experiment 5

    Survey 8

    Market data 4

    Press report 2

    Published pricing 6

    Vendor-reported 12

    Practitioner guidance 2

    Incident 3

    Regulation and standards 3

    AWALI illustrative scenario 1

    Prior AWALI briefing 2

    AWALI platform 1