Let the numbers land one at a time. Stop on 37%: everyone uses AI, few see value. Then the line that frames the whole briefing: the constraint has moved inside the enterprise.
For the CEO and executive team · the briefing on one page
What you can do over the next 90 days
Eight shifts are changing how companies get value from AI. Six call for a leadership decision this quarter: a commitment of budget, policy, ownership or workflow change. Two are watch areas: set a baseline and a trigger now, and decide later. Each one is collected in your 90-day agenda.
The best AI model depends on the task, not on the leaderboard.
In this illustrative routing scenario, sending every task to the top model costs up to 299× more than sending each task to the least expensive model that meets its quality requirement. Choose a task to see where it goes and what it costs.
Figure 01
Five tasks, four models, three providers
An AWALI illustrative routing policy applied to Artificial Analysis benchmark data from 2 October 2026. Each task goes to the least expensive model that meets its quality requirement.
Model strategy is becoming a routing and economics discipline. A single standard model tends to overpay for routine work and underperform on complex work, and at enterprise volumes that gap can be a large share of the AI budget.
Decision 1 of 6 · 90-day agenda
Set up task-level model routing
For your three highest-volume AI workflows, define the quality each one requires. Then test at least two providers, three effort settings and one open-weight model, using your own data. Review the routing every month.
Suggested owner: CTO
Go deeperWhat's new since Q3, and the research behind this chapter
Two trends are running at the same time. Epoch AI estimates that reaching a given level of AI performance gets about 13 times cheaper each year. A separate academic study finds that running the very newest models keeps getting more expensive. Routing is how a company captures the first trend without paying for the second.
A single model's cost can vary by up to 15 times depending on how much reasoning effort it is allowed to use. Often the cheapest option that meets the quality requirement is a frontier model at a low setting, or a smaller model at a high one.
Open-weight models now score about 12 points below the top of the benchmark at a small fraction of the price. MiMo-V2.6-Pro, for example, scores 46 at $0.13 per task, compared with 58 at $5.98 for the top setting. Because these models can run on your own infrastructure, they also address most data-residency concerns and make costs more predictable at high volume. On Vercel's gateway in June, they handled 29% of traffic for less than 4% of spending.
Start with ‘Classify support tickets’ and finish with ‘Fix a production bug’. The same policy sends five tasks to four different models.
02From loops to graphsLayer: Loops · WorkflowsDecision
Prompts improve an interaction. Loops automate a job. Graphs redesign a workflow.
A large share of the reliability gain comes from adding a validation step: the system checks its own output against something it knows is correct and tries again if the check fails. In one study, that step alone raised accuracy from 19% to 44% using the same model.
Figure 02
As capability increases, so does complexity
The same disputed supplier invoice handled three ways: a single prompt, a loop and a full workflow. Keep scrolling to add each layer, or drag and use the control below.
Much of the reliability gain comes from the loop, not from the workflow around it. If loops are connected into a larger workflow before each one works reliably, their errors compound.
Decision 2 of 6 · 90-day agenda
Redesign one cross-functional workflow
Choose one workflow that spans at least two teams. Make sure each automated step checks its own output against a written definition of a correct result, and only then connect the steps into a single workflow.
Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter
A prompt produces an answer. A loop checks that answer against a reliable reference, such as the purchase order, the contract or a test, and tries again until it passes or reaches a set limit. A workflow, sometimes called a graph, connects several loops with systems of record, approvals and people.
The evidence supports building in that order. In one coding study, adding a test-and-retry loop raised accuracy from 19% to 44% with the same model. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027 because of cost, unclear value or weak risk controls.
Nearly three quarters of the organizations getting the most value from AI have redesigned their workflows. Among all other organizations, about a quarter have.
Read the three statements as you move through the states. In the loop state, point to the check and the stop condition. That is where the reliability comes from.
03Routine decisions can run on AILayer: Context · DecisionsDecision
AI can now make routine business decisions, like matching an invoice or routing a support ticket, for about two cents per thousand.
For routine decisions, model cost is rarely the binding constraint now. What leaders need to decide is how sure the AI must be before it acts on its own: a higher bar sends more cases to people, and a lower bar accepts more mistakes.
Figure 03
What happens to 10,000 decisions a day
The model rates its confidence in each decision. Above the threshold, the decision is automated; below it, the case goes to a person. Move the threshold to see the effect on workload, cost and errors.
In every scenario in this simulation, the labor saved is hundreds to hundreds of thousands of times larger than the model cost. The real limits are how many errors the business can tolerate and how much review capacity it keeps.
Decision 3 of 6 · 90-day agenda
Automate one high-volume decision process
List your 20 most frequent recurring decisions and choose one. Test the model on 1,000 to 2,000 past decisions, set the threshold based on the error rate you measure, and assign someone to own the queue of cases that go to people.
Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter
At list prices, 1,000 short decisions (about 500 tokens of input each) cost $0.02 with Jev, $0.06 with GPT-6 Luna, $0.60 with Claude Haiku 4.5 and $2.40 with Claude Opus 5.5. These prices say nothing about accuracy, which has to be measured on your own cases.
Accuracy depends largely on how the decision is set up. In an independent test, Jev was 62.6% accurate when asked a single question and 95.0% accurate when the same decision was broken into five questions weighted using labeled data. The same analysis recommends testing on 1,000 to 2,000 past decisions and setting a separate threshold for each question.
Where to set the threshold is a business decision about acceptable error. It should be set, documented and reviewed in the same way as a credit limit.
Start at ‘Balanced’. Move to ‘Cautious’ and point out that errors fall while review hours rise. Then switch models: the labor numbers don't change and the model cost barely moves.
04The interface recedesLayer: VoiceDecision
Voice is becoming a practical way to get work done in enterprise systems.
Speech models now respond in about a second, close to the pace of normal conversation. The most valuable use is a single spoken request that gets a task done across company systems.
How people use voice today
Personal productivityMost business value →
DictationEmails, documents and AI prompts, spoken. About three times faster than typing.
The foot pedalHold to talk, release to send, with hands kept on the keyboard.
Thinking out loudVoice mode in ChatGPT, Claude or Gemini to prepare and draft on the move.
Hands-busy workTechnicians, clinicians and drivers update records without a screen.
Customer voice agentsService lines that resolve simple requests, with a person on hand.
Requests that get doneOne instruction that updates several systems. Shown below.
Simulated request: "Move tomorrow's client meeting to the afternoon and send them the updated prep package." Timings are representative, not measured.
Step
Stage
Start (ms)
End (ms)
What happens
Listen
Transcription
0
3500
Audio streams in while the person speaks.
Transcribe (streaming)
Transcription
200
3600
Speech becomes text as it arrives, so the system can start before the sentence ends.
Understand intent
Conversation
3600
3780
The sentence is turned into two specific actions, each with the details needed to carry it out.
Find the right records
Conversation
3780
3990
The calendar, the CRM and the document library are searched, so ‘the client’ and ‘the prep package’ are matched to real records.
Check permissions
Transaction
3990
4060
The assistant is allowed to reschedule meetings the company organizes and to share documents already approved for this client. Anything outside those rules would stop here and go to a person.
Update calendar
Transaction
4060
4330
The meeting moves to the first afternoon slot that works for every attendee.
Retrieve documents
Transaction
4330
4470
The latest approved version of the prep package is selected, not an old draft.
Prepare communication
Transaction
4470
4760
A short note to the client is drafted from an approved template, with the new time and the package attached.
Complete and confirm
Transaction
4760
5200
The message is sent, every step is logged, and the person hears a summary of what was done.
Evidence
125,000+
businesses use Wispr Flow for voice dictation, including most of the Fortune 500
Speaking takes 3.6 of the 5.2 seconds. The value is in the 1.6 seconds after, when the system finds the records, checks permissions and makes the updates. Voice pays off mainly where those connections already exist.
Decision 4 of 6 · 90-day agenda
Choose one transaction to run by voice
Pick one transaction where typing or navigating screens slows people down, in field service, sales, operations or customer service. Build it end to end, including the permission check and a handoff to a person.
Suggested owner: Chief Customer Officer
Go deeperWhat's new since Q3, and the research behind this chapter
The leading speech models now start responding within 1.1 to 1.4 seconds and score above 90% on a spoken-reasoning benchmark. That is close to the pace of normal conversation, which makes voice practical for customers, field technicians and sales teams.
Most of the work happens after the person stops talking. The request has to be turned into a structured instruction, matched to the right records, checked against policy, recorded in the system of record and confirmed back to the user. The trace above shows each step and how long it takes.
Customer service leaders increased AI spending by 38% while their overall budgets grew 2%. Customers have one condition: 87% want to be able to reach a person.
For differentiated workflows, custom software is economically viable again, and it increasingly deserves to be tested against SaaS.
AI has significantly reduced the cost of writing software. It has not reduced the cost of testing it, maintaining it and owning it over time.
Figure 05
Where should we stop standardizing?
An AWALI illustrative framework. Move a workflow from standard to differentiating and compare the relative five-year cost of buying it with building it. The bars show relative cost, not measured cost.
AWALI framework. Cost bars in the figure are relative, not measured.
Workflow type
Examples
Strategy
Why
Systems of record
General ledger, payroll, ERP core
Buy
Every company needs the same thing here, and the vendor carries the compliance burden. Buy it and manage it through procurement.
Standard workflows
HR service desk, expense approval
Configure
The process is much the same across companies. Buy a platform, configure it to fit, and make sure your data stays portable.
Competitive workflows
Pricing approval, field dispatch, underwriting
Bespoke
How you run this process is part of how you compete, and off-the-shelf software forces you to work like everyone else. With AI, building it yourself now makes financial sense.
Core judgment
Your models of risk, demand and customer value
Own
This is where your advantage comes from. Own the whole system, keep it out of vendors' training data, and invest most heavily in testing.
Evidence
32%
of organizations have dropped at least one software purchase because they built the capability themselves
The real difference is what the software touches. Tools that sit on top of existing data are low risk to build in-house. Anything that writes to systems of record, touches customers or money, or has to behave the same way across the business needs a proper data and governance foundation first.
Decision 5 of 6 · 90-day agenda
Re-evaluate one differentiated workflow for build versus buy
Pick one strategically differentiating workflow that a SaaS product constrains, and compare building and buying on full cost, including testing, maintenance and ownership. Decide which parts your team builds and which foundation layers, such as data and governance, come from a specialist partner or platform. Apply the same review to software renewals above a set value.
Suggested owner: CIO
Go deeperWhat's new since Q3, and the research behind this chapter
The research on developer productivity is mixed, and the differences are informative. On new, well-defined tasks, developers were 21% to 55% faster. In METR's 2025 trial, experienced developers working in their own mature codebases were 19% slower, even though they believed they were 20% faster. Google's DORA research links AI use to faster delivery but less stable systems.
The practical conclusion is to build where the logic is specific to your business and the scope is clear, and to fund the review, testing and ownership needed to keep it reliable.
The productivity gain depends on the task, and developers tend to overestimate it
Reported effect of AI coding tools in controlled trials. Positive values mean faster work or more output. Each study measures something slightly different.
Move from ‘Buy’ to ‘Own’. Then turn on ‘Show the build cost before AI’: the dashed outline shows why building wasn't viable a year ago. Notice that the testing segment didn't shrink.
06Creative becomes a production systemLayer: CreativeWatch area
Creative AI is becoming an orchestrated production system, and Higgsfield is one of the clearest examples.
Higgsfield puts the leading video and image models, camera control, consistent characters and ad templates into one workflow. A small team can now produce, and test, far more creative than a traditional production schedule allows.
Figure 06
Higgsfield at a glance
AWALI example · made with Claude, Blender and Higgsfield
Scene
The Teton Range from Driggs, Idaho, in winter, at 5,000 ft
Terrain
Built from USGS elevation data, so the mountains match the real topography
Tools
Claude and Blender built the 3D scene; Higgsfield generated the finished video
Build
A quick build that combined several models in one workflow
Write a short shot list and set the look with one key frame from an image model.
Use saved camera presets instead of describing the move in words.
Send wide shots and B-roll to a lower-cost model; keep the best model for the one or two hero shots.
Train a Soul ID if the same presenter or character appears across the campaign.
Generate versions by format and message, then cut, grade and caption in your usual editor.
Test the versions and track the cost of each ad that beats the current one.
Clips run up to 15 seconds per generation, final assembly still happens in an editor, and the capability descriptions are Higgsfield's own.
What it does well
Many models, one placeVeo 3.1, Kling 3.0, Seedance 2.0, WAN 2.6 and others sit alongside Higgsfield's own models, so each shot can use the model that suits it.
Camera controlDolly, orbit, tracking and crane moves are set when the clip is generated, rather than hoped for in a text prompt.
Consistent peopleSoul ID is trained once from 20 or more photos and keeps the same face across every clip and model.
From product page to adMarketing Studio reads a product image or link and fills templates for product shots, Meta and Google ads, short UGC-style video and motion graphics.
Editing in placeInpainting, relighting, upscaling and layers change a scene without regenerating it from scratch.
When producing a campaign's worth of versions takes days rather than months, the useful measures change: the cost of each ad that beats the current one, and how quickly results feed the next brief.
Watch area · baseline and trigger, no major commitment yet
Baseline creative testing speed
Measure how many ad versions you test each month and how many days it takes to act on results. Then run one bounded test in a single channel with a production platform such as Higgsfield. No enterprise commitment is needed yet.
Suggested owner: CMO
Go deeperWhat's new since Q3, and the research behind this chapter
Generated video lists at $0.10 to $0.40 a second, so 30 seconds costs $3 to $12 in model fees before retries, editing and review. Production cost is no longer the main constraint.
Audience reaction still has to be tested rather than assumed. In IAB's survey, 82% of advertising executives believed young consumers feel positive about AI ads; 45% of those consumers do.
Play the video first. Then walk through the five strengths and the six steps. The point is throughput and learning speed, not cheaper assets.
07AI is changing faster than annual plans can respondLayer: OperationsWatch area
AWALI tracked 44 significant AI developments in the first nine months of 2026, including 25 major model releases. That is more change than a once-a-year plan is designed to absorb.
Under an annual plan, each of those developments waited an average of 204 days before anyone could decide how to respond. With a monthly review, the average wait drops to 13 days.
Figure 07
How long each AI development waits for a decision, by planning cycle
Significant AI developments in 2026, by date, compared with three planning cycles. Choose a cycle to see how long each change waits before anyone can act on it.
Material AI changes in 2026 to 2 October. Kling 3.0 is dated to February 2026 by Kuaishou; the day shown is approximate. Anthropic's >80% figure is as of May 2026, published June 2026. Model releases are the flagship releases of OpenAI, Anthropic, Google, xAI and Meta plus the leading open-weight releases; dates follow public release trackers, using the first public availability where trackers differ.
Annual planning still works for setting strategy, and quarterly cycles still work for allocating capital. AI operations increasingly need a decision point every month.
Watch area · baseline and trigger, no major commitment yet
Instrument your AI planning cadence
Log material AI developments and how long each waits for a decision. Set a trigger: if more than one a month affects your priorities, add a 30-day AI operating review alongside annual strategy and quarterly capital allocation.
Suggested owner: COO
Go deeperWhat's new since Q3, and the research behind this chapter
The length of software task AI can complete went from about an hour in early 2025 to about 12 hours by February 2026. Since 2023 it has doubled roughly every 4.3 months. METR notes that its tests can barely measure the leading models, and that the task length at 80% reliability is much shorter.
AI labs report that AI now writes most of their own code: more than 80% of merged code at Anthropic and 75% of new code at Google. Both figures are self-reported and measured differently.
Something that was out of reach when the budget was set can become routine before the end of the year. A monthly review needs a clear trigger for revisiting work when that happens.
The length of task AI can complete doubles every few months
METR's measure of the length of software task AI can complete with a 50% success rate, by model release date, on a log scale with 95% confidence intervals.
Start with ‘Annual’ and read the average wait: 204 days. Switch to the 30-day review: 13 days. That difference is the point of the chapter.
08AI agents need limits built inLayer: GovernanceDecision
Enterprise exposure is determined less by an agent's raw capability than by the authority, access and controls surrounding it.
Agents can now send email, move money and change records on their own. A written policy is unlikely to stop one in the middle of a task; controls built into the systems it uses can.
Figure 08
What one AI agent can do on its own, with and without controls
One operations agent with access to seven systems. The rings show how serious each action is, from reading data to actions that can't be undone. This is an illustration; nothing is executed.
Illustrative scenario: one operations agent with access to Email, CRM, Finance, Documents, Calendar, External tools, Customer comms. Nothing is executed.
Control
When on
When removed
Identity
Every action is signed by the agent's own identity, with a named owner.
Actions run on a shared login, so nobody can tell which agent did what.
Permission scope
The agent can reach only the systems and actions this job requires.
The agent has broad access, so deleting records, wiring money and calling any API all become possible.
Approval thresholds
External emails, file shares and customer messages need a person's approval.
Actions that can't be undone go ahead without anyone approving them.
Spending limits
Payments are capped at $1,000, and the agent can't make purchases or issue refunds.
There is no limit on how much money the agent can move.
Rate limits
The agent can take at most 20 actions a minute, so mistakes stay small.
The agent can take thousands of actions a minute, so one bad instruction can repeat many times before anyone notices.
Audit logging
Every action is recorded with who did it, what went in and what happened.
Nothing is recorded, so incidents can't be reconstructed.
Human escalation
Unclear or out-of-policy cases go to a named person.
The agent decides for itself how to handle exceptions.
Kill switch
An operator can stop the agent and revoke its access in the middle of a task.
Once a task is running, it can't be stopped.
Evidence
34%
of organizations apply the same security controls to AI agents as to employees
Only 34% of organizations apply the same controls to AI agents as to employees. The gap between a typical pilot and a governed agent is four controls, and each is far easier to put in place before launch than after an incident.
Decision 6 of 6 · 90-day agenda
Set real-time controls for AI agents
Make an inventory of every AI agent and automated account. Give each one a named owner, its own identity, spending and activity limits, approval thresholds, logging, an escalation path and a tested way to shut it down.
Suggested owner: CISO
Go deeperWhat's new since Q3, and the research behind this chapter
The necessary controls are well understood: a separate identity for each agent, only the access it needs, approval thresholds, logging, ongoing testing and the ability to revoke access immediately. Recent incidents show why they need to be enforced while systems run rather than written into policy. In 2025 a coding agent deleted a live production database during a code freeze, and AI carried out 80% to 90% of the work in a state-linked hacking campaign. In 2026, evaluation agents reached a third party's production systems.
Practice is behind. Only 34% of organizations apply the same controls to AI agents as to employees. OWASP's first list of top risks for agentic applications includes identity abuse, misuse of tools and agents acting outside their instructions, and NIST has started standards work on agent identity and authorization.
Start with ‘Fully governed’. Switch to ‘Typical pilot’: identity, limited access and a log are on, yet the agent can message customers and issue refunds unsupervised and can't be stopped. Then ‘No runtime controls’: deleting records and wiring money open up too.
Next steps
Turn these six decisions into your 90-day AI operating agenda.
Each of the six decisions has a suggested owner. In an AWALI workshop, we turn them into named owners, dates and a first operating review within 30 days.
Finish on the agenda. Assign owners before the meeting ends, and send the brief the same day.
Sources and method · 63 sources
Where every figure comes from
Figures are reported as published by each source as of the date shown. This edition was compiled on 2 October 2026, and model and pricing data reflect that date.
Where a shorter summary would overstate a finding, we use the source's original, narrower wording. AWALI has not independently reproduced these studies.
Self-reported figures, press reports, forecasts and AWALI calculations are labeled wherever they appear.
Every chart is drawn from the values shown in its table view. Figures we could not confirm were left out rather than estimated.
The interactive scenarios (the routing policy, the decision simulator, the voice trace, the build-versus-buy comparison, the ad-testing example and the agent controls map) illustrate how these systems behave. They are not measurements.