For the past two or three years, the face of generative AI in the enterprise has been the "copilot" — a prompt-and-response assistant you ask and it answers. In 2026, the conversation moves a step further. The talk is of AI agents: systems that take a goal, plan toward it on their own, call external tools, update systems, and escalate to a human when they get stuck. Anthropic's "The 2026 State of AI Agents Report," released in December 2025, surveyed more than 500 technical leaders and reports that 80% of organizations already see measurable economic returns from their AI agents [Source: Anthropic, 2025].
Why is the whole world watching this shift now? Several research firms point to 2026 as the inflection point where the center of gravity moves from copilots to agents. Yet the same data also shows a sober reality. Most companies' agents still handle single, low-risk tasks, and the biggest barrier is not the model's intelligence but the work of integrating it with existing systems. This article sets out what has actually changed, how to think about autonomy in tiers rather than absolutes, and how to tell announced expectations apart from verified results.
A word on method before the numbers begin. Three kinds of statements get mixed together in this field. There is what has been measured — survey results from a named sample, where organizations report on themselves. There is what has been projected — research-firm forecasts about years that have not happened yet. And there is what has been assessed — an analyst's judgment about the field's maturity, informed by data but not a count. One caveat covers the measured tier throughout: it is self-report. No independently audited dollar figure for agent returns appears here, because the sources contain none.
Table of Contents
- Copilots vs. agents — what actually differs
- Five tiers of autonomy — where 2026 sits
- Are the returns real — and what was measured
- The real barrier isn't the model
- Deployment is fast, coordination is slow
- What to watch
Copilots vs. agents — what actually differs
Two different jobs: responding and executing
Start with the terms. A copilot is an assistant you direct at every step and whose output you review. It drafts a document or suggests a snippet of code, but the next move is always yours. An AI agent (agentic AI), by contrast, takes a single goal, breaks it into sub-tasks and plans them, calls the tools it needs (search, databases, internal APIs), and executes several steps in sequence. Where it is unsure or exceeds its permissions, it hands the decision back to a human. In short, the difference is between responding and executing.
Two pieces of that definition carry most of the weight. A "tool call" is the agent reaching outside its own text output — running a search, querying a database, writing to an internal system — so that its reasoning ends in a change to the world rather than a paragraph on a screen. "Escalation" is the stop condition: the point where the agent judges a question beyond its confidence or permissions and returns it to a person. With a copilot the human sits inside the loop; with an agent, at the edge of it, handling exceptions.
Why the shift is happening now
This shift owes less to models suddenly getting smarter than to the maturing of the plumbing that wires models to tools and data across multiple steps. Anthropic's report makes the point plainly: "agent adoption is no longer limited by model capability." In fact, 57% of surveyed organizations already run multi-step agent workflows, and 16% have gone as far as cross-functional processes spanning multiple teams [Source: Anthropic, 2025].
Those two numbers say more together than apart. A majority have crossed from single answers into sequences of steps, but only a small minority have carried agents across team boundaries — and that gap is not mainly technical. Running several steps in a row is an engineering achievement; running a process that spans teams means touching several systems, several sets of permissions, and several groups of people who each own part of the workflow. The first a capable team can build; the second an organization has to agree to.
What it looks like in practice
Concretely, it looks like this. An agent handed a support ticket triages the problem, queries internal knowledge bases and logs, changes the settings it needs to, and passes only what it cannot resolve to a human. Gartner expects task-specific agents to handle work such as automating development, managing incidents, and resolving support cases "without human involvement" [Source: Gartner, 2025]. Where a copilot hands over a draft and stops, an agent tries to carry that draft into action and check the result. That, though, is the destination, not today's norm.
Note which layer that Gartner statement belongs to. "Without human involvement" describes what a research firm expects task-specific agents to do, not a measured tally of tickets closed unaided today. The example is still useful, because it shows what execution adds to response: the agent has to read state, change state, and know when to stop. A copilot only ever has to produce text about all three.
Autonomy is a spectrum, not a switch
One thing is worth stating clearly, though. "Agent" is not an on/off switch but a spectrum. Under the same label, one agent may automate a single fixed task while another reasons across several systems. As we will see, most of today's deployments cluster at the lower end of that spectrum.
Organizations appear to read it as a ladder too. In the same Anthropic survey, 81% said they plan to expand to more complex use cases during 2026 — 39% toward multi-step workflows and 29% toward cross-functional processes spanning teams [Source: Anthropic, 2025]. Those are stated intentions — the projected tier, not the measured one. But their shape is informative: companies describe their next move as one rung up the same ladder, not a jump to full autonomy.
Five tiers of autonomy — where 2026 sits
Gartner's five stages
To talk about autonomy without hype, you need tiers. In an August 2025 forecast, the research firm Gartner projected that AI inside enterprise apps would evolve through the following five stages [Source: Gartner, 2025]. Note up front that this is a forecast, not a measurement.
- Late 2025 — nearly every enterprise app embeds an AI assistant (a copilot).
- 2026 — 40% of apps carry task-specific agents dedicated to particular jobs.
- 2027 — about one-third of implementations use collaborative agents for complex tasks.
- 2028 — multi-agent ecosystems dynamically collaborate across multiple apps.
- 2029 — at least half of knowledge workers build, govern, and deploy agents on demand.
Read down the list and you can see what each rung adds. The first embeds an assistant beside the work; the second hands over a bounded job end to end; the third lets several agents cooperate on one complex task; the fourth removes the application boundary; the fifth moves authorship to the worker. Each step adds either scope or authorship — and the sequence, not the calendar, is the durable part of a forecast like this.
Why the 2026 marker matters
Gartner put the 2026 milestone at 40% of enterprise apps carrying task-specific agents — a jump from less than 5% in 2025 [Source: Gartner, 2025]. What matters is the placement. The 2026 marker is the task-specific agent, not the "multi-agent" system that headlines favor. On Gartner's own roadmap, multi-agent ecosystems mature closer to 2028 (stage four). Calling 2026 "the year of the multi-agent" thus gets the direction right but the timing early.
The long-range number, and how to hold it
Gartner also put a number on the far end of the ladder. In a best-case scenario, it suggested agentic AI could account for roughly 30% of enterprise software revenue by 2035 — more than US$450 billion — up from about 2% in 2025 [Source: Gartner, 2025]. Two labels belong on that figure and should not come off: it is the optimistic scenario, and its horizon is a decade away. That makes it the least verifiable number here — a picture of the scale one firm thinks possible, not a plan. The 40% figure, by contrast, describes the end of 2026 and can be checked within a year.
Are the returns real — and what was measured
What the 80% actually says
So are the returns real? The most striking figure in Anthropic's report is 80%: that share of organizations reports already seeing measurable economic returns from their AI agents, and 88% expect those returns to hold or grow through 2026 [Source: Anthropic, 2025]. This number should be read for what it is, though — a self-reported survey result, not an externally audited accounting figure. "Reported a return" and "independently verified a return" are statements at different tiers.
It is worth being concrete about what the figure leaves open. A self-reported "measurable economic return" tells you an organization believes it has a number it can stand behind — not how large that number is, how it was calculated, or whether anyone outside has checked it. The survey measures how many organizations answered yes, which is a real measurement and a different thing from a measurement of value created. That absence is part of the picture: the strongest broad evidence available today is a large, consistent self-report.
Coding is the frontier
The clearest frontier for those returns is software development. In the same report, roughly 90% of organizations said they use AI to assist development, and 86% said they deploy agents on production code [Source: Anthropic, 2025]. Code is fertile ground for agents because results can be checked immediately with tests.
The gap between those two figures is small, and that is the interesting part. Using AI to assist a developer and letting an agent touch code that ships are different levels of trust, yet the second sits only a few points below the first [Source: Anthropic, 2025]. In most domains the drop from "we use it" to "we let it act on the real thing" is far steeper. The likely reason is verification: code arrives with tests, review, and version history — a ready-made way to catch a mistake before it ships, and to undo it if it does.
How to read the return, honestly
In sum, the direction of the returns is clear: many companies feel a real benefit, and they report it first in specific domains, coding above all. But how large that benefit is deserves the caution owed to any self-report.
One further distinction sits inside the same pair of numbers. The 80% reports something that has already happened; the 88% who expect returns to hold or grow through 2026 describe something that has not [Source: Anthropic, 2025]. Expectations are the easiest thing for a survey to capture and the hardest to hold anyone to. Keeping the two apart prevents a common error: treating a strong outlook as a strong result.
The real barrier isn't the model
Four barriers, and what they share
If adopting agents is hard, why? The answer, interestingly, is not "because the model isn't smart enough." In Anthropic's survey, the top barrier organizations named was integration with existing systems (46%), followed by data access and quality (42%), security and compliance (40%), and change management (39%) [Source: Anthropic, 2025]. The crucial point is that every leading barrier lives outside the model — in data, in systems, in the organization.
Notice how close together those four figures sit: 46, 42, 40 and 39 [Source: Anthropic, 2025]. There is no single wall everyone runs into, which would be the easier problem — there is a cluster of obstacles of roughly equal weight. Connecting to systems, trusting the data inside them, satisfying the rules that govern them, and persuading the people who use them are four faces of the same job. None is fixed by a better model.
The bottleneck moved
This dovetails with the line quoted earlier: model intelligence is no longer the primary bottleneck to getting agents into production [Source: Anthropic, 2025]. The bottleneck has moved from model performance to operational infrastructure. For an agent to actually do work, it has to connect safely to internal systems, read trustworthy data, and operate inside permissions and audit trails. The distance between a slick demo and a workflow that runs every day usually opens up right here.
"Operational infrastructure" sounds abstract, but each piece is concrete. An agent needs a credentialed route into internal systems rather than a person copying data by hand. It needs permissions narrow enough to be safe and broad enough to be useful. It needs data current enough to act on, since an agent reading a stale record does not hesitate the way a person might. And it needs a record of what it did, because an action nobody can review is one nobody can approve. A demo bypasses nearly all of that — prepared data, a scope chosen to work, a person present to intervene — while production must satisfy every requirement with nobody watching. The capability on display is genuine; what the demo omits is the apparatus around it, and that is where the months go.
Deployment is fast, coordination is slow
Twelve agents, half of them alone
Speed of deployment and maturity of coordination are two different things. According to Salesforce's 2026 Connectivity Benchmark Report, which surveyed 1,050 IT leaders worldwide, the average company already runs about 12 AI agents, a number projected to grow roughly 67% within two years. Yet about half of those agents operate in isolation — disconnected from one another and, in many cases, from any central oversight [Source: Salesforce, 2026].
The survey behind those figures is worth naming, since methodology governs how much weight a number carries. It is Salesforce's eleventh annual Connectivity Benchmark, run with the research firm Vanson Bourne among 1,050 global IT leaders and fielded in October and November 2025 [Source: Salesforce, 2026]. On the growth rate it reports, the typical company would run close to 20 agents by around 2027 — a projection carried forward from a measured base [Source: Salesforce, 2026]. The striking part is not the count, though, but that a company can reach a dozen agents one at a time and still have no picture of them as a system.
The governance gap
This picture lines up precisely with Anthropic's numbers. Many run multi-step workflows (57%), but only 16% have reached cross-functional processes spanning teams [Source: Anthropic, 2025]. Deployment is fast; coordination lags. Salesforce's finding that only 54% of organizations have a formal governance framework covering their agent deployments points the same way. Deployment is outrunning control, creating the management blind spot some call "shadow AI" [Source: Salesforce, 2026].
Two of Salesforce's figures make the gap plain side by side. Some 89% of organizations say they are deploying agents across most or all of their teams, while 54% report a formal governance framework covering those deployments [Source: Salesforce, 2026]. The two questions do not track the same respondents item by item, so the distance between them is not a precise count of ungoverned companies. As a description of the field, though, it is hard to miss: broad deployment is close to universal, formal governance is not. That is what "shadow AI" names — agents nobody has a complete list of.
The optimistic and the cautious readings
Here a balance of perspectives is needed. Where Gartner paints optimism in app-penetration rates and market size, Forrester is more cautious. In its 2026 assessment, Forrester noted that roughly three-quarters of enterprise leaders say they are adopting agentic AI, but only a minority have it running in meaningful production, and true scaled multi-agent systems are rarer still [Source: Forrester, 2026]. Forrester framed it as "the technology is a runaway train; the enterprise is the heavy load it has to pull," arguing that the real constraint is not the model but risk management and governance. Indeed, 49% of security decision-makers named agentic AI as a concern [Source: Forrester, 2026]. Adoption and maturity are not the same word.
These readings contradict each other less than they appear to, because they answer different questions. Gartner is counting software: how many applications ship with an agent inside. Forrester is assessing organizations: how many have agents doing meaningful work at scale. Both can hold at once — an application can arrive with a task-specific agent built in while the company using it is nowhere near a coordinated multi-agent system.
What to watch
To sum up, 2026 is the year the center of gravity genuinely began to move from copilots to agents. The verified facts are these: many organizations report measurable returns from agents (80%), more than half now run multi-step workflows (57%), and coding is the frontier (90% assist development). At the same time, the open questions are just as clear. The biggest barrier is not the model but integration, data, and governance (integration at 46%); cross-functional automation spanning teams is still a minority (16%); and half of deployed agents run in isolation from one another.
If you follow this field, three questions will do most of the sorting when the next headline number arrives. Who measured it — the organization reporting on itself, or an outside party? Which tier is it — measured, projected, or assessed? And does it describe adoption or scale: that agents were deployed, or that they are doing meaningful work across teams? Most of the disagreement between optimistic and cautious accounts turns out to be a disagreement about which question was answered.
Three things are worth watching from here. First, whether today's single-task agents mature into coordinated multi-agent systems — on Gartner's roadmap, 2027–2028 is the turning point. Second, whether the real bottlenecks of integration, data, and governance get solved, or whether deployment keeps outrunning control. Third, whether today's self-reported gains hold up over time as durable, independently verified results. The copilot assisted the human; the agent has begun to execute the work. How dependable and how well-governed that execution proves to be will be the next chapter of the story.