One disclosure before the argument. This publication is produced with support from Docket, which sells software in this category. Nothing below names or scores a product, and every question in it can be put to any vendor in the market.
Whether a system is called an agent or a chatbot tells you almost nothing about how much supervision it will need. Three things do: what it can reach, how far its output steers a decision, and whether it can execute without a person in the loop. Those vary independently, and none of them is fixed by the noun on the contract.
The label does not price the supervision
We have made the definitional case separately and it holds. What a definition cannot settle is the operating question, because two systems can both qualify as agents and still owe you completely different amounts of oversight.
Gartner puts the failure in all-or-nothing governance
In a press release dated 26 May 2026, Gartner argues that applying uniform governance to all AI agents, regardless of their autonomy level and scope, leads to failure. It predicts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps found only after a production incident. Both are guidance and forecast rather than measured results.
Its stated mechanism is specific: "Failures are most likely to occur when organizations fail to distinguish between an agent's ability to act and the scope of access it is granted."
Shiva Varma, a Senior Director Analyst at Gartner, locates the cause in an all-or-nothing habit. "Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure." Over-restriction, Varma says, "slows delivery and drives shadow development"; under-restriction "increases operational, security and compliance risk."
Gartner's recommendation is proportional governance, classifying agents by autonomy level and treating each level as its own trust boundary. Its published ladder has four rungs.
Observe has read-only access to defined data sources, with output visible only to the person who asked. Advise drafts, recommends or proposes while a human reviews everything and executes by hand; Gartner is explicit that this rung also retains read-only access, with no write access to any system. Act with approval can write data, send communications or change configuration, but only after explicit human approval for every action. Act autonomously executes inside defined guardrails, while humans review exceptions, audit logs and aggregate outcomes rather than individual decisions.
Note that access scope and execution authority are separate dials. Widening what a system may read does not promote it up that ladder, and a system can move up the ladder without gaining any new data.
Read-only still carries a risk
The second rung is the one that complicates any account resting on execution alone, including the one this piece opened with.
Advisory agents, Gartner says, "can anchor judgment, creating downstream risk when inaccurate outputs are trusted due to automation bias." The governance it prescribes there is heavier than for pure retrieval: accuracy and hallucination testing, domain-specific quality evaluations, and user training on appropriate reliance levels.
So a system with no write access anywhere can still carry real risk, through a person who stops re-deciding. Decision influence is its own dial, and it is the one least visible in a demo.
Execution authority changes what the reviewer needs to see
Gartner's warning about the third rung is the most operationally useful passage in the release. "At this level, human review is effective only if it remains a meaningful control," Varma says. Without security testing, clear approval workflows with audit trails and agent-specific incident response, "approvals can degrade under time pressure or approval fatigue, creating a false sense of safety while expanding the attack surface."
What follows is our inference, worked through a hypothetical rather than an observed case, and the conditions matter to the argument.
Take a lifecycle email system that already drafts, and give it the ability to send behind a per-send approval. Assume the same reviewer, roughly twenty approvals a day, an approval screen that renders the full email with a segment name and a recipient count, and a log that records approver, timestamp and outcome.
Under those conditions the log cannot separate a careful approval from a superficial one, because the two produce identical rows. Nothing in the screen forces the difference either: the reviewer sees a rendered email, and a rendered email that resembles last week's approves easily.
The information that would change the decision is not on the screen. Which contacts entered the segment since the previous send, what changed in the copy, whether a suppression rule was edited. Those are the facts that make a send wrong, and they are the ones a full preview buries.
The case for a compact approval carrying a risk summary, and for treating reviewer time as a real cost, is one we have already made. What is new here is narrower. Adding execution authority to a workflow does not just add a click. It changes what the reviewer has to know, and the approval surface built for the drafting rung will not carry it.
The tests to demand of a vendor, and the question of who answers when it goes wrong, are set out in the agent ownership problem. Use those rather than a new checklist.
Where this leaves an approval-dependent workflow
For workflows that depend on a human approving individual actions, and only those, one requirement stays unproven in most deployments: that the approval screen surfaces the deltas which could make the action wrong, and that the reviewer has the time to read them at the volume the system actually generates.
That is a checkable requirement, and it is checkable before the pilot rather than after the incident Gartner's forecast is about. Where a vendor cannot describe what its approval screen shows beyond the artifact itself, the control being sold has not yet been specified.
