Skip to content
Novyant
All insights
  • AI Automation
  • Governance

Enterprise AI agents doubled in four months. Governance did not.

Deployments roughly doubled between December and April. Monitoring coverage moved about five points, and 85% of organizations still cannot name a person accountable for what an agent does. The gap is the whole story.

By Roberto Sanson7 min read
A records archive receding into darkness with a single shelf lit from within

Something changed in enterprise AI between December 2025 and April 2026, and it was not the models.

According to Gravitee's State of AI Agent Security 2026, a survey of 750 executives and technical leaders taken twice four months apart, the typical enterprise went from running between 26 and 50 agents to running between 76 and 100. That is a doubling inside a single quarter, in a category that barely existed two years ago.

Over the same four months, the share of those agents anyone was monitoring went from 47% to roughly 52%.

Read those two numbers together and the picture is not "AI adoption is accelerating." It is that the absolute number of unwatched agents in production grew faster than at any point in the history of enterprise software. Confidence, meanwhile, went the other way: stated confidence in agent visibility rose from 82.6% to 91.8% while coverage stayed roughly flat.

This is vendor-sponsored research, and it is worth reading with that in mind. A company selling agent governance has an interest in finding a governance gap. But the shape matches what we see walking into client environments, and the direction of the confidence number is the part no vendor would have made up. It is too specific and too unflattering to everyone.

What the survey found

MeasureFinding
Typical deployment size26 to 50 agents → 76 to 100 agents in four months
Plan to deploy significantly more within 12 months81.7%
Mean monitoring coverage46.96% → ~52%
Monitor more than 80% of their agents9.5%
Fully secure and govern agents before deployment19.7%
Experienced or suspected an agent incident in 12 months54%
Have no formal accountability structure for agent behaviour85%
Can name an individual responsible for agent actions7.2%

The last row is the one to sit with. In 92.8% of these organizations, if an agent did something expensive on a Tuesday, there is no name attached to it.

Why this happened now, and why it is not carelessness

Agents got dramatically easier to deploy, and nothing else changed at the same speed.

The specific unlock was interoperability. Anthropic's Model Context Protocol became the de facto way to connect a model to real systems, and it is now supported by every major model provider along with thousands of enterprise tools. Anthropic handed MCP to the Linux Foundation in December 2025 to cement it as neutral infrastructure. TechCrunch's own year-ahead piece called 2026 the year AI moves from hype to pragmatism, with MCP as the reason agents would finally reach day-to-day practice.

It worked. Connecting an agent to a calendar, a database, a ticketing system or a payments API stopped being a six-week integration project and became an afternoon.

What did not get easier by the same factor: deciding which agent is allowed to touch which system, logging what it did, noticing when it stops working, and answering for the result. Those remained ordinary organizational work: slow, political, and nobody's favourite quarter.

So the survey's finding that 81% feel pressure to deploy quickly is not a finding about recklessness. It is a finding about asymmetry. One side of the work got ten times cheaper and the other side did not.

An agent is not a feature. It is an employee with API keys.

The mental model most organizations imported is wrong, and it is the root of the measurement problem.

A feature is deterministic. You test it, ship it, and it does the same thing on day 300 as on day one. Monitoring it means watching for errors and latency.

An agent holds credentials, makes decisions under uncertainty, takes actions with external consequences, and behaves differently as its inputs drift. That is not a feature. Functionally, it is a junior employee who works instantly, never sleeps, has been handed production access, and has no manager.

Once you accept that framing, the survey numbers stop being surprising and start being obvious. No organization would let 100 new hires operate for a quarter without a reporting line, an access review or a record of what they did. Nearly all of them did exactly that with agents, because agents arrived through a software procurement path rather than a hiring one.

Only 30% of respondents described themselves as very prepared to manage agents as authenticated actors. That is the single most honest number in the report, because "authenticated actor" is precisely what an agent is.

A wall-mounted key cabinet holding rows of identical keys, with one hook empty
Identical keys, one board, no register of who took what. This is what a shared service account looks like from the perspective of an audit.

The four questions we ask before an agent goes to production

We build these systems for clients in regulated industries, where "we are not sure what it did" is not an acceptable sentence. Four questions decide whether an agent is ready, and none of them are about the model.

Whose credentials is it using? If the answer is a shared service account, every action it takes is anonymous by construction, and no amount of logging downstream will recover the attribution. An agent needs its own identity for the same reason a person does.

What is the blast radius of one bad decision? Not the worst case anyone can imagine, but the realistic worst case within the permissions it holds. If that number is larger than you are comfortable saying out loud, the fix is narrower permissions, not a better prompt.

How would you know it broke? Silent degradation is the characteristic failure of these systems. They do not crash. They keep returning confident output while quality drifts, and the drift is invisible unless someone sampled the output deliberately. If the detection mechanism is "a user will complain," the agent is unmonitored regardless of what a dashboard says.

Who is the name? One person, on the org chart, accountable for this agent's behaviour. Not a committee and not a team. The 7.2% figure exists because this question feels bureaucratic until the first incident, at which point it is the only question.

What to do if you already have agents in production

You almost certainly do, and probably more than you think. That is what a modal count of 76 to 100 means. Shadow deployment is normal here.

Start with an inventory, not a policy. Policies written before you know what is running describe an imaginary estate. Find every agent, what credentials it holds, what systems it can reach, and who built it. In most organizations this exercise alone takes a week and produces at least one genuine surprise.

Then rank by blast radius rather than by volume. The agent that drafts internal summaries can stay unmonitored for another month. The one with write access to a financial system cannot. Most governance programmes fail because they try to cover everything at once and stall; the ones that work start with the five agents that could hurt you.

Then instrument before you expand. The survey's most repeatable finding is that 81.7% plan to deploy significantly more agents in the next year. Adding to an estate you cannot see is how a manageable problem becomes an unmanageable one, and it is considerably cheaper to build observability at 80 agents than at 300.

The honest summary

The agent capability is real, the productivity is real, and none of this is an argument for slowing down.

But there is a specific failure mode ahead, and it is legible in the data already: 54% have had an incident, 85% have no accountability structure, and 81.7% are about to add more agents. Those three facts do not resolve on their own.

The organizations that come out of this well will not be the ones that deployed the most agents. They will be the ones that can answer, on any given Tuesday, what their agents did yesterday and who is responsible for it.

That is not an AI problem. It is an operations problem wearing an AI costume, and it is solved with the same unglamorous work that has always solved operations problems: identity, permissions, logging, and a name against each thing that can act.

If you are deploying agents into an environment where being wrong has consequences, how we approach AI automation describes the review queues, audit trails and confidence thresholds we build first, and where automation pays and where it quietly does not covers how to decide what to automate at all.

Working on something like this?

We spend the first conversation understanding what you run on today. No pitch, and no obligation to build anything.

Book a call