Back to Blog
Last updated Sep 11, 2026.

76% Think They Can Stop an AI Agent in 15 Minutes, 33% Have a Kill Switch

minutes read
Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Share:
76% Think They Can Stop an AI Agent in 15 Minutes, 33% Have a Kill Switch
Harness asked 700 technology professionals: 76% believe they can disable an agent in 15 minutes, 33% have a kill switch. A second survey agrees.
AI agentsgovernanceenterprise AIsurvey

TL;DR

Harness published The State of Agent DLC 2026 on 10 September. Sapio Research, 700 technology professionals, US, UK, France, Germany and India, fielded July 2026.

Every confidence question outruns its matching control question. The kill-switch pair is the worst: 76% believe they can disable an agent in 15 minutes, 33% have the mechanism.

7 in 8 organisations had at least one agent-related issue. Of the 75% who say their agents are secure, 88% have already had a security incident.

A separate survey by Gravitee, different vendor and different sample, finds 91.8% claim visibility against 52% actual monitoring coverage.

Both land near a 40-point gap. That is the number to take away, because two independent instruments agreeing is worth more than either one.

What did Harness actually measure?

Confidence and capability, as two separate questions, which is why the result is useful.

The survey was run by Sapio Research in July 2026 across 700 technology professionals, screened to organisations with at least 1,000 employees, 100 developers and $100 million in annual revenue, and with AI agents already in production, pilot or proof of concept. Five countries: United States, United Kingdom, France, Germany, India.

Asking "are you confident" and "do you have the mechanism" separately produces this:

What they believe. What they have

77% confident they have a complete agent inventory. 44% run active discovery

74% confident testing catches failures. 19% have automatic blocks

76% believe they can disable an agent in 15 minutes. 33% have kill switches

75% say their agents are secure. 88% of that group have had a security incident

74% say they have complete spend visibility. 60% still overran budgets

Keith Mann, Field CTO at Harness, names the mechanism: "Controls that worked in testing can still miss something because agents don't behave the same way every time."

The spend row deserves its own note, because it is the one with a hard edge already in the market. Three quarters believe they can see what agents cost, and 60% overran the budget anyway. When we read the Copilot Credit rate card behind Business Central's managed AI resources, the tiers spanned 100x for the same ten answers and enforcement disabled custom agents at 125% of prepaid capacity. Spend visibility that arrives on the invoice is not visibility, and the penalty for discovering it late is an agent that stops answering.

Which gap is the dangerous one?

The kill switch, and it is not close.

An inventory gap means you do not know what you have. A testing gap means problems reach production. Both are bad and both are recoverable. A kill-switch gap means that at the moment you decide something must stop, you find out whether stopping it was ever built.

believe they can disable an agent in 15 minutes 76%

actually have a kill switch 33%

believe it without the mechanism 43 points

Forty-three points of an enterprise population are holding an incident-response assumption that has no implementation behind it. That assumption is not tested on an ordinary day. It is tested on the worst day, under time pressure, by whoever happens to be on shift.

Yesterday we covered an attack that went from cloud compromise to a running mass campaign in under six hours. Set that against a fifteen-minute stopping assumption that two thirds of organisations cannot execute, and the arithmetic of a bad afternoon writes itself.

The security number that inverts

There is one pair in this report that does not merely show a gap. It shows the confidence pointing the wrong way.

75% say their agents are secure. Of that group, 88% have experienced a security incident.

Read it twice. Among the organisations most confident in agent security, nearly nine in ten have already had the thing they are confident will not happen. That is not a gap between belief and capability. It is belief that has become detached from a company's own incident history.

The report adds the aggregate: 7 in 8 organisations had at least one agent-related issue. So the incidents are not rare, and the confidence is not rare either. They are simply not talking to each other.

Does a second survey agree?

Yes, and that is the part worth more than either survey alone.

Gravitee's State of AI Agent Security Report surveyed 750 senior technology leaders, 500 in the US and 250 in the UK, through Opinion Matters in April 2026, against a December 2025 baseline. Different vendor, different sample, different countries, different question set.

8% said they had visibility into the agents running inside their organisation. Actual monitoring coverage sat near 52%.

Harness average of the 3 clean confidence-vs-control pairs

(77-44) + (74-19) + (76-33) = 33 + 55 + 43 = 131 / 3 = 43.7 points

Gravitee 91.8 - 52 = 39.8 points

Two independent instruments, and the gap comes out around 40 points in both. When two surveys built by competitors, in different quarters, with different samples, converge on the same magnitude, the number is probably describing something real rather than either vendor's marketing.

The number that moves the wrong way

Gravitee ran the same questions twice, and the trend is the finding.

Between December 2025 and April 2026, stated confidence in agent visibility rose from 82.6% to 91.8%, a gain of 9.2 points in four months. Monitoring coverage moved from 46.96% to roughly 52%, a gain of 5 points.

So the gap did not close. It widened:

December 82.6 - 46.96 = 35.6 points

April 91.8 - 52.00 = 39.8 points

change +4.2 points

And the fleet grew underneath both numbers. Gravitee reports the modal enterprise deployment jumping from a band of 26 to 50 agents to a band of 76 to 100 in a single quarter. Take the midpoints and run what that does to the absolute count of unwatched agents:

December 38 agents x (100 - 46.96)% unmonitored = 20.2 unmonitored

April 88 agents x (100 - 52.00)% unmonitored = 42.2 unmonitored

Monitoring coverage improved by 5 points and the number of unmonitored agents doubled. That is the trap in measuring this as a percentage: the denominator is growing faster than the coverage. A team can genuinely improve every quarter and still be losing ground, and the dashboard will show green while it happens.

7% plan to deploy significantly more agents in the next twelve months. The denominator is not done growing.

Why do the old controls not fit?

Because the thing being controlled is not deterministic, and almost every control in a release pipeline assumes it is.

Harness puts it plainly: the same agent can produce different outputs from one run to the next. A test that passes proves the agent behaved once, not that it behaves. Trevor Stuart, SVP and GM, describes the sequencing that produced this: "Teams moved fast to build agents, now circling back to ask how to govern what they've shipped."

The pipeline data shows exactly that shape:

53% run agent changes through standard pipelines, the ones built for deterministic code

37% run less than half their changes through any pipeline at all

42% route prompt edits through code pipelines, and only 34% have a dedicated system for AI behaviour configuration

58% report increased production incidents per 100 changes

That third bullet is the one engineers will recognise. A prompt is a behaviour change shipped as a text edit. It carries the risk of a code change and, in most organisations, the review process of a copy fix.

What should a team do about this?

Four things, in this order, because the order is the whole point.

Build the kill switch first, before the next agent ships. It is the cheapest control in the list and the only one whose absence is discovered during an incident. If you can only do one thing this quarter, it is this one, and the test is a drill rather than a design document.

Run discovery rather than asking. The inventory gap is 33 points wide because the inventory is assembled by asking teams what they run. Active discovery returns a different list, and the difference is the finding.

Treat a prompt change as a deploy. Same review, same staged rollout, same rollback path. Harness's recommendation is progressive rollout for agents, canary and blue/green, and it notes those remain far less common for agent changes than for code changes.

Count agents, not percentages. As the Gravitee arithmetic shows, coverage can improve while exposure grows. Report the absolute number of agents with no monitoring attached, every month. It is a harder number to feel good about, which is the point.

FAQ

Who ran these surveys?

Harness commissioned Sapio Research, July 2026, 700 technology professionals across five countries. Gravitee commissioned Opinion Matters, April 2026, 750 senior technology leaders in the US and UK.

Are these vendor surveys?

Yes. Both companies sell into this problem. That is a reason to read the framing critically and a reason the convergence matters: two vendors with different products landed on the same magnitude.

What is an agent kill switch?

A mechanism to stop a running agent immediately, independent of the agent itself, without redeploying the application it lives in.

Does 88% mean agent security is hopeless?

No. It means self-assessed security is not a measurement. The organisations reporting incidents are also, in many cases, the ones with enough instrumentation to see them.

How many agents does a typical enterprise run?

Gravitee's modal band moved from 26 to 50 agents in December 2025 to 76 to 100 by April 2026.

The last mile

The striking thing in both reports is that nobody is being careless. These are organisations with real revenue, real engineering teams and real pipelines. They shipped agents quickly, which was the correct commercial call, and the governance is arriving afterwards, which is the normal order for every technology that ever mattered.

What is different this time is the shape of the lag. A confidence number that rises faster than a control number is not a slow rollout. It is a widening belief that the problem is handled, growing on top of a fleet that is doubling. The dashboards will keep reading green because percentage coverage is improving, and the absolute number of agents nobody is watching will keep going up.

Which makes the useful work unglamorous and specific: knowing exactly which agents are running, what they can reach, what they cost, and how to stop one inside a system a business already depends on. That instrumentation is the layer Cognilium builds in, and on this evidence it is roughly forty points ahead of where most estates think they already are.

Share this article

Share:

Weekly AI engineering brief

One email a week. New model releases, agent patterns, and lessons from production systems we ship.

No spam, no client data sales. Unsubscribe any time.

Ali Ahmed

Ali Ahmed

AI Solutions Engineer, Cognilium AI

Ali Ahmed is an AI Solutions Engineer at Cognilium AI.