AI Agents · Guardrails · Security

Excessive agency has three causes, not one: functionality, permissions, autonomy

Published · Updated

OWASP splits agent overreach into three distinct root causes. Treating them as a single problem means building the wrong guardrail for the actual failure.

When an agent does something it shouldn't, the instinct is to add a review step and move on. OWASP's Gen AI Security Project frames the problem more precisely: excessive agency has three separate root causes, and each one needs a different fix, not a single blanket control bolted on afterwards.

Three root causes, not one

The three causes are distinct: excessive functionality, excessive permissions, and excessive autonomy. An agent can have too many tools available to it, those tools can have broader access than the task requires, or the agent can be left to act without any checkpoint. A system can fail on any one of these axes while the other two are perfectly sound, which is why a single fix rarely covers the real exposure.

Mitigations map to the cause, not to the agent as a whole

Three controls address the three causes separately. Human-in-the-loop review requires a person to approve high-impact actions before they happen, which targets autonomy. Authorisation enforced in downstream systems, rather than left to the LLM's judgment, targets permissions. Rate-limiting the number of actions an agent can take in a given period targets the scale at which functionality can be misused. Applying only one of these while ignoring the other two leaves the remaining causes unaddressed.

Autonomy also needs a hard stop built into the loop

Autonomy control is not only a review gate at the end. Anthropic's guidance on building agents recommends including stopping conditions, such as a maximum number of iterations, directly inside the agent's loop so it cannot run indefinitely even without a human watching. This is a structural limit on autonomy, distinct from a post-hoc approval step, and it works alongside human-in-the-loop review rather than replacing it.

Tools are code execution, and their descriptions can lie

Functionality risk extends beyond how many tools an agent has, to what those tools actually are. The Model Context Protocol specification states that tools represent arbitrary code execution and must be treated with appropriate caution, and that descriptions of tool behaviour should be considered untrusted unless they come from a trusted server. A tool's self-reported description is not a safety guarantee, which is a reason to scope functionality tightly rather than trust the tool's own claims about what it does.

Compounding errors are a cost problem too

Anthropic's engineering notes on building effective agents point out that the autonomous nature of agents means higher costs and the potential for compounding errors, and recommend extensive testing in sandboxed environments alongside appropriate guardrails. Their multi-agent research system write-up adds a concrete figure: agents typically use about 4 times more tokens than chat interactions, and multi-agent systems use about 15 times more. Excessive autonomy is not only a safety exposure; it is a cost exposure that scales with every unchecked iteration.

Recovery architecture is the other half of autonomy

Letting an agent run unsupervised only works if failure is recoverable. Anthropic describes building systems that resume from where the agent was when errors occurred, combining the adaptability of agents with deterministic safeguards like retry logic and regular checkpoints. They also implemented patterns where agents summarise completed work phases and store essential information in external memory before moving to new tasks. Autonomy granted without this kind of recovery design turns a single failure into a full restart.

What bounded autonomy looks like in production

The Tatano Energy platform runs a daily SEO autoblog across four country domains in seven languages, publishing eight SEO articles a day with manual intervention required at zero. The Matthieu Pesesse media platform publishes articles autonomously every day using three model fallback levels, with weekly content operations time at approximately zero hours. Both run unsupervised, but on a narrowly scoped, repeatable task, not open-ended decision-making, which is consistent with limiting functionality and autonomy rather than removing the limits altogether.

Before adding another review step to an agent, identify which of the three causes is actually present: too many tools, too much access, or too little oversight on the loop itself. The fix for one is not a fix for the other two, and an agent that is safe on functionality and permissions but unbounded on autonomy will still run away, token cost and all, until something stops it.

Sources

OWASP Gen AI Security Project — LLM06:2025 Excessive Agency — https://genai.owasp.org/llmrisk/llm062025-excessive-agency/

Anthropic — Building effective agents — https://www.anthropic.com/engineering/building-effective-agents

Model Context Protocol — Specification (version 2026-07-28) — https://modelcontextprotocol.io/specification/2026-07-28

Anthropic — How we built our multi-agent research system — https://www.anthropic.com/engineering/multi-agent-research-system

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Working on a project where these methods apply?