Retry the step, not the run: what an agent re-executes when it starts over
Published · Updated
When an unsupervised agent fails halfway, re-running it looks free. It isn't a replay: it re-executes tool calls, re-spends tokens and re-enters the loop that failed.
A failed agent run invites the cheapest-looking fix: run it again. That instinct comes from stateless web requests, where a repeat is harmless. An agent run is not a request. It is a sequence of real tool calls, each of which may already have changed something, plus a token bill you pay a second time. These are field notes on designing recovery at the step boundary instead of the run boundary.
A retry re-executes, it does not replay
The Model Context Protocol specification is blunt about what a tool is: tools represent arbitrary code execution and must be treated with appropriate caution. Re-running a step means running that code again. The specification also warns that descriptions of tool behaviour such as annotations should be considered untrusted unless obtained from a trusted server. So you cannot lean on a tool's own metadata to decide whether repeating it is safe.
That puts the judgement back in your code. Anthropic's distinction helps: workflows are systems where LLMs and tools are orchestrated through predefined code paths, while agents are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks. The more dynamic the run, the less you can predict which tools a second attempt will reach for, and the less a restart resembles a replay of the first.
Checkpoints make resuming cheaper than restarting
Anthropic describes building systems that can resume from where the agent was when the errors occurred, combining the adaptability of AI agents built on Claude with deterministic safeguards like retry logic and regular checkpoints. Note where the intelligence sits. The retry logic is deterministic, not a decision the model makes about its own failure. The checkpoint is the record that lets recovery begin at the step that broke rather than at the first step.
The companion pattern is memory. Anthropic implemented patterns where agents summarise completed work phases and store essential information in external memory before proceeding to new tasks. For recovery, that summary is the payload. A resumed run reads what the earlier phases concluded instead of reconstructing context by repeating the work that produced it, which is the expensive half of any restart and the part most likely to drift.
Starting over is a billing decision
The multipliers make this concrete. In Anthropic's data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats. A full restart pays that multiplier again for the portion that already succeeded. Anthropic is direct about the trade: the autonomous nature of agents means higher costs, and the potential for compounding errors, which is why they recommend extensive testing in sandboxed environments along with appropriate guardrails.
Ask about the operation instead of relaunching it
The Model Context Protocol specification documents Tasks: asynchronous execution of long-running operations, with polling, mid-flight input, and durable handles. A durable handle changes the recovery question. Rather than asking whether to run the step again, the orchestrator asks what became of the step already in flight. Long operations are exactly where a timeout is easiest to mistake for a failure, and where a blind second attempt does the most damage.
Caps keep a retry from becoming a loop
Anthropic notes that it is common to include stopping conditions, such as a maximum number of iterations, to maintain control. OWASP approaches the same risk from the permissions side: the root cause of excessive agency is typically excessive functionality, excessive permissions or excessive autonomy, and one recommended control is rate-limiting to reduce the number of undesirable actions that can take place within a given time period. A rate limit caps repetition without needing to know why it repeated.
Two further OWASP controls matter specifically for retries. Implement authorisation in downstream systems rather than relying on an LLM to decide if an action is allowed, and use human-in-the-loop control to require a human to approve high-impact actions before they are taken. Both are enforced outside the agent, so a second attempt at a high-impact action meets the same gate as the first, however convincingly the model argues for it.
Fallback is not the same move as retry
On our own platform, the autonomous publishing engine uses resilient multi-model fallback with 3 fallback levels, publishing articles daily without supervision at roughly 0 h of weekly content operations time. Fallback and retry answer different failures. Retry repeats the same call and assumes the fault was transient. Fallback changes who is asked, which is the better response when the fault is the model or provider rather than the moment.
Treat recovery as a design surface, not an exception handler. Write down which steps are safe to repeat and which are not, checkpoint between phases so a resumed run can skip what already succeeded, keep the retry logic deterministic and outside the model, and put a hard iteration or rate cap around the whole loop. The question worth answering before launch is not whether a run can fail, but what exactly it re-executes when it does.
Sources
Anthropic — Building effective agents — https://www.anthropic.com/engineering/building-effective-agents
Anthropic — How we built our multi-agent research system — https://www.anthropic.com/engineering/multi-agent-research-system
Model Context Protocol — Specification (version 2026-07-28) — https://modelcontextprotocol.io/specification/2026-07-28
OWASP Gen AI Security Project — LLM06:2025 Excessive Agency — https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
Neurolinks case study — A personal brand that publishes itself — https://neurolinks.be/work/matthieu-pesesse-media
Working on a project where these methods apply?