· via dev.to (home feed)
Timeouts don't mean failure: why agent retries need idempotent operation IDs
A dev.to post explains why an agent timeout says nothing about whether the remote write failed, and how operation IDs, explicit unknown states and reconciliation prevent duplicate effects.

A post on dev.to examines a failure mode that agent deployments hit as soon as their tools write to real systems: a request that times out after the server has already committed the change. The agent sees the timeout, retries, and the result is two published documents — while the logs tell the story of a clean recovery. The author frames this as a design scenario rather than a benchmark result, and argues that one question must be answered before an agent gets tools that mutate external state: how does the system behave when it cannot determine whether an action took effect?
A timeout proves ambiguity, not failure
According to the post, a timeout only establishes that the caller never received a timely result; it says nothing about whether the remote operation failed. The author points to the Amazon Builders' Library material on idempotent APIs, which covers exactly this ambiguity and the standard mitigation: caller-provided request identifiers that make retries safe. For agent workflows, that uncertainty should be made explicit in the application contract rather than left to the model.
Give uncertainty its own state
A boolean success flag cannot describe every outcome of a remote write. The post proposes five states: Ready (authorized and durably recorded), In flight (a worker has claimed it and it may have reached the destination), Succeeded (the destination confirmed the outcome), Failed (conclusive evidence it did not take effect) and Unknown (it may have taken effect, but evidence is missing, so it goes to reconciliation). Expired worker leases deserve special attention, since a worker can crash after the remote commit. On the UI side, the author suggests honest wording — the request was submitted but completion is unconfirmed — which is less satisfying than "Done" but reflects reality.
One identity per logical operation
A tool-call ID identifies a single attempt; the application needs an identity that spans attempts. The post sketches a record combining an operation ID, actor ID, action, target, content version and a request fingerprint, created and persisted in application code and carried across retries and restarts. Where the destination supports idempotency keys, the same key should be reused for the same operation under that API's contract. The fingerprint has a separate job: catching changed parameters, so reusing an operation ID with a different payload should raise a conflict. A hash alone cannot express intent, the author notes — two deliberate operations can carry identical payloads. Identifiers should be scoped to the actor or tenant, uniqueness enforced atomically, and the provider's retention window checked, because an expired key no longer protects a retry.
Keep authorization attached to the approved action
When a workflow requires approval, record exactly what was approved: the action, target, payload version and applicable limits. A retry of that exact operation can reuse the recorded authorization where policy allows, but a model that alters the payload has proposed a different action and should not silently inherit the old approval. Permissions should also be rechecked at execution time, since a durable approval record should not override a later revocation.
A local ledger cannot close a remote transaction
The hard sequence is: record the operation locally, send the remote write, the destination commits, and the worker crashes before recording the result. No local transaction can make those steps atomic with an unrelated service. Recovery then depends on the destination: with proper idempotency support, retry under that contract; with an authoritative lookup by operation ID, query and reconcile; with neither, keep the unknown outcome and route it for investigation before another consequential write. The post adds two caveats: an empty search result may be inconclusive under eventual consistency, and a durable queue improves delivery without removing the consumer's duplicate-attempt problem.
Test the inconvenient boundary
Before shipping, the author recommends injecting the awkward conditions: a response lost after the remote commit, two workers dispatching the same operation, the same ID arriving with a different payload, a worker dying after the commit, a temporarily stale lookup, and an expired idempotency record. These exercise the application around the model — they are not prompts asking the model to be more careful.
Why it matters
Agents are being wired into systems with real side effects, and this failure mode is quiet: the log looks healthy, the retry looks like resilience, and the duplicate only appears downstream. The post's practical review question is worth adopting: if a write commits and its response disappears, what evidence does the next worker use to decide what happened? If the only answer is that the agent will work it out, the recovery protocol is unfinished — idempotency, explicit unknown states and reconciliation belong in the application contract, not in the model's discretion.
- #ai-agents
- #idempotency
- #distributed-systems
- #reliability
- #api-design