The agent sends the request. The service performs the action. The connection breaks before the receipt arrives. From the agent’s perspective, the request failed. From the customer’s perspective, it may already have changed something. Retrying without understanding that gap can create a second effect.
Idempotency means that repeating the same intended operation does not create an additional effect within its defined boundary. Automated systems need stable identities and retained state because retries are ordinary, not exceptional. The difficult question is deciding which attempts represent the same operation and how to discover what happened when the result is uncertain.
This chapter of The AI Software Factory uses a hypothetical paid diagnostic job. It examines duplicate prevention separately from durable execution, which preserves progress, and safe writes, which binds authority to an action. None of those mechanisms alone guarantees that every distributed operation happens exactly once.
Separate an operation from an attempt
A customer asks for one diagnostic report and agrees to pay the stated amount. The application creates an operation. A worker may make several attempts to complete it because of network failures or process restarts. Those attempts should share the operation’s identity.
If the customer deliberately orders another report from a revised input, that is a different operation. Deduplicating it against the first would deny legitimate work. A stable key therefore needs to represent the intended business action rather than merely the customer, date or current worker.
The identity can be an opaque random value associated with the authorized request. Keep customer information out of the key when it is unnecessary. Store the relationship between the key, actor, payload version and expected effect in the application record. A worker should receive that retained identity instead of generating a new one every time it starts.
This distinction is simple to explain and easy to break in code. A retry helper that constructs a fresh identifier inside its loop makes every attempt look new. A key based only on customer email makes every purchase look the same. The identity rule belongs in the operation contract and its tests.
Define the effect boundary precisely
A job can contain several effects: generating a report, recording entitlement, charging payment and notifying the customer. “The job is idempotent” is too broad unless each relevant boundary is defined.
Generating a deterministic calculation twice may be harmless if the result replaces the same version. Charging twice is not harmless. Sending a notification twice may be annoying or confusing even when it has no additional monetary effect. Publishing two differently worded versions can create uncertainty about which result the customer bought.
Give consequential suboperations their own identities linked to the parent job. The payment identity can remain stable across payment retries; the publication identity can remain stable across delivery retries. A deliberate revised report gets a new version and an explicit relationship to the prior publication.
The boundary should also name what is outside its guarantee. A payment provider may deduplicate a particular API request while the application still creates two separate requests with different keys. A database uniqueness constraint may prevent duplicate local records while a notification service receives two messages. End-to-end behavior depends on the connections between these guarantees.
Retain the payload that the identity authorizes
An operation identity without a stable payload can hide a changed action. Suppose a retry uses the same key but a different amount or destination. The system needs to reject the mismatch rather than quietly select whichever version arrived first.
Store or derive a canonical representation of the authorized fields. For a paid diagnostic, that might include the customer account, report input version, price, currency and delivery destination. The approval record should identify the same intended action. Changing a consequential field requires a new authorization or other explicit revision path.
The representation should distinguish missing values from empty values where the product does. It should normalize only what is semantically equivalent. Treating “USD” and an absent currency as interchangeable would defeat the very validation the app promised.
A payload fingerprint can help detect changes, but the application still needs the underlying fields for explanation and recovery. A hash proves equality under the chosen representation; it does not prove that the represented action is valid or authorized.
Make claiming work atomic
Two workers can receive the same job at the same time. If both read “not started” and then perform the effect, a later duplicate record check is too late. The operation needs an atomic way to claim or transition the work.
Within a database boundary, a uniqueness constraint or conditional state transition can ensure only one worker acquires a particular execution claim. The exact mechanism depends on the storage system. The invariant is that the decision and state change cannot be separated into a race in which both workers proceed.
A claim may need an expiry or lease so a crashed worker does not block the job forever. Expiry creates another race: the old worker can resume after a new worker acquired the job. A generation number or other fencing mechanism can help reject stale ownership when committing local state. External services may require additional safeguards because they do not automatically understand the lease.
Do not assume a lock is sufficient across every boundary. A worker can hold a local lock and still lose the remote receipt. A second worker can arrive after the lock expires and repeat the action. Stable external operation identity and reconciliation remain necessary.
Local transactions cannot cover arbitrary remote systems
A database transaction can atomically update records within its supported scope. It cannot generally include an unrelated remote payment or email service merely because the application calls that service between two database statements.
Consider the sequence: mark the job processing, send the notification, mark the job complete. If the worker fails after sending but before completing, the local state invites a retry. Reversing the order creates a different failure: the database says complete before the notification is sent.
One pattern is to retain an outgoing operation record in the same transaction as the local business change, then let a worker deliver it. That record provides a durable identity and a place to retain attempts and receipts. It does not remove the remote ambiguity; it makes the obligation explicit and recoverable.
For incoming events, retain the event identity and associated processing state. Repeated delivery of the same event should return or preserve the prior result. Events arriving out of order need a state rule as well; deduplicating each event does not guarantee their sequence produces the intended customer state.
Use provider idempotency within its documented scope
Stripe documents idempotency keys for safely retrying create or update requests. Its idempotent request reference explains that repeated keys return the saved result and mismatched parameters can be rejected. The current documentation also describes key-retention limits. Those details were read October 7, 2026 and should be rechecked for an integration.
The application should retain the provider key with its own operation record and reuse it for the same intended request. A timeout should not trigger a new key automatically. The worker should inspect whether the provider returned a definitive result or whether reconciliation is needed.
Provider-specific behavior matters. A saved error result can behave differently from an error that occurs before execution begins. A key reused after the provider no longer retains it may no longer prevent a new effect. The application’s retention and retry policy should reflect those boundaries rather than assume a key protects the operation forever.
This is an API capability, not financial guidance or a claim about the economics of the app. Payment authority, pricing and customer consent still require their own controls. An idempotent unauthorized charge remains unauthorized.
Reconcile uncertainty before creating a replacement
After an ambiguous timeout, ask what evidence can establish whether the effect occurred. A provider receipt, operation lookup, current remote state or controlled event stream may help. The best route depends on the external system and the operation.
A lookup based only on amount and date can be ambiguous. Two legitimate orders may share both. Use the retained operation identity or other authoritative reference where available. Preserve an unresolved state when the evidence does not support a unique answer.
A person may need to investigate when the system lacks a reliable lookup. That is a real operating cost, and it belongs in the product’s exception design. The alternative should not be to send another action and hope the earlier one failed.
The reconciliation result should advance the same operation record. If it discovers success, retain the receipt and complete the appropriate state. If it establishes no effect and authority remains valid, a retry can proceed under the defined policy. If it discovers a partial or conflicting effect, move to compensation or review rather than invent a clean success.
Idempotent reads can still produce changing evidence
A read does not necessarily create a side effect, but repeating it can return a different result. A supplier quote may change, an account may lose permission or a document may be revised. A job that re-reads everything on retry can silently change the evidence underlying its output.
Retain an input snapshot or version reference appropriate to the task. A diagnostic approved from one snapshot should not be published as if it described a later one. A fresh read may be appropriate before commitment to confirm authority or current conditions; that purpose should be explicit.
Distinguish evidence freshness from repeat-safe execution. The calculation can reuse its retained input while the write step rechecks that the destination and permission remain valid. Those checks can stop the operation if the current state no longer permits the approved action.
This is especially important for model output. Regenerating an explanation can produce different wording or findings. If the original version received review, a new version should not inherit that review automatically. Independent verification establishes the evidence needed for the artifact actually used.
Design customer-visible states for incomplete work
A user should know whether the request was accepted, is processing, needs attention or completed. Ambiguous remote effects should not be reported as a simple failure if that encourages the customer to order again and creates a second operation.
For the hypothetical report, the interface could say that payment confirmation is being reconciled and that the customer should not resubmit. The wording should be accurate and provide a support route. The system should avoid exposing internal identifiers unnecessarily while retaining them for authorized investigation.
A repeated click on the same submitted request should return the current operation state rather than create a new purchase. A deliberate new purchase should require a clear new action. The user interface and backend identity rule need to agree on that difference.
Support staff need the same view. If the dashboard shows only an error, an operator may manually resend or recharge without seeing the ambiguous first attempt. Provide the state, evidence and allowed next actions so support cannot accidentally undo the duplicate-prevention design.
Test races and gaps, not only repeated calls
A test that submits the same request twice sequentially is useful but incomplete. Run concurrent submissions and interrupt execution at the points where state and effect can diverge. Expected outcomes should include both the retained result and the absence of duplicate effects.
For the hypothetical report, protected cases could include two workers claiming the same operation, a lost payment receipt, repeated webhook delivery, a stale lease holder resuming and a changed payload using the same identity. Another case should verify that an intentional second order is accepted as new.
A controlled fake service can count effects precisely and simulate timeouts after success. That establishes local behavior under the simulated condition. A separate safe integration check is needed to confirm actual provider behavior. Neither result should be described as a universal exactly-once guarantee.
Retain failed cases and configuration. If the identity rule changes, rerun cases that distinguish retries from new operations. If a provider changes its retention or API semantics, revalidate the relevant boundary rather than assume the old test covers it.
Retire operation records without reviving old effects
The application cannot retain every operational detail forever, yet removing an identity can make an old retry look new. Define a retention window that reflects the provider’s behavior, customer dispute needs and the maximum plausible retry delay. Separate sensitive payload retention from the smaller record needed to recognize a completed operation where policy permits that distinction.
Old events should have a deliberate rule. A delayed message arriving after the job has been retired may be rejected, reconciled against an archived receipt or sent for review. It should not automatically create a fresh effect because the worker no longer finds the original row.
The same issue appears in manual exports and restores. Restoring business records without their operation identities can revive work the system already performed. A recovery exercise should include a completed operation and verify that a repeated event remains recognized after restoration. Backups need the state that preserves the customer’s obligation, not merely the table that displays the result.
These policies are application choices requiring evidence about actual timing and obligations. A short provider key window does not settle the product’s own retention needs, and indefinite retention can create unnecessary privacy exposure. Document the selected boundary and revisit it when the integration or customer contract changes.
Include duplicate prevention in the economics
Idempotency can reduce repeated charges, duplicate notifications and unnecessary compute. It also costs storage, engineering and reconciliation work. Those costs should be evaluated against the consequences of the operation.
A read-only preview may tolerate recalculation, although repeated model calls still consume resources. A paid purchase or customer record update needs stronger controls because the effects matter beyond the application. The mechanism should be proportionate to the customer obligation.
Track attempts per accepted operation, ambiguous outcomes and manual reconciliation time. These measures can reveal a flaky integration that technically recovers but consumes too much support effort. Contribution margin depends on the real delivery burden, including exceptions that a clean demonstration omits.
Avoid claiming savings without a baseline. A proposed design may prevent a class of duplicate effects in tests; a production cost claim needs observed attempts, costs and customer consequences over a defined period. Retaining that distinction makes the engineering result useful without inflating it into an unsupported business result.
What Would We Do at Salars?
A proposed Forge template would include an operation record for consequential actions. It would retain identity, actor, authorized payload version, state, attempts and receipts. Application-specific rules would determine whether a new request is a retry or a new customer decision.
For a proposed Supplier Margin Guard diagnostic, input processing would be separate from any paid entitlement or customer notification. The first pilot would remain read-only where possible. A later paid delivery workflow would require safe provider integration tests and a reconciliation route before launch.
The evaluation would deliberately lose receipts and repeat events in a controlled environment. It would also test legitimate second orders so duplicate prevention does not become accidental denial of service. Success would be scoped to the tested operations and configuration, not declared for every future Forge app.
An operation should have one meaning across its attempts. Stable identity, atomic local transitions and honest reconciliation let automated software keep that meaning when the network does not provide a clean answer.
Sources
- Stripe: idempotent requests, read October 7, 2026, for provider-specific key and retry semantics.
- Temporal: Activity overview, read October 7, 2026, for idempotent external work on retry.
- Salars AI library, for related agent and verification guides. State designs, customer examples and evaluation cases are proposed or hypothetical, not observed Salars results.
Loading comments…