AI · Article 29 of 72 · Part 6

How to Build Durable Agent Workflows

Explain persisted state, checkpointing, replay, timers and compensation through one multi-step job.

A report is ready, but the customer has not approved publication. The worker shuts down overnight. In the morning, another worker starts. Should it regenerate the report, send the old one, wait for approval or assume the whole task failed? Without a retained state, every answer depends on reconstructing the past from incomplete clues.

A durable agent workflow preserves enough execution history and decision state to continue a defined job after interruption. It needs explicit steps, stable operation identities, valid events and clear terminal outcomes. Persistence makes progress recoverable. It does not make the decisions correct or the external effects reversible.

This chapter of The AI Software Factory develops a hypothetical diagnostic-publication job. Agent orchestration owns dependency scheduling; Cloudflare runtime selection owns service mapping. Here the central question is how a job can survive time and failure without silently changing its meaning.

Begin with the job’s state, not the agent’s memory

A conversation transcript can describe what happened, but it is not a reliable execution record by itself. A model can misunderstand a previous message, omit an unresolved condition or treat a suggestion as approval. The workflow needs structured state that distinguishes completed work from planned work.

For the hypothetical job, define accepted input, validated input, diagnostic prepared, awaiting approval, approved, publication attempted, published and failed. Each state has an entry condition. Prepared requires a retained diagnostic version. Approved requires a valid authorized decision about that version. Published requires evidence that the intended destination contains the accepted artifact.

The states should describe the business job rather than the model’s feelings about progress. “Thinking,” “nearly done” and “looks good” can be interface messages, but they cannot determine whether a customer write is allowed. A worker can update explanatory text while the governing state remains awaiting approval.

The application should also identify the job. A customer asking for a revised diagnostic may create a new version within the same job or a new job linked to the old one. The identity rule should be explicit so a retry is not mistaken for a fresh request and a fresh request is not deduplicated away.

Choose steps that retain useful boundaries

A step should represent a unit of work whose result can be retained and whose failure can be understood. “Do everything” gives the runtime little useful information. An extremely fine split can create excessive state and make the job hard to read.

For the diagnostic, input validation, evidence preparation, deterministic calculation, explanation generation, approval wait and publication are natural boundaries. Each has different inputs, failure causes and authority. The explanation step can be retried without repeating a publication. The approval step can wait without keeping a processor busy.

Step results should be versioned and associated with the input they describe. If a supplier quote changes, the old explanation no longer applies automatically. Retain the relationship between input snapshot, calculated findings and generated text. This lets a reviewer see exactly what was approved.

External actions deserve particularly clear boundaries. A publication attempt should have its own operation identity and result record. Combining report generation and delivery in one opaque step makes it harder to determine whether a failure occurred before or after the customer received something.

Replay requires stable decisions

Some durable systems reconstruct progress by replaying execution against recorded history. Temporal’s workflow execution overview explains that generated commands are checked against event history and progress resumes from recorded events. That design imposes deterministic constraints on workflow logic.

The practical consequence is that nondeterministic work needs an appropriate boundary. A model response, current clock reading or external service result can change between attempts. The workflow should use the runtime’s prescribed mechanisms to record results and control decisions, rather than assume rerunning arbitrary code will reproduce the same path.

Suppose the model first recommends “flag three rows” and later recommends “flag four.” If replay regenerates the explanation casually, the job may no longer contain the artifact the reviewer saw. Retaining the accepted output keeps later progress tied to the original evidence. A deliberate revision can create a new version with a new approval requirement.

Runtime semantics differ. Do not copy an implementation pattern from one system into another merely because both advertise durable execution. Read the chosen SDK’s rules for state, replay, external calls and version changes. The conceptual requirement is stable retained decisions; the implementation must fit the actual service.

Activities and external work have different failure semantics

The workflow’s control logic can be recoverable while a remote operation remains uncertain. Temporal’s Activity overview recommends idempotent activities to avoid duplicate side effects on retries. That distinction is central to any durable agent design.

Imagine publication succeeded at the destination but the response was lost. The workflow sees no receipt. Restarting the step can publish again unless the destination or application recognizes the same operation. Durable state cannot infer a remote fact it never recorded.

Use an operation identity, a retained payload and a reconciliation route. If the destination can return the result for that identity, inspect it before creating another action. If it cannot, design a conservative ambiguity state and an authorized investigation rather than assume the first request failed.

Idempotent software agents explains the detailed duplicate-effect mechanism. This chapter’s concern is preserving the distinction between workflow progress and remote consequences so the recovery path remains honest.

Approval is an event with meaning

An approval event should identify who approved, what version was approved, what action is allowed and when that decision remains valid. A generic message saying “yes” is insufficient when several pending artifacts or destinations exist.

The hypothetical merchant diagnostic could bind approval to the report hash, input snapshot and publication destination. If any of those change, the old approval should not silently authorize the new action. The workflow can return to review or create a revision path.

Validate the event at the boundary. A webhook body claiming approval needs authentication and an association with the correct job. A forwarded email or an instruction embedded in an uploaded document is not automatically an authorized event. The workflow should not let an agent’s interpretation bypass the application’s authority checks.

Approval can expire for practical reasons. The input may become stale, the customer may revoke access or the intended destination may change. The commitment step should recheck applicable conditions even when an earlier event was valid. Safe AI writes treats this binding between preview and actual commitment.

Waiting should have a policy

A durable workflow can wait for a long time, but the product still needs a rule for what happens while waiting. Is the job visible to the customer? Does the report expire? Will reminders be sent? Who owns unanswered review requests?

Define a timeout or other stopping condition appropriate to the obligation. A diagnostic based on a current supplier quote may lose usefulness after the quote changes. A wait can end in expired rather than failed, preserving the prepared artifact while preventing stale publication.

Reminders are external actions with their own permission and duplication concerns. A workflow waking periodically should not send a new message every time it resumes. Record reminder state and respect customer communication preferences. The ability to schedule work does not grant permission to contact someone indefinitely.

Cloudflare’s Workflows overview documents sleep and external-event waits. Those runtime capabilities help implement the policy. The builder still decides which events count, what waiting means and how the customer sees the outcome.

Retrying is a decision about the failure

A transient network failure can justify a retry. Invalid input usually needs correction. A missing authorization requires a different path. Treating all exceptions as retryable can turn a small defect into repeated spending or repeated customer disruption.

Classify failures at the step boundary. A validation error can return actionable fields to the customer. A temporary dependency outage can retry with a bounded policy. An ambiguous external write can enter reconciliation. A revoked credential should stop consequential work until valid authority is restored.

The retry budget should include attempts, elapsed time and cost where relevant. Model calls can be expensive even when each one succeeds technically. If the model repeatedly produces an unsupported explanation, another attempt without a changed constraint or input may have little value.

A worker should retain the reason it stopped. “Failed after retries” needs the step, relevant error and remaining obligation. A person can then decide whether to correct input, change the workflow or end the job. Durability should make exceptions more understandable rather than hide them behind endless background activity.

Cancellation is a workflow outcome

Stopping a job does not undo everything it has already done. A customer may cancel while a diagnostic is prepared, while approval is pending or after publication begins. Each condition has a different consequence.

Before publication, cancellation can prevent the write and retain or delete intermediate data according to policy. During an ambiguous write, the job may need reconciliation to learn whether an effect occurred. After publication, the product may offer withdrawal or correction, which is a new authorized operation rather than an erasure of history.

The workflow should acknowledge cancellation in a state the customer can understand. A button that says cancelled while the worker continues sending messages breaks the promise. Check the stopping signal at relevant boundaries and make irrecoverable work explicit.

The rollback chapter covers reversal and compensation. Durable workflow design needs to preserve the events that tell the recovery mechanism which situation it faces.

Version the workflow while old jobs remain alive

A short-lived request finishes before the next deployment. A durable job can outlive several releases. Changing the workflow definition can affect jobs started under an older rule, especially when replay expects a historical sequence.

Plan an upgrade strategy using the chosen runtime’s supported mechanisms. You may keep older workers available, route jobs by version or use a compatibility path. The correct technique depends on the SDK and service. The important requirement is that a new deployment does not invalidate progress simply because an old job is still waiting for approval.

Separate business policy changes from technical refactoring. If the customer approved a report under one publication rule, a new release should not reinterpret that approval casually. A migration of workflow state needs the same attention to obligations as a database migration.

Keep version metadata where an operator can inspect it. When a job fails after deployment, the investigator needs to know which definition and input version governed it. A moving branch name is insufficient. Retained artifact identity connects the event history to the code that interpreted it.

Bound retained history and sensitive state

Durable jobs accumulate events, results and evidence. That history supports recovery and audit. It can also grow large or retain customer information longer than needed.

Store references to large authorized objects where the architecture permits it, rather than duplicating full documents across every event. Define which intermediate results must remain for replay or accountability and which can expire. Read the runtime’s payload and history limits before relying on unlimited accumulation.

Privacy obligations can conflict with casual retention. If a customer requests deletion, the app must know which records contain their information and how deletion affects unresolved jobs. A workflow cannot remain authorized to publish data after the relevant access has been revoked.

The retained state should make uncertainty visible. A missing remote receipt should remain unresolved, not be overwritten by a later summary claiming success. An audit trail is useful when it preserves the actual evidence rather than simply narrating a confident ending.

Test interruptions at the awkward moments

A workflow’s happy path proves little about durability. Interrupt it after calculation but before saving the result, after saving but before returning the receipt, after approval and during publication. Those boundaries reveal whether the state and effect model work together.

A proposed test suite should include repeated events, stale approvals, changed input, unavailable dependencies and cancellation. Expected outcomes should be written independently of the implementation. The suite should assert both the visible result and forbidden effects, such as duplicate publication or use of an outdated approval.

Retain safe fixtures and operation records so a failed case can be reproduced. A passing test in an isolated environment does not establish that the production provider or external API is healthy. Separate controlled recovery tests from live integration checks and report the limits of each.

A useful local result may establish that one diagnostic workflow survives named interruptions under a specific configuration. It does not establish that every possible agent job is recoverable. Revalidate when the workflow adds a new external effect, changes runtime or expands its data responsibilities.

Keep operator repair within the same evidence model

An operator may need to repair a job whose automatic path cannot proceed. That repair should identify the affected state, the evidence justifying the action and the resulting transition. Editing a database row directly without retaining the reason can make later replay or investigation confusing.

Provide a limited repair interface where the product warrants it. It might allow an authorized operator to attach a verified remote receipt, expire a stale approval or retry a corrected input version. Each action should preserve who acted and which obligation changed. A broad command to “mark complete” can conceal unfinished customer work.

Test the repair path with safe failures. Confirm that an operator cannot approve a different customer’s job or overwrite a published version silently. Recovery authority can be powerful; it should remain scoped even when the ordinary worker is unavailable. The retained repair record then becomes part of the job’s history rather than an unexplained exception outside it.

Decide whether durability earns its cost

Not every model call needs a durable workflow engine. A small read-only request with a clear retry path may work well as an ordinary handler. Durable orchestration adds services, state, versioning and operational knowledge the team must maintain.

The case becomes stronger when jobs span time, need approvals, involve several dependent effects or must resume without repeating expensive work. Compare the expected failure and recovery burden with the complexity of the durable design. A scheduled task can be simple until it creates a customer obligation that must survive interruption.

Use a bounded pilot to inspect that comparison. Record recovery effort, duplicate effects, missed jobs and operating cost for a baseline and the proposed design. These measurements can support a scoped architecture decision. A platform feature checklist cannot provide the same evidence.

What Would We Do at Salars?

A proposed Forge pilot would choose one read-only merchant diagnostic with a publication review step. It would specify state transitions, input and report versions, operation identities, authority checks and terminal outcomes before choosing the runtime implementation.

The initial recovery exercise would stop workers at named boundaries and inspect what remained. It would repeat an approval event and submit one for an outdated report. Success would require an understandable job state and no unauthorized or duplicate publication in the protected cases. This is a proposed evaluation, not a completed Salars result.

If the design met those cases within the pilot’s cost and maintenance budget, Forge could retain the pattern for similar jobs. A later workflow changing prices or billing would need new permission, reconciliation and recovery evidence. The earlier diagnostic would not authorize the broader action.

A durable job should resume with the same obligation it began with. Retained progress is valuable when it preserves that meaning across time, workers and failure.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home