An agent app has three different needs. It must answer a request quickly, remember the right customer state and continue a job after the user closes the page. Calling all three “an AI agent” hides the runtime decisions that determine whether the product can meet its promise.
Cloudflare offers building blocks for request execution, persistent coordination and durable workflows, but the application still owns its rules, permissions and customer obligations. Choose a service for a requirement it satisfies. A feature list is a starting point for architecture, not evidence that a proposed app is ready for production.
This chapter of The AI Software Factory maps those decisions. Official documentation was read October 7, 2026; service limits and feature maturity should be rechecked before implementation. The merchant app used below is hypothetical. No scale, price or operating result is claimed for Salars.
Write the workload before selecting the service
Begin with the customer’s job. A merchant uploads a supplier quote, supplies cost assumptions and asks which lines need attention. The first version returns a bounded diagnostic with source rows and unresolved fields. It does not change prices or place an order.
Describe the workload in operational terms: input size, processing duration, state, external dependencies, concurrent access and acceptable delay. Does the customer need a response in the same request? Can the job continue asynchronously? Does more than one participant need to inspect and approve the result? Is there a consequential external write?
These questions distinguish service requirements. A simple request may need only an execution endpoint and a database. An ongoing session may need durable identity and coordinated state. A job waiting for an approval event needs persisted progress across time. A bulk import may need buffering and controlled processing.
The cloud development chapter describes the complete environment and release path. Here the narrower job is deciding which runtime capabilities belong in the product and which would create unnecessary complexity.
Separate the harness from the runtime
An agent harness controls how the model is called, how tools are selected, how results are interpreted and when execution continues. A runtime provides the environment in which that logic runs and retains state. Confusing the two can make an infrastructure feature look like a behavioral guarantee.
Cloudflare’s Agents overview explicitly distinguishes communication channels, harness, runtime and tools. Its current runtime description includes durable identity, local SQL state, real-time connections, scheduling and recovery. Those capabilities support application design; they do not determine whether the model selects the right tool or follows the customer’s actual authority.
A proposed merchant harness might first validate input, perform deterministic arithmetic and ask a model to explain flagged rows using a bounded evidence record. The runtime can retain the record and stream the explanation. The harness must still reject unsupported claims and keep missing information visible.
Do not give the model a broad store-management tool because the runtime can expose tools. The customer’s authorization is narrower: prepare a diagnostic. Runtime convenience should not convert a read-only analysis into permission to change the merchant’s business.
Use request execution for request-shaped work
A request handler accepts an input, checks identity and authorization, performs bounded work and returns an appropriate response. This is a useful starting point because it keeps the product’s first version small and its success state clear.
Cloudflare Workers supplies the execution platform around which its other services can integrate. The Workers overview describes the platform; workload-specific limits and API compatibility require deeper checking for the chosen implementation. This chapter does not assume that arbitrary existing server code can run unchanged.
For the merchant diagnostic, the handler can validate the upload metadata and create a job record. If the processing fits the request’s intended duration and resource limits, it may return the result directly. If the work needs more time, return an honest accepted state and a way to inspect progress. Do not hold a connection open merely to create the impression that the job is synchronous.
Authentication and tenancy belong at this boundary. A valid job identifier should not give any visitor access to the result. The handler needs to establish that the requester can view that job and its evidence. The agent’s durable memory must not substitute for an application authorization check.
Choose coordinated state when the job needs it
Ordinary application records and coordinated session state solve related but different problems. A list of completed reports may fit a conventional database. A shared session in which several participants update one live state may need a coordination primitive.
Cloudflare’s Durable Objects overview describes uniquely named compute with attached durable storage for stateful coordination. It identifies SQLite storage APIs as generally available in the current documentation. That service-specific statement does not make every related agent feature generally available.
A hypothetical approval session could use one durable identity for a merchant job. It would retain the proposed diagnostic version and current review state. Two users trying to approve different versions should receive a clear result rather than overwrite each other’s decision silently.
The application must define what the identity represents. One object per user, organization, job or conversation has different implications for access, load and lifecycle. “One agent per customer” is not enough if the customer has several users and many independent jobs. State ownership should follow the operational entity whose transitions need coordination.
Durable state still needs a lifecycle
Persistence is useful because it survives interruption. It also creates an obligation to manage retained information. Define which state is authoritative, which can be reconstructed and which should expire.
The merchant session might retain a sanitized input reference, calculated findings, review decisions and operation identities. It may not need to retain every model token or every uploaded document indefinitely. Long conversational memory can accumulate stale instructions, private information and conflicting assumptions.
A lifecycle policy should include creation, retention, export, deletion and failure recovery. If a merchant requests deletion, the application must identify the relevant records across storage services and derived outputs. Deleting the visible chat does not necessarily delete all supporting state.
The privacy design chapter owns that broader data lifecycle. Runtime selection should make the required lifecycle possible rather than leave it for a later launch checklist.
Use a queue to decouple bursts from processing
A queue lets an application accept work separately from performing it. This can help absorb bursts, batch operations and control how much processing runs at once. It also introduces states the product must explain: queued, processing, completed and failed.
Cloudflare’s Queues overview documents batching, retries, delays and dead-letter handling. Those features can support a bulk import. They do not mean the customer receives a correct result merely because a message was delivered.
For a hypothetical merchant upload containing many supplier files, the request creates authorized jobs and enqueues references. Consumers process them within a defined concurrency and resource budget. Invalid files should reach a visible terminal state rather than retry endlessly. A dead-letter path needs an owner and a way to investigate why work could not proceed.
Keep the queue message small and purpose-specific. A reference to an authorized stored input can be easier to govern than embedding the entire sensitive document in each message. The consumer should validate the referenced operation and its current permission before writing a result.
Workflow durability is about progress through steps
A multi-step job can need to wait for external events or continue after failure. Cloudflare Workflows documents durable steps, retries, sleep and event waits. Its overview supplies the current capability description.
The merchant example could prepare findings, wait for a reviewer and then publish the approved diagnostic to the account. That sequence needs persisted progress and an approval bound to the intended version. A workflow can provide the execution structure; the application defines which event counts as valid approval.
The dedicated durable workflows chapter examines step boundaries and recovery. Runtime selection should not duplicate that guide. The relevant choice here is whether the app has enough long-running coordination to justify a workflow service instead of a simpler job record and worker.
Queue and workflow designs can coexist, but each should have a distinct job. A queue buffers many independent imports; a workflow coordinates the steps of one import. Adding both without that distinction can create two overlapping sources of progress state and confusing recovery behavior.
External effects need their own identities
A worker may update a record successfully and lose the response. A retry can repeat the effect. Durable execution and reliable messaging do not remove this uncertainty at an external boundary.
Give each intended operation a stable identity. Store the identity with the authorized payload and resulting receipt. If the task retries, use the same identity where the receiving API supports it, and reconcile ambiguous outcomes before creating a new operation.
The application should distinguish generating a new diagnostic version from publishing an already-approved version. A retry of publication should not create multiple customer notifications. An intentional revised report should have a new version and clearly related operation state.
Idempotency for agents supplies the detailed mechanism. It matters to the Cloudflare design because a runtime can safely preserve its own state while the external system still receives duplicated actions.
Read maturity labels feature by feature
A platform can have a mature runtime and a beta integration on the same documentation page. The current Agents navigation labels Models and Pi features as beta. The Durable Objects overview separately describes SQLite storage APIs as generally available. These are narrower statements than “Cloudflare agents are all production-ready.”
Record the exact feature, documented maturity, relevant limits and account availability for the proposed design. If the app depends on a preview feature, identify the alternative or the consequence of change. A beta label does not automatically prohibit a pilot, but it changes the durability of the assumption.
Vendor scale claims also need scope. An overview may describe very large instance counts. That is the vendor’s platform assertion, not a load test of the merchant app, its model provider or its external APIs. The product can hit another bottleneck long before the runtime reaches its advertised capacity.
A small factory should test its own workload at plausible volumes and failure conditions. Its evidence should identify the configuration and dependencies used. That produces a useful local result without pretending to establish a general platform benchmark.
Inspect limits before the architecture depends on them
Resource limits include more than request count. CPU time, wall-clock duration, payload size, persisted state and concurrency can constrain different parts of the design. A workflow that waits for hours may consume little active CPU, while a brief transformation can exceed a CPU allowance.
The current Workflows limits distinguishes several of these categories. Its concurrency table and narrative are not entirely consistent in the read material, so this chapter does not publish a specific paid concurrency guarantee. A real implementation should resolve the applicable allowance before relying on it.
Design tests around the relevant boundaries. What happens to an input larger than allowed? Can the app reject it before accepting an obligation? Does a partial import preserve completed work? Can the customer see which records remain unresolved?
Limits should become product behavior. A hidden platform exception is poor communication. A clear supported-size statement and an actionable rejection message let the customer choose another route. The product’s promise should fit the architecture it can actually deliver.
Observability must connect state and consequence
A runtime dashboard can show invocation counts and errors while the customer waits on a job whose state never advances. Track the customer’s operation across handler, queue, workflow, model and storage boundaries.
Use a correlation identifier and a compact state record. Retain when work was accepted, what input version it used, which step failed and whether an external effect occurred. The operator should be able to distinguish “no result produced” from “result produced but notification failed.” Those conditions require different recovery actions.
Avoid logging sensitive content simply because it helps the first debugging session. A redacted evidence reference and a structured error can be more useful over time than an unrestricted transcript. The retention policy should preserve the facts needed for support without making general logs a second customer-data archive.
Observability also reveals cost. A repeated model call may look like a harmless retry until it multiplies per-report expense. Count attempts and usage against the intended operation so the COGS ledger reflects what delivery actually consumed.
Plan an exit that preserves useful state
A runtime choice should leave a practical path for exporting the product’s authoritative records. Define a versioned format for job state, customer-owned inputs and accepted results where retention permits. Check whether another authorized environment can read that format without the original live session. The point is not to promise a painless migration. It is to avoid discovering that the only usable representation of a customer obligation exists inside an opaque runtime instance. An export rehearsal with safe records can reveal missing identifiers, undocumented schemas and derived fields that cannot be reconstructed.
Test the smallest useful runtime design
A proposed pilot should start with one customer workflow and a limited set of services. Define the expected result independently of the implementation. For the merchant diagnostic, that might mean correctly rejecting ambiguous currency, retaining source references and returning a completed or explicit unresolved state.
Test normal, boundary and recovery cases. Interrupt processing after a result is stored but before the completion receipt is returned. Repeat a message. Submit two conflicting approvals. Remove an external dependency temporarily. Confirm that the application state remains understandable and that no unauthorized write occurs.
These are proposed experiments rather than executed Salars tests. Their value is in exposing which runtime assumptions the app relies on. If a simpler request handler and job record satisfy the requirement, there is no need to add an agent harness merely because the product uses a model.
Retain the configuration, test inputs and observed results. Revalidate when the workload changes materially, the service API changes or the app adds a consequential operation. Evidence for a read-only diagnostic does not establish safety for automatic pricing updates.
What Would We Do at Salars?
A proposed Salars Forge evaluation would begin with a read-only merchant diagnostic. It would list the minimum required Cloudflare services and the reason each is included. GitHub would retain the accepted specification and release evidence; the runtime would retain customer operation state under explicit access and retention rules.
Supplier Margin Guard would remain a product hypothesis until customer tests establish value and economics. Its first runtime test would inspect valid inputs, missing currency, repeated submissions and a stopped worker. The report would distinguish advertised capabilities from observed app behavior.
A later proposal might add an approval step for a customer action, but that would require a new authority and recovery design. The presence of a tool integration would not authorize changing live merchant records. The customer promise and permission boundary would determine the next step.
Cloudflare can provide useful execution and state primitives. The builder’s task is to assemble only the ones the product needs, then establish that the resulting application meets its own bounded promise.
Sources
- Cloudflare: Agents overview, read October 7, 2026, for harness/runtime distinction, state capabilities and feature-specific beta labels.
- Cloudflare: Durable Objects overview, read October 7, 2026, for uniquely named stateful coordination and SQLite API maturity.
- Cloudflare: Queues overview, read October 7, 2026, for batching, retries, delays and dead-letter paths.
- Cloudflare: Workflows overview and limits, read October 7, 2026, for durable-step capabilities and workload boundaries.
- Salars AI library, for related app design guides. Platform documentation establishes documented capabilities; all merchant designs and tests here remain hypothetical or proposed.
Loading comments…