The laptop closes before the build finishes. The build continues. A second operator opens the project from another computer and sees the same runtime, dependencies and test commands. That is a useful cloud development capability. It does not tell either operator whether the release is correct, who authorized it or whether the customer’s data survived.
A cloud software factory puts development compute, shared artifacts and release controls in remotely managed environments while preserving accountable decisions. The operating goal is a reproducible path from a customer problem to a verified product change. Browser access makes that path portable; careful state, access and recovery design make it dependable.
This cornerstone chapter of The AI Software Factory follows that path end to end. The architecture is a proposed design for a small operator, not a claim that Salars already runs it. The detailed roles of GitHub as control plane and Cloudflare as runtime platform belong in their own chapters. Here the decision is how those pieces should fit, what must remain independent and when a cloud-only workflow is unsuitable.
Define what “entirely in the cloud” means
The phrase concerns where computation and authoritative working state live. A person still needs a client device, a network connection, an accessible interface and a way to authenticate. A browser can be the primary development interface without being the place that owns the repository or runs the production service.
Separate five environments: the development workspace, the automated check runner, the preview deployment, the production runtime and the operational record. They can share a provider, but their purposes and permissions differ. A development workspace should allow experimentation. Production should execute an accepted artifact under narrower authority. The operational record should preserve what happened when either environment fails.
A cloud-only policy should name its scope. Does it prohibit local compilation, or does it simply avoid requiring a powerful local machine? Can a person download a patch during an outage? Which customer data may enter a development environment? A vague policy can make routine recovery appear forbidden or let sensitive production records spread into temporary workspaces.
For a small factory, the useful requirement may be that any authorized operator can recreate the development environment from a repository and continue work without recovering a particular laptop. That is a concrete property to test. “Everything is cloud-native” is a description too broad to verify.
Start with the authoritative repository
The repository should contain the application source, environment definition, test entry points and release configuration needed to reproduce the product. It should also contain or reference the specification and decisions that explain important behavior. An agent’s conversational memory is a poor place to keep the only explanation of why a billing state exists.
A fresh environment should answer ordinary questions without archaeology: which runtime version is required, how dependencies install, which tests establish acceptance and which command builds the intended artifact. If setup depends on a private undocumented action, the factory is portable only to someone who remembers that action.
Keep secrets outside committed source. The environment definition can name required credentials and their purpose without containing their values. Production credentials should not become development defaults because that makes a demonstration convenient. A test account and a restricted development key are often enough for the actual task.
Repository authority also requires a clear release branch or other accepted state. Workers may explore on separate branches; the product must have one identifiable state intended for deployment. Parallel agent development explains why several successful working copies do not establish a successful combined release.
Make the development environment disposable
A disposable workspace is one that can be recreated without losing authoritative work. It may have valuable temporary state, but that state is deliberately distinguishable from the product record. Committed code, retained patches and approved data fixtures should survive the workspace’s removal.
GitHub Codespaces provides hosted development environments configured through repository files and accessible through a browser or development clients. Its remote containers run Linux. Those facts support a browser-based development route, while also limiting workloads that require a different operating environment. The current Codespaces introduction was checked October 7, 2026.
The proposed environment would pin runtime versions, define the installation process and expose named tasks for checks and builds. It would start from safe fixtures rather than a copy of production customer data. An operator should be able to create a fresh workspace, run the baseline checks and obtain the same intended application behavior.
Reproducibility has degrees. A pinned runtime can still install a changed dependency if the dependency graph is not locked. A locked graph can still download a compromised artifact. An identical build can still contain an incorrect policy. Environment definition, supply-chain controls and behavior verification answer different questions; none should be mistaken for a replacement for the others.
Give each environment its own authority
The easiest early prototype gives one account access to the repository, cloud runtime, customer data and billing system. That convenience makes mistakes travel far. A command copied into the wrong terminal can change production when the operator thought it was testing a preview.
Use environment-specific identities and clear labels. A preview should have preview data, preview URLs and credentials that cannot change production. The release job should receive only the authority required for the accepted deployment. A coding agent should not inherit a permanent administrator key merely because it needs to inspect an application configuration.
Access should be revocable without destroying the work. If an agent task ends, its temporary capability can expire while the patch remains available for review. If a person leaves the project, the repository and release process should continue under owned accounts and documented recovery procedures.
Human approval should occur where an action becomes consequential. Approving a general goal is not the same as approving an exact production migration or a customer-facing price change. The proposed factory would attach approval to a specific artifact and operation when the existing release process requires it. Preview, approve, apply and verify develops that transaction boundary.
Distinguish development compute from product runtime
A cloud workstation is not automatically the right place to host the product. It is optimized for tools, interactive editing and temporary experiments. A production service needs a deployment model, operational limits, observability and predictable customer access.
The runtime decision begins with the workload. A short request that classifies a small document differs from a large file conversion, a persistent interactive session or a workflow waiting days for approval. List execution duration, data size, connection needs, state, latency expectations and failure recovery before choosing a service.
Cloudflare’s current Workflows documentation describes durable steps, retries, sleep and external events. That provides a potential home for some long-running orchestration, not a promise that every operation runs once or that every workload fits its limits. Its Workflows overview and limits reference should be read for the intended operation.
A product can combine services without making the whole design unnecessarily complex. A simple request handler may invoke a queued job and store a result. The extra service earns its place when it solves a defined requirement, such as recovering progress after interruption. Adding queues, workflows and persistent objects to an app that needs none of them creates more state to inspect and pay for.
Follow one change through the factory
Consider a hypothetical merchant app that compares a supplier quote with the merchant’s own cost assumptions. A user reports that imported prices without a currency label are being interpreted as dollars. The accepted requirement is to reject ambiguous amounts and show the missing information clearly.
The issue record contains the observed behavior, a safe reproduction fixture and the intended result. A coding agent receives a bounded task to change the parser and associated tests. It works in an isolated environment with no live merchant records. The agent returns a diff, test evidence and a note about related fields it did not change.
An independent review checks the requirement and the patch. The automated checks verify parsing behavior, the user-visible error and the existing unambiguous paths. A preview deploys the combined artifact with safe example data. The reviewer confirms that an ambiguous row is rejected and that the interface tells the merchant how to correct it.
The release proceeds through the repository’s authorized workflow. The deployed artifact is identified. A live verification checks the public interface and a safe operational path. If production behavior differs from preview, the operator examines configuration, data and integration differences before declaring completion. The issue closes only when the promised behavior has evidence at the relevant boundary.
This sequence is intentionally ordinary. The factory’s value is in making ordinary changes repeatable and inspectable, including the unglamorous work between “agent finished” and “customer can use it.”
Build once and know what you released
An accepted source revision and a deployed artifact are related but distinct. A build can depend on environment variables, dependency artifacts and generated assets. Rebuilding the same source later may not yield the same result if those inputs changed.
Where the release tooling supports it, retain the artifact or its identifier and promote the accepted result through environments rather than recreating it casually at each stage. Record the source revision, relevant build inputs and checks. This gives an operator a more useful answer to “what is running?” than a branch name that has moved since deployment.
The exact implementation depends on the existing framework and host. A static site may promote generated files; a containerized service may promote an image; a serverless runtime may package code and configuration. Reuse the system’s normal release architecture instead of inventing a parallel publishing mechanism for an agent experiment.
Artifact identity does not establish artifact correctness. It establishes which thing received review and which thing was deployed. That connection is essential when investigating a regression, and it complements the behavior checks described in testing in the AI era.
Keep external writes repeatable and recoverable
Cloud jobs fail between steps. A worker can send a billing request, lose the connection and retry because it did not see the receipt. Without an operation identity, the second request can create a second effect.
The application should distinguish the intended business operation from an individual execution attempt. Persist its identity, prepare the payload, perform the authorized write and reconcile the remote result. A retry uses the same identity when it represents the same operation. A different customer decision receives a new identity.
Durable workflow state helps preserve progress, but external systems have their own semantics. A runtime’s replay feature does not undo a sent email or restore a changed supplier record. Idempotency explains the duplicate-effect problem; recovery design should also name compensating actions and irrecoverable consequences.
For the merchant example, generating a diagnostic report can be retried safely if its output is versioned and associated with the input snapshot. Publishing that report to a customer requires a separate delivery identity and authority. Changing the merchant’s store prices would require a different, more consequential operation. One convenient “run” button should not erase these distinctions.
Observability should follow the customer job
A service dashboard reporting healthy machines does not prove that a customer received a useful result. The factory needs evidence across the job: request accepted, input validated, operation started, external calls completed, result stored and customer notified where appropriate.
Use a correlation identifier to connect those events. Record timestamps and states that help explain delays and retries. Keep sensitive content out of general logs where it is unnecessary. An operator should be able to investigate a failed job without copying the customer’s whole document into a diagnostic channel.
Measure outcomes that correspond to the product promise. For a quote comparison, that might include accepted input, completed comparison and explicit unresolved fields. A fast model response that silently omitted a row is not a successful result. A retry that eventually succeeds may still create an unacceptable delay or cost.
Logs should have a retained purpose and an owner. Unlimited logging can increase both expense and privacy exposure. Too little logging can make a customer dispute impossible to investigate. Decide what evidence is needed, how long it remains useful and who can access it before a crisis supplies the motivation.
Calculate the cost of the operating path
Cloud compute makes resources available without buying a powerful workstation. It does not remove their cost. A factory has development workspace usage, automated checks, preview services, production runtime, storage, model calls, external tools and human exception work.
A hypothetical monthly budget might assign $40 to development environments, $30 to checks and preview services, $80 to production infrastructure and $150 to model and tool usage. Those assumed amounts total $300 before support and acquisition. They are arithmetic examples, not current vendor prices or observed Salars spending.
If three agents run redundant full builds throughout the day, the check cost may rise without increasing accepted output. If a preview remains active after a task ends, it can continue consuming resources. If a model retries an impossible task twenty times, the bill grows while the result remains unavailable. Cost controls belong in task dispatch and lifecycle cleanup.
Separate development expense from customer delivery expense. A one-time implementation experiment is different from inference required for every paid report. The first influences the investment decision; the second influences unit economics and price. AI software COGS supplies the customer-cost ledger, and the factory should retain enough usage evidence to populate it.
Treat budgets as controls with failure behavior
A budget alert informs someone that a limit is near. A hard stop changes what the system can do. The two are useful for different purposes and should not share an ambiguous label.
For an agent task, define a maximum number of attempts, a resource ceiling and the action taken when it is reached. The worker can return a partial artifact with a clear unresolved state. It should not invent a success because the budget is exhausted or continue indefinitely because it interprets the goal as urgent.
For production, a hard stop can deny customers a promised service. Design graceful degradation where possible: queue work, disable an expensive optional feature or route an exception to a person. The customer should see an honest state rather than a fabricated result produced to avoid an error page.
Budget policies should consider the whole job. A per-call cap does not prevent thousands of small calls. A per-agent limit does not prevent many agents from creating a large combined bill. The coordinator needs an aggregate view that matches the project commitment, while the runtime needs controls matching customer workload.
Plan for the network and provider to fail
A cloud-only workflow can lose access because of a network outage, identity problem, exhausted budget or provider disruption. The operator needs a recovery plan that distinguishes “cannot edit” from “customers cannot use the product.” Those failures may happen independently.
Keep the repository recoverable through its normal backup or replication arrangements. Retain environment definitions and data export procedures. Know which production obligations continue during an outage: billing, scheduled jobs, customer support and data retention do not disappear because the development interface is unavailable.
A recovery rehearsal can start small. Recreate a development environment from a fresh authorized session. Confirm that an operator can find the latest accepted artifact and the rollback procedure. Restore a safe test dataset and inspect the result. A written backup policy provides little confidence if nobody can use the backup format.
Provider diversity can reduce some dependencies while adding integration work. A small app does not necessarily need multiple active runtimes. It may need a documented export format, a recoverable repository and a realistic migration route. Choose resilience against actual consequences and tolerable downtime, rather than constructing an expensive architecture around every imaginable failure.
Migration and rollback need data boundaries
Rolling back code can restore an earlier implementation. It may not restore the data structure that implementation expects. A release that removes a field or changes its meaning can leave the previous version unable to read new records.
Use compatible transitions where practical. Add a new representation, populate it, verify readers and only later retire the old one. This can permit an older application version to operate during a deployment reversal. The exact technique depends on the data model and migration tooling, and it needs tests against representative safe records.
Some external changes require compensation rather than reversal. A notification cannot be unsent. A subscription change may need another authorized change. A published explanation may need correction with a retained history. The factory should say which recovery action applies to each operation rather than attach “rollback supported” to the product as a general comfort phrase.
The dedicated rollback chapter develops those mechanisms. In the overall cloud workflow, the important requirement is that release planning includes the data and customer obligations that survive a code change.
Cloud development can be the wrong fit
A workload requiring specialized local hardware, a disconnected environment or an operating system unavailable in the selected cloud workspace may need another route. Some organizations impose data residency, access or procurement conditions that the proposed stack cannot satisfy. A browser interface can also be inaccessible or inconvenient for a particular operator.
These are requirements to investigate, not signs that the operator lacks ambition. The useful question is whether the cloud route improves the complete delivery process within the actual constraints. If it adds latency, cost and access complexity without solving a meaningful problem, a local or hybrid environment can be more appropriate.
Avoid converting a preference into a technical claim. “We want no required local installation” is a product or operating choice. “This platform supports our workload” requires evidence. “This stack is cheaper” requires a comparable cost boundary. “We can recover within a day” requires a rehearsal or other credible basis.
A proposed pilot should retain those claims separately. The result may support cloud development for one small web app without supporting every future product. That limited conclusion is useful because it identifies where the operating method has evidence and where a new workload requires revalidation.
Design a fresh-environment acceptance exercise
The first pilot can test portability directly. Ask an authorized reviewer who did not prepare the workspace to create a fresh environment from the named repository state. Give that person the same written setup instructions a future operator would receive. Record which undocumented facts are needed to finish the task.
The exercise should include more than opening the editor. Install dependencies, run the baseline checks, create a small harmless change, build the application and inspect a preview. Confirm that fixtures are available without access to customer production data. Verify that the preview identifies itself clearly and cannot submit production writes.
Suppose the reviewer can run the application but cannot execute one integration test because a credential’s purpose is undocumented. That is a useful failure. It means the setup contract is incomplete. The response is to document the credential scope and provide an approved test route, rather than copy an administrator secret into the environment to make the exercise pass.
The success criterion should be independent of the person who designed the setup. A proposed criterion could require that the reviewer complete the named safe workflow from a fresh environment, using retained instructions, within an agreed operating budget. The threshold would be a pilot choice, not an industry standard. Retain the failures and changes so a later run can establish whether the fix solved the actual gap.
Account for preview drift
A preview can differ from production in ways that invalidate its apparent success. It may use a test database with tiny records, a permissive access policy, a stubbed external service or a missing scheduled job. Those differences can be appropriate, but they must be known.
Keep a short environment comparison beside release evidence. Identify which configuration values change, which services are substituted and which production behaviors cannot be exercised safely in preview. The reviewer can then choose targeted live verification after release rather than infer equivalence from a attractive screenshot.
For the currency-label change, a preview can prove that the interface rejects the provided ambiguous fixture. It cannot prove that every real supplier format has been covered. Production monitoring may reveal a previously unseen format, which should become a retained regression case after sensitive details are removed. Preview success and coverage completeness remain different claims.
Drift also arises over time. A preview created early in a task can remain on an old revision while the repository advances. The review surface should identify the revision or artifact it displays. If a change enters after approval, the release process should establish which checks and approvals need renewal. Otherwise the person approves one behavior and receives another.
Keep the human work legible
A cloud workflow often makes machine steps conspicuous and human decisions invisible. Dashboards show build duration and deployment status. They may not show the hour spent resolving a customer-policy ambiguity or the support obligation created by a new feature.
Record consequential decisions in a compact form: the question, chosen behavior, evidence, responsible owner and conditions that would change it. This can live in the issue, a design note or another existing repository convention. The point is to preserve the decision where future workers can find it, without building a second bureaucracy.
An agent should be able to distinguish a settled rule from a suggestion. If a document lists several pricing options, it should not interpret the first option as the live policy. If a review requests a future enhancement, it should not expand the current release silently. Explicit decision state helps both human and machine contributors stay within the accepted scope.
Human attention belongs in the cost record too. A task that uses little compute but requires repeated clarification can be expensive for a sole operator. Track interruptions, review time and exception handling when comparing workflows. These measurements can reveal that a better specification creates more value than adding another model or deployment service.
Leave a useful artifact when a task stops
The factory should have an honest incomplete state. A coding task may stop because a dependency is unavailable, a budget is reached or the specification exposes an unresolved customer rule. The result can still contain useful evidence: a patch, reproduction fixture, tested portion and a precise remaining question.
Require the handoff to identify what is safe to retain and what has not been accepted. A partially completed migration should never be confused with a deployable release because the agent attached a confident summary. A preview URL should not be described as production. A passing unit test should not be reported as a live integration check.
The next operator should be able to resume from that state without rereading the whole conversation. Preserve the exact baseline, artifact locations, checks performed and unresolved conditions. If the task resumes under a different owner, revoke or narrow the earlier worker’s authority so two execution paths do not compete.
This continuity is part of portability. Moving the environment to the cloud solves little if the only useful state remains in an interrupted chat. A retained, bounded handoff lets the factory carry work across people, sessions and service failures while keeping the acceptance decision visible.
What Would We Do at Salars?
A proposed Forge pilot would use one repository, one reproducible development definition, one app template and the existing authorized release path. It would begin with safe fixtures and a narrow merchant diagnostic. The first milestone would be an accepted change that another authorized operator could reproduce from a fresh environment.
GitHub could hold specifications, patches, checks and release records. A Cloudflare runtime could be evaluated against the actual request and workflow requirements. The proposal would not assume that every Cloudflare feature has the same availability or maturity; the service-specific documentation and account configuration would determine what can be used.
The pilot would collect accepted completion time, resource spending, failed handoffs, recovery time and customer-visible correctness. It would compare the cloud route with a clear baseline. Any benefits would be reported within that scope rather than converted into an unsupported claim about running dozens of profitable apps.
Supplier Margin Guard and Merchant Revenue Guard remain proposed product examples here. They would need customer evidence, pricing tests and ongoing support ownership before becoming commercial commitments. An available runtime establishes a place to execute software, not a market for the result.
A dependable factory should make the next change easier to understand and recover. The meaningful milestone is a customer obligation met by an identifiable release, with evidence that survives the machine and conversation that produced it.
Sources
- GitHub: What are Codespaces?, read October 7, 2026, for hosted environment configuration, client access and Linux container scope.
- GitHub: pull request reference, read October 7, 2026, for review and check surfaces.
- Cloudflare: Workflows overview, checked October 7, 2026, for documented durable-step capabilities.
- Cloudflare: Workflows limits, checked October 7, 2026. Numeric quotas must be rechecked for a real design; this article does not rely on a specific concurrency allowance.
- Salars AI library, for related workflow guides. Architecture, budgets and merchant examples in this chapter are proposed or hypothetical rather than operational results.
Loading comments…