A software factory needs to know more than how to generate code. It needs to know which job deserves a build, who authorized the work, which evidence supports the product promise, what remains unresolved, and who will maintain the result. Without that control, a faster coding process can create more half-supported apps rather than a stronger business.
Salars Forge would be a proposed control plane for evidence, delivery, and maintained obligations. No implemented Forge platform, customer base, operating revenue, or measured productivity gain is established here. This article specifies a first viable design and the tests that would determine whether it deserves further investment. It belongs to the AI Software Factory series, within our AI guides, and connects the series components without assuming that naming a factory makes them operate together.
Give Forge one bounded first job
The first job would be to move one approved app experiment through a visible lifecycle with clear ownership. It would not begin by launching a marketplace, discovering every opportunity, or managing an unlimited agent workforce. A useful initial system could consist of structured records, a repository template, explicit review steps, and a small status interface. The design should earn automation by making a manual process understandable first.
Consider a proposed invoice-comparison experiment for Supplier Margin Guard. The experiment would have a buyer hypothesis, eligible document conditions, permitted inputs, an independently reviewed answer set, budget, stop date, and a read-only output promise. Forge’s job would be to connect those records to the implementation and review. Supplier Margin Guard itself remains a proposal; this scenario does not establish customer demand or an executed pilot.
Define what Forge would improve relative to the current baseline. A reasonable proposed baseline is a maintained manual checklist and repository record. The question would be whether Forge reduces lost decisions, stale evidence, or unclear handoffs without creating excessive administration. Counting generated repositories would miss that purpose. A factory can create no new apps during a period and still improve a consequential maintenance decision.
Keep the first scope small enough to evaluate. One candidate, one template family, one reviewer path, and one deployment target could expose the core coordination problems. Expansion to unrelated app types would require evidence that their rules and obligations fit the system. The control plane should not impose a universal process simply because a database schema allows another row.
Separate records, interpretations, proposals, and actions
A factual record might say an official API supports a documented capability as of a checked date. An interpretation might say that capability could support the buyer’s job. A proposal might request a compatibility test. An action might run that authorized test or apply a reviewed change. Forge would preserve those distinctions so a model’s plausible summary cannot silently become an approved deployment.
The existing website as an AI business system guide uses a similar distinction for maintained information and approved business actions. Forge would apply it to software delivery. Public article information, private customer material, product decisions, and executable actions need appropriate boundaries. A common interface would not make their rights or authority interchangeable.
Every consequential proposal would identify the evidence it uses, the action it requests, the owner, scope, budget, and completion check. An approval would apply to that proposal version. If the implementation changes the data it collects or the external actions it performs, the system should require reconciliation with the approved scope. A decision to run a synthetic test would not authorize processing customer records.
Actions would produce an outcome record rather than merely a success badge. The record would include what was attempted, what completed, what remains unknown, and any relevant external state. If a deployment command returned successfully but the service failed its behavior checks, Forge should show the unresolved condition. Technical command success and product readiness answer different questions.
Define identities that survive a renamed project
Each candidate would receive a stable app identifier, separate from its public name, repository, route, and deployment. An experiment identifier would connect a question to a versioned plan and result. A release identifier would bind acceptance evidence and applicable authorization to the exact source revision and built artifact. Routine changes within an established authorized scope would receive proportionate verification; a materially changed scope would require reconciliation with that authority. A customer job identifier, where appropriate and permitted, would connect service events without exposing unnecessary source material throughout the control plane.
Those identifiers would prevent a rename from erasing history. A failed pilot should not appear as a fresh investment because its repository moved or its product label changed. A release should retain the rules and test conditions that justified it. A retirement should preserve residual obligations even after the public offer disappears. Stable identity supports accountability rather than merely making search easier.
The app registry guide develops the maintained inventory. Forge would use that registry as the source of lifecycle state, ownership, and dependency visibility rather than create a competing list. Records would distinguish proposal, discovery, development, pilot, maintained service, pause, retirement in progress, and retired. Exact transitions would follow the app’s actual obligations.
Avoid putting every detail in one status field. An app could be maintained while a new release is blocked. A pilot could be paused for data access while its buyer hypothesis remains interesting. A retired app could still have a records obligation. Separate the product lifecycle, experiment state, release state, and unresolved operational items so the interface can describe those conditions honestly.
Make the opportunity handoff reviewable
Opportunity Radar, itself a proposed system, would eventually provide candidate evidence. Forge would accept a bounded opportunity brief rather than a raw popularity score. The brief would identify a buyer job, current alternative, pain evidence, rights and access questions, likely delivery constraints, source dates, counterexamples, and the next experiment. An issue count or a model-generated market estimate would not be sufficient authorization.
The receiving reviewer would decide whether the proposed next step fits current resources and service obligations. They could approve discovery, request evidence, narrow the scope, or decline the candidate. Approval would not mean the business opportunity has been proved. It would mean a specified next activity is justified within its budget and permission boundary.
In the Supplier Margin Guard scenario, the first approved activity might be inspecting independently created quote and invoice fixtures and preparing interview questions. It might later become a permitted concierge diagnosis. These activities produce different evidence. Forge would record what was actually performed and which uncertainty remains, rather than allow the candidate to inherit a vague validation-complete label.
A handoff would include acceptance criteria for the receiving stage. The developer should know the supported input family, required refusal behavior, external-action boundary, and questions that require a domain reviewer. A beautiful brief that omits those conditions leaves the implementation team to invent the product promise. Forge’s first benefit would be making that invention visible and preventable.
Build a template with explicit contracts
A first app template would include a minimal project structure, configuration validation, identity and entitlement boundaries where required, input validation, diagnostic states, test fixtures, deployment instructions, and an operating record. It would not include a broad set of unused integrations merely to look complete. Each component should exist because the supported app family needs it.
The app template guide examines reuse at this level. GitHub’s official repository template documentation explains that templates copy structures and files, while template-created repositories start with a different history structure from forks. A copied template is a starting point, not an automatic future update channel. Forge would need an explicit process to identify affected apps when a shared component changes.
Define the template contract in plain terms. For the read-only invoice experiment, the output would preserve the source reference, show the comparison rule, distinguish supported mismatches from unresolved cases, and require human interpretation before an external action. A component that produces a number without those conditions would not satisfy the contract merely because its tests passed.
Keep product-specific rules outside generic infrastructure. Supplier invoice interpretation differs from merchant refund reconciliation. They may share upload handling and audit mechanisms while needing different domain answers, permissions, and recovery practices. Template reuse would be reviewed against the new job instead of inheriting confidence from an unrelated app.
Give agents bounded roles and outputs
Agents could help research, draft specifications, implement code, propose tests, and inspect changes. Their roles would be bounded by the stage and evidence available. A research agent would return source-specific claims and limitations. An implementation agent would work from an approved contract. A review agent would inspect behavior and risk without automatically becoming the authority that ships the result.
The multi-agent software development guide addresses task decomposition. The official OpenAI multi-agent documentation describes manager-style agents-as-tools, handoffs, and code-controlled orchestration. Those patterns offer implementation choices; they do not establish that more agents improve this factory or that model behavior becomes deterministic.
A proposed Forge workflow could use code to sequence gates and allow parallel work only when tasks are sufficiently independent. Researching a source and inspecting an existing interface might proceed concurrently. Editing the same rule and its dependent acceptance criteria would require coordination. The system would identify ownership of files and artifacts to reduce conflicting work rather than rely on agents to infer shared boundaries.
Every role would return a structured result with status, evidence, changes, uncertainties, and next required owner. A missing source or unresolved test would remain a blocker or an explicit limitation. An agent should not complete a record by inventing a result. Forge would favor honest partial evidence over a polished handoff that conceals what was not verified.
Prevent review from becoming self-approval
The agent that implements a change can help explain and test it, but it may reproduce its own assumptions in the review. A second agent can share the same error if both rely on the same flawed brief. Independence is not established by opening another model session. The review process needs distinct evidence, protected cases, and a qualified human decision where consequences require it.
For the invoice experiment, an independent answer set would define legitimate price changes, ambiguous line matches, missing information, and confirmed discrepancies. Review would include cases the implementer did not use to construct the rule. A system that flags every line would fail the buyer’s task even if it finds all known discrepancies. Review effort and unsupported conclusions matter alongside detection.
The OpenAI evaluation best practices recommend task-specific evaluation and calibration with human judgment. Forge would use that guidance to design evidence, not import illustrative thresholds as universal release criteria. Model judges can have biases and shared errors; consequential domain labels need an appropriate independent basis.
A release proposal would identify failures, exclusions, and uncertainty. The reviewer could accept a narrow scope, require another test, or reject the change. A code review approval would not automatically establish buyer demand, privacy readiness, or economic viability. Forge would track those decisions separately so one green check cannot stand in for every form of readiness.
Control tools by the action they can perform
Reading public documentation, editing a local draft, deploying a service, sending a message, charging a customer, and changing an external record have different consequences. Forge would assign tools and permissions accordingly. An agent working on a synthetic parser test would not need live billing credentials or the ability to contact customers. Capability would be limited to the approved task.
External material would be treated as evidence, not authority. A repository issue, web page, document, or customer file can contain instructions directed at a model. Those instructions should not override the approved workflow. OWASP’s prompt injection guidance describes indirect injection and emphasizes controls such as least privilege and human involvement for high-risk actions. It does not offer a prompt that eliminates the risk universally.
Require a concrete action proposal before a consequential write: target, intended change, evidence, expected effect, limits, and recovery path. The approver should see the result they are authorizing rather than an abstract request to let the agent continue. User authorization already established for a bounded task should be preserved; the system should not ask repeatedly for routine reversible steps inside that scope.
Secrets would remain in appropriate controlled configuration, with access separated by environment and role. Logs would avoid exposing credentials or unnecessary customer material. A permission policy must be tested through denied actions as well as allowed ones. The proposed design would include cases where a malicious document requests an unrelated write and the agent cannot perform it.
Choose runtime after defining recovery
Forge’s first workflow could run as a simple application with durable records and explicit jobs. More advanced durable execution would be considered if retries, waits, and external events create a real coordination burden. Architecture should follow the required behavior rather than a desire to use every available agent platform.
Cloudflare’s Agents documentation describes durable identities, state, scheduling, and recoverable execution. Its feature documentation distinguishes maturity at the feature level; the Models and Pi areas are explicitly labeled beta in the reviewed navigation. Those capabilities are candidates, not evidence that a Salars deployment can meet its actual requirements. Account access, current compatibility, limits, and operating costs would need verification before implementation.
A job would have a known state: proposed, authorized, running, waiting, completed, failed, or uncertain external outcome. If a process crashes after an external request, recovery should determine what happened before retrying a consequential action. Durable orchestration does not by itself make every outside operation exactly once. The action adapter needs appropriate deduplication, reconciliation, and ownership.
For a harmless source fetch, retrying might be straightforward. For a deployment or billing change, a retry can have a material effect. Forge would define idempotency boundaries and a manual resolution path where the external state cannot be established automatically. A failed orchestration job should not leave the operator believing that no external action occurred.
Keep evidence current without creating an endless research task
A source record would contain URL, checked date, exact supported claim, relevant conditions, and revalidation trigger. The system would distinguish a source passage actually read from a search result or an inferred summary. If a product depends on a current vendor capability, the claim would be rechecked before the implementation decision that relies on it.
Not every record needs continuous polling. A stable method explanation can remain useful until its conditions change. A price, platform requirement, legal obligation, or beta feature can need timely verification. Assign revalidation based on the consequence and volatility of the claim rather than impose one arbitrary expiry period on the entire factory.
When a claim changes, identify dependent decisions. A vendor restriction might affect a candidate’s feasibility, a deployed integration, a public promise, and a support procedure. Forge would produce a review queue for those items. It should not automatically rewrite the product promise or terminate service without the appropriate decision and obligation assessment.
Keep research cost visible. A candidate can consume many hours of browsing without reducing the uncertainty that determines investment. The reviewer should ask what next observation would change the decision and whether it can be obtained within the remaining budget. Evidence discipline includes stopping when further collection no longer supports a useful choice.
Track operating obligations after launch
Forge would not mark an app complete at deployment. A maintained service has support, billing, dependencies, data handling, and response commitments. The registry would connect those obligations to owners and current queues. A release can succeed technically while the operator lacks the capacity to serve the new customers it attracts.
The first proposed maintained-app review would inspect task outcomes, support effort, recurring costs, unresolved incidents, permissions, and dependency changes. Economic records would distinguish cash, contribution, and owner capacity. A provider’s recurring-revenue metric would not become proof of profit or customer success merely because Forge can display it.
A pause would stop specified intake or actions while existing obligations are handled. Retirement would remain in progress until customer resolution, billing, export, data, infrastructure, and residual contact checks are complete. The control plane would make those states legible rather than archive a repository and assume the business commitment disappeared.
The factory should support fewer launches when maintenance evidence calls for it. A dashboard that rewards only completed builds would distort the goal. Maintained usefulness, resolved uncertainty, and fulfilled promises are better decision objects. Quantifying them would require a defined method and actual observations; this proposal supplies no measured Forge score.
Design the interface around decisions that need attention
A first dashboard would show the next decision, its owner, the evidence available, the deadline or trigger, and the consequence of delay. It would not begin with a wall of agent activity. A founder needs to know that a release lacks a protected test or that a customer export remains unresolved, not merely that twenty tasks ran today.
Each app page would show its current promise, supported scope, lifecycle state, active experiments, latest accepted release, operating obligations, and relevant dependencies. Evidence would be reachable from the decision it supports. Unknown values would appear as unknown, with a proposed next action where appropriate. Empty fields would not be filled with model estimates to make the page look finished.
An action review would show the exact target and change, including differences from the approved proposal. The interface would explain terms a nontechnical operator needs to decide. It could summarize implementation details while retaining access to the underlying artifact for appropriate review. Product language should follow the user’s task rather than expose internal orchestration jargon unnecessarily.
Notifications would be tied to meaningful changes: required decision, failure, completion, or an obligation at risk. Repeated unchanged status messages can obscure the item that needs attention. The owner would be able to inspect progress without being asked to approve the same scope at every step. The aim is calm control over actual commitments.
Measure the control plane against a manual baseline
A proposed evaluation would compare the maintained manual process with Forge on protected scenarios. Include an ordinary app handoff, missing source evidence, conflicting edits, a failed deployment check, a changed permission boundary, an uncertain external result, and retirement with a pending billing item. The criteria would assess whether the right owner receives a clear decision and whether unsupported actions are prevented.
Measure administrative time, decision completeness, stale records found, unresolved obligations, and errors introduced by automation. Do not count faster coding as the sole benefit. Forge could reduce handoff confusion while taking longer to prepare a release proposal; that tradeoff might still be useful. It could also create a polished interface without improving decisions, which should weaken the investment case.
Protect evaluation cases from the implementation team where feasible, establish answers independently, and retain counterexamples with scope. A simulated workflow test would demonstrate behavior under those conditions, not real-world business profitability. A later bounded pilot would need appropriate participants, permissions, and operating coverage. Label proposed, simulated, and observed evidence separately.
Set a stop condition for the control plane itself. If its maintenance and administrative burden exceeds the decision value it provides, simplify it or return to the manual baseline. Forge should earn the next increment just as an app should. A factory is not exempt from capital discipline because it coordinates other projects.
Walk one failure through the proposed system
Suppose an implementation agent adds automatic invoice updates to the read-only Supplier Margin Guard experiment. The code might be technically competent, but the behavior exceeds the approved contract. Forge would compare the release proposal with the authorized scope and route the discrepancy to the owner. It would not reinterpret the original approval as permission merely because the agent describes the new behavior as useful.
The owner could reject the action, require removal, or request a separate proposal with appropriate evidence, permissions, and recovery design. If removal is chosen, the reviewer would verify that scheduled jobs, credentials, and external adapters cannot still perform the write. Deleting the button alone would not establish that the action is impossible. The test would include an attempted write through a permitted evaluation scenario and inspect the resulting denied state.
Now suppose a deployment job fails after sending its request. Forge would retain an uncertain outcome until the actual deployment state is reconciled. A second attempt might be safe under an appropriate identifier and adapter contract; otherwise the job would require investigation. The system would show the difference between no request sent, request rejected, and effect unknown. Those states determine what recovery is justified.
Finally, imagine the revised read-only release passes its protected cases but requires unexpectedly extensive domain review during the pilot. Forge would record the assistance and operating cost, preserving the distinction between technically accepted output and a viable service. The owner might narrow eligible inputs or stop the pilot. A green release check would remain in the history, while the product decision could still be negative. That combination is precisely why release and business states need separate identities.
This walkthrough is a proposed evaluation case, not a report of an actual Forge incident. Its value would be established by running it against the manual baseline and proposed implementation under independently defined criteria. If the manual checklist handles the case more clearly and cheaply, Forge should improve or simplify rather than receive credit for the existence of a dashboard.
Budget the first increment honestly
Suppose an illustrative first implementation requires twenty hours for records and interface, twelve for template and handoff integration, eight for evaluation fixtures, and ten for review and repair. That is fifty hours before any additional operating pilot. At an illustrative internal resource value of $50 an hour, the capacity commitment is $2,500, plus direct expenses that would need separate verification.
If the owner has five discretionary hours weekly, those fifty hours represent ten weeks under the simplified schedule. The business could narrow scope, obtain suitable assistance, or choose a faster manual improvement. It should not publish a one-week factory launch based solely on the speed of code generation. Reviewing the control mechanism and testing failure conditions are part of the work.
The possible benefit might be fewer lost handoffs, faster diagnosis, or clearer obligations. Assigning a precise dollar return before measurement would be speculative. Define which observations could support further investment and which would stop it. A local reduction in administrative work can be useful without proving a general productivity multiplier.
Keep current customer service funded during development. Forge work competes with apps and other business tasks for the same resources. The first version should solve a recurring coordination problem already visible in the approved workflow, rather than create a new platform simply because a broad architecture diagram looks compelling.
What Would We Do at Salars?
We would propose a records-first Forge pilot for one bounded candidate, with customer outreach, spending, deployment, and external financial actions controlled by their established authorization. We would begin with a manual baseline, explicit ownership, a minimal template, and reviewable artifacts. No Forge implementation or pilot result is claimed by this design.
The first candidate would carry a buyer hypothesis, allowed input conditions, independent answer cases, budget, stop date, and maintenance estimate. Agents could prepare research, code, tests, and review proposals inside assigned boundaries. A responsible owner would decide whether the evidence justifies the next stage. Another agent’s confidence would not replace that evidence or the necessary approval for consequential action.
We would evaluate missing evidence, malicious source instructions, uncertain external effects, shared template changes, and retirement obligations alongside the ordinary path. A favorable simulated result would support only the tested coordination behavior. A later operating pilot would need its own success criteria and service coverage. Supplier Margin Guard, Opportunity Radar, and other named components remain proposals throughout.
We would expand Forge only when the first increment demonstrably improves decisions enough to justify its operating burden. The result could be a small internal tool rather than a broad platform. Its success would be a clearer path from justified opportunity to supported product, with fewer unowned commitments. The factory would be valuable because it helps the business keep useful promises, not because it produces an impressive number of apps.
Sources
- GitHub repository templates: copied project structures and history distinctions; future updates need a maintained process.
- OpenAI multi-agent orchestration: manager, handoff, and code-orchestration patterns, not demonstrated Forge productivity.
- OpenAI evaluation best practices: task-specific tests and human calibration.
- OWASP prompt injection guidance: indirect injection and layered controls, not universal prevention.
- Cloudflare Agents: runtime capabilities and feature-level maturity; actual account suitability remains to be checked.
Sources checked October 7, 2026. Forge architecture, workflows, budgets, and evaluations are proposals. No deployed platform, executed experiment, customer outcome, productivity multiplier, or business return is asserted.
Loading comments…