A software factory can generate three prototypes in a week and still leave the business worse off. None may have a reachable buyer. One may depend on an unavailable integration. Another may create a support obligation the owner cannot maintain. Speed has increased the number of unresolved decisions rather than the number of useful assets.
The venture engine proposed in this final chapter has a different unit of progress: a justified next commitment, followed by an observed result and a maintained obligation. Connect opportunity evidence, delivery, distribution and economics through a lifecycle that can continue, narrow, pause or stop. An app earns the next stage; it does not inherit it because the previous task finished.
This is a synthesis of The AI Software Factory, not a claim that Salars already operates the system. Salars Forge, Opportunity Radar, Supplier Margin Guard, Merchant Revenue Guard and the software store remain proposals. No customer base, deployment, revenue, benchmark or product-market fit for them is established here.
Define the engine’s output as a maintained asset
A maintained asset is a supported capability connected to a useful customer job, current evidence and an owner who can fulfill its obligations. It may be an internal tool, a paid service, a commercial app or retained knowledge that improves another workflow.
The definition deliberately allows a useful stop. An experiment that rejects an expensive idea can preserve resources and produce evidence worth retaining. It creates a decision improvement even though no public app launches.
A prototype can be part of the path without being the final output. Its code, demo and technical checks answer particular questions. They do not establish outside demand, sustainable delivery or a repeatable acquisition route.
The engine should record what each stage actually produced. A source claim, compatibility finding, paid pilot, accepted release and maintained customer result are different artifacts. Keeping them distinct prevents a persuasive narrative from converting one kind of success into every kind of readiness.
Begin with the current business obligations
Before evaluating new ideas, identify the service and operating commitments already accepted. Existing customers, support, maintenance, billing, data handling and transition work consume resources regardless of how attractive the next proposal looks.
Define the available cash and owner attention for a stated period. Separate expected receipts from resources actually available, and committed funds from discretionary exposure. The same remaining budget cannot be promised independently to several agent workstreams.
A proposed engine review would protect existing obligations first, then assign a bounded learning budget. That budget is not a universal percentage of revenue; it depends on the actual business, uncertainty and capacity.
If no resources remain for a new test, the next useful action may be reducing a recurring exception or resolving an old obligation. A system that only celebrates new candidates would misread that work as stagnation even when it improves the business materially.
Give Opportunity Radar a constrained intake job
Opportunity Radar would collect candidate evidence, not authorize builds. A useful record would name the buyer’s job, current alternative, source, observed pain, access questions and the uncertainty a next test could resolve.
A public issue can be a signal, but its meaning needs interpretation. GitHub describes issues as records for ideas, feedback, tasks and bugs. An issue count is therefore not a paying-customer count or proof of a commercial opportunity.
The intake record should preserve counterexamples. A problem may be solved adequately by an existing tool, matter only to a few maintainers or require rights the proposed app cannot obtain. Evidence against the candidate belongs beside evidence for it.
Set an intake limit appropriate to the decision capacity. Gathering hundreds of ideas can create a research backlog without changing the next commitment. The Radar proposal becomes useful when it produces a small number of reviewable questions, not merely a larger list.
Separate eligibility from ranking
Some conditions determine whether a candidate can proceed at all: lawful and permitted data access, a safe action boundary, feasible delivery and an available owner. Other dimensions help choose among eligible candidates.
The software opportunity score develops ranking with uncertainty and sensitivity. The engine would use it for the stated decision, such as which investigation to fund next, rather than as an automatic predictor of startup success.
An attractive score cannot waive a missing permission or turn a beta capability into a verified production dependency. Unknown values should remain visible and influence the next test. A model should not fill them with plausible estimates to complete the table.
Record why a candidate is deferred or declined. A missing buyer route may justify a distribution test. A prohibited data use may end the candidate. A maintainer absence may require a capacity change. These reasons create different conditions for reconsideration.
Translate the candidate into a falsifiable question
A broad proposition such as “merchants need AI margin protection” is too flexible to guide a bounded experiment. Specify the customer, job, conditions and expected useful result.
For a proposed Supplier Margin Guard investigation, the question might concern whether intended merchants can supply eligible quote and invoice records with identifiers sufficient for a reviewable comparison. That is a compatibility question. It leaves willingness to pay and service economics unresolved.
Write the baseline and independent success criteria. The baseline could be the merchant’s existing manual comparison. Criteria could address correct matches, unresolved cases and review burden, established before implementation.
Include counterexamples and stopping conditions. Missing identifiers, conflicting records and a legitimate price change should not be forced into the same discrepancy label. The proposed test should identify what would narrow or reject the hypothesis, not only what would allow the team to announce success.
Fund the next observation rather than the entire story
A candidate can require interviews, a fixture review, a concierge job or a technical compatibility check before a build is justified. Choose the smallest activity that can change the next decision.
State its owner, permitted inputs, budget, time boundary and result artifact. A vague instruction to “validate the idea” leaves the worker to decide how much research, coding and external action it authorizes.
Funding one stage should not imply funding the next. A negative result can complete the approved task successfully and end the candidate. A positive result supports a specified continuation, not an unlimited product roadmap.
This staged commitment also preserves option value without inventing a precise financial return. The business pays a bounded amount to reduce an uncertainty and retains the ability to stop. The benefit depends on whether the observation actually changes a consequential choice.
Preserve the difference between discovery and sales
A customer interview can reveal language, workflow and alternatives. A paid pilot can reveal a consequential commitment and delivery behavior. Neither should silently stand in for the other.
Record what the participant did, what assistance was supplied and what conditions applied. A friendly person agreeing that a tool sounds useful does not establish purchase demand. A paid custom service does not establish unassisted SaaS readiness.
The engine would connect each observation to its supported claim. It could say that one consented workflow contains compatible fields, that one buyer accepted a bounded service or that a particular output met an independently defined criterion. It should avoid a single “validated” badge covering all of them.
This discipline makes positive evidence more useful. A narrow, honest result can justify a clear next experiment. A broad label can conceal which commercial and technical questions still need work.
Make Forge a control plane for commitments
Building Salars Forge specifies the proposed control plane. Its role in the engine would be to connect identities, evidence, authorized work, release records and operating obligations.
A candidate would keep a stable identity through renaming, repository changes and lifecycle transitions. Its current name would not erase failed pilots or unresolved commitments. An experiment and release would have their own identities so an app could remain maintained while a new feature test is paused.
Forge would prepare a reviewable next action, not transform a model recommendation into authority. A consequential proposal would show target, intended effect, supporting evidence, scope and recovery. Existing authorization inside a bounded task would remain valid for routine steps rather than be requested repeatedly.
The first implementation could be structured records and a maintained manual process. Automation would need to demonstrate an improvement over that baseline. A elaborate interface that hides decisions behind agent activity would miss the control plane’s purpose.
Keep agents inside stage-specific contracts
Research, implementation and review agents can work on independent bounded tasks. Each should know its inputs, owned artifacts, allowed tools, completion conditions and unresolved questions.
A research agent can return source passages and limits. It cannot establish customer demand by summarizing public complaints. An implementation agent can prepare a change from the approved specification. It cannot expand the product’s data use or release authority merely to finish the code.
A review agent should inspect evidence and protected cases, not only repeat the implementer’s explanation. Another model session may share the same assumptions, so independent evidence matters more than the number of agents.
The engine would record proposed and executed actions separately. A generated test plan is not a passed test. A successful deployment command is not a verified customer workflow. A worker should return honest uncertainty rather than complete a status field with invented evidence.
Treat the specification as a customer contract boundary
Before implementation, define the supported inputs, accepted useful result, refusal behavior, retained customer decision and external-action boundary. This specification gives the builder a target and the reviewer an independent criterion.
For a proposed read-only supplier comparison, the product might show source-linked changes and unresolved matches without contacting suppliers or editing purchasing records. The scope can be useful while leaving business interpretation with the merchant.
A new automatic action would be a different commitment, requiring authority and recovery appropriate to its consequence. It should not be introduced as a minor convenience inherited from a read-only pilot.
Keep the specification version tied to the release. If the implementation changes material data collection or acceptance behavior, reconcile the difference before release. A green test against an old promise cannot establish the readiness of a newly expanded one.
Reuse capabilities without inheriting unsupported confidence
Shared identity, validation, logging, billing and deployment mechanisms can reduce repeated work. Their contracts still need review in the new product’s context.
A parser proven on one source family may fail on another. A billing adapter can execute charges correctly while the new commercial unit remains confusing. A template’s secure default can be undermined by a product-specific integration.
Track which versions each app inherits and who owns updates. Copying a repository is a starting point, not an automatic future maintenance channel. Shared capability creates both useful reuse and correlated obligations.
The engine would retain accepted improvements with scope. A new protected test, clearer input rule or corrected action adapter can help future products. It should not become a universal standing instruction simply because one local case succeeded.
Require independent acceptance before release
Release evidence should inspect the promised customer job, not only compilation and ordinary function tests. Include supported difficult cases, refusal cases and the authority boundary.
The OpenAI evaluation guidance recommends task-specific cases and human calibration of automated judgments. The engine would use those principles without importing illustrative thresholds as universal gates or assuming a model judge is independent proof.
The review should state failures and excluded conditions. A narrow release can be justified if it meets its explicit promise and makes unsupported cases clear. A broad release with unresolved critical behavior should remain a proposal.
Evidence about accuracy, privacy, security, buyer demand and economics remains distinct. One successful check cannot substitute for the others. The release decision would show which obligation each artifact supports and which owner accepts the remaining scope.
Separate release, publication and acquisition
Deploying a service makes it available under a technical path. Publishing an offer communicates a commercial promise. Acquiring customers tests whether suitable people discover, understand and commit to that promise.
The distribution system examines the maintained route from discovery to retained value. The engine would connect channel evidence to the product’s actual acceptance and economics rather than celebrate traffic alone.
A software store could organize supported offers and explain their readiness. It would not make an unvalidated candidate ready by displaying it beside maintained products. Proposal, pilot and available service states should be legible.
Marketplace access also carries current requirements and dependencies that need verification for the actual release. A distribution option is not a guarantee of installs. The engine should preserve channel uncertainty and platform concentration when deciding how much to commit.
Measure the whole customer path
A visitor, signup, first accepted result, paid relationship and useful repeat cycle are different events. Choose definitions appropriate to the product’s cadence and preserve non-completing customers in the relevant denominator.
Record assistance. A founder-supported account can reveal valuable learning while consuming substantial delivery work. The engine should not count that path as unassisted activation or hide its human cost.
Connect customer jobs to permitted evidence without centralizing unnecessary private content. A control plane can use job identities, state and scoped metrics while the sensitive source remains in the appropriate product boundary.
When a result changes, ask what mechanism changed. Better input readiness, faster processing, clearer interpretation and a different traffic mix can each affect completion. A before-and-after total alone may not establish causation.
Build economics from accepted work and real obligations
The engine would reconcile retained revenue, resource cost, refunds, support and owner time under explicit definitions. Recurring-revenue totals remain useful but do not become cash or profit merely because they appear in a dashboard.
A cost ledger should include failures and retries that consume resources without creating billable work. An owner-time record should include maintenance and exceptions rather than only planned development hours.
The operating review should inspect cohorts and expensive tails. An average can conceal a group whose workload exceeds the plan’s obligation. A price change can improve per-account contribution while reducing total contribution if customer response differs from the scenario.
Keep formal accounting treatment separate from the operating decision model and obtain appropriate review for the actual business. The engine’s purpose is traceable resource decisions, not a universal financial certification or a promise that every maintained app will be profitable.
Allocate the next increment with current evidence
Software portfolio capital allocation compares the next dollar and hour, not only historical product totals. The engine would bring current obligations, discretionary capacity and scoped proposals into that review.
An operating app can deserve maintenance funding while losing the contest for expansion. A new candidate can deserve a small compatibility test without being declared the next winner. A stopped app can still need transition resources.
Compare the increment and its spillovers. A shared parser improvement may affect several products and require verification across them. A new app may increase dependence on the same channel or reviewer rather than diversify the business.
Funding should end in an owner, scope, budget and review trigger. A ranking meeting that does not specify the actual commitment leaves authority ambiguous. A forecast remains dated and conditional until observed delivery and customer behavior support it.
Design pause as a useful operating state
Pause can stop new intake, an uncertain external action or a candidate experiment while preserving existing obligations. Define what stops and what continues.
A proposed merchant app might pause new reports after an integration change while preserving access to accepted results and resolving queued jobs. A candidate might pause because data permission is unavailable while retaining its discovery evidence.
The state should have an owner and a condition for reconsideration. “Paused” without either can become a hiding place for an unresolved commitment. Repeated reminders without changed evidence do not improve it.
A pause should not be interpreted as proof the idea failed or permission to erase customer work. It is a scoped decision about current uncertainty and capacity. The engine’s records should make that distinction understandable to the next reviewer.
Retire obligations as carefully as repositories
A weak product may need to stop, but retirement includes customer communication, billing, exports, retained data, infrastructure and residual support. Archiving source code alone does not end the commercial relationship.
Fund transition work and identify its completion evidence. A customer export, pending refund or prepaid period can require resources after new sales stop. The money becomes available for another project only when the obligation is genuinely resolved.
Preserve useful evidence from the retired candidate: supported inputs, rejected hypotheses, observed cost mechanisms and remaining revalidation conditions. Remove or clearly label stale public promises so the stopped product is not rediscovered as a current offer.
The engine should make stopping ordinary rather than shameful. A product does not need an unlimited rescue build because the factory has already invested effort. The relevant question is the next obligation and its supported value, not the desire to justify sunk work.
Create a review cadence tied to decisions
A proposed weekly operating review could focus on customer failures, unresolved jobs, support capacity and evidence needed for current commitments. A less frequent allocation review could compare experiments, expansion proposals and retirement obligations.
The cadence should fit the actual service and team. A daily meeting is not automatically disciplined, and a monthly review may be too slow for a consequential failure. Set event-triggered reviews for changes that affect the accepted promise or authority boundary.
Every review should produce a decision or a clear next observation. A meeting that merely recites agent activity adds administration without reducing uncertainty. Record continuation, narrowing, pause, stop or a bounded additional test with its owner.
Do not reopen a deferred proposal without new evidence or a changed condition. The retained decision record should save attention, not create an endless loop of discussing the same attractive idea.
Evaluate the engine against a manual baseline
The engine itself is a proposal needing validation. Compare it with a maintained manual checklist and registry on representative scenarios before expanding automation.
Protected scenarios should include an ordinary handoff, missing source evidence, an unsupported customer input, a denied cross-account action, a failed release check, an uncertain external effect and retirement with a pending obligation. The expected result is the correct owner receiving a clear decision and the appropriate boundary holding.
Measure lost decisions, unclear handoffs, unresolved obligations and administration under stated definitions. Do not count completed agent tasks as the independent success criterion when the engine is designed to complete agent tasks.
A local exercise can support a finding about the tested workflow. It does not establish broad venture returns or production resilience. Preserve counterexamples, test conditions and a revalidation trigger when the system’s scope changes.
Keep public knowledge and private action separate
The article series can explain methods, proposals and evidence limits. A live control plane may handle private customer records, credentials and authorized actions. Their shared subject matter does not give them interchangeable permissions.
A public page should not expose private pilot material or imply deployment that has not occurred. A private agent should not interpret a published proposal as authority to contact customers, spend money or change accounts.
When evidence becomes suitable for publication, review its scope, permission and representation. A hypothetical example remains hypothetical. A measured local result should state the conditions and avoid claiming general causal validity beyond its design.
This boundary also protects the user’s trust in the work. Readers should be able to tell what exists, what was tested and what is proposed. Naming a venture engine can make a useful design coherent; it cannot supply the missing operating evidence.
Follow a candidate through a disagreement
Consider a hypothetical candidate whose compatibility test passes but whose prospective buyers do not commit to the proposed result. The technical worker has completed its task successfully. The commercial hypothesis remains unsupported. The engine should show both findings rather than choose one broad success label.
The next decision might be to narrow the customer segment, test a different useful result or stop. It should not automatically fund a more elaborate interface on the theory that enough features will convert technical feasibility into demand. That is a new hypothesis requiring its own bounded evidence.
Now consider the reverse case: a buyer accepts a pilot, but the supported source lacks the identifiers needed for a reliable comparison. Payment demonstrates a commercial commitment under its conditions; it does not repair the technical boundary. The provider may need a narrower input requirement, manual service or a refund under the actual arrangement.
These disagreements are normal reasons for a lifecycle to exist. They show why evidence, ownership and readiness need separate records. A well-designed engine helps the owner see which next commitment is justified without forcing every result into a comforting story.
Resolve uncertain external effects before repeating work
A proposed deployment or billing operation can finish externally while its acknowledgment fails. The engine then needs an uncertain state rather than a simple failed badge. Repeating the action automatically can create a duplicate effect or a misleading record.
The action record should preserve its target, operation identity, approved version and any receipt evidence. The relevant adapter or owner can reconcile the external state before deciding whether a retry is safe. If it cannot be established, the system should expose the uncertainty and a bounded resolution path.
This is an engine-level handoff question, not a claim that durable orchestration guarantees exactly-once behavior everywhere. The detailed component mechanism belongs to the action adapter and its tests. The lifecycle’s obligation is to prevent an uncertain result from disappearing into the next stage.
A proposed evaluation would include a completed external action with a lost acknowledgment. Success would mean the next reviewer receives a correct uncertain-state record and cannot accidentally authorize an unrelated repeated action from a stale proposal.
Revalidate the engine when its authority expands
A control plane that only prepares research records has a different consequence boundary from one that can deploy, charge or publish. Adding such a capability should trigger review of tool scope, approval representation, recovery and evidence retention.
The original manual-baseline evaluation may remain informative, but it does not establish the expanded system’s safety or usefulness. Add protected scenarios for the new action and inspect the actual enforcement outside the model.
Keep the revalidation proportional. A cosmetic display change may need only a straightforward check. A new external write needs evidence about the authority and effect it creates. The engine should help the owner recognize that difference rather than require the same ritual for every edit or treat every change as already covered.
What Would We Do at Salars?
We would begin with one proposed customer job and a maintained manual decision record. Supplier Margin Guard could be a candidate if independent discovery established a relevant buyer, permitted eligible inputs and a useful comparison result. The name alone would not select it.
Opportunity Radar would prepare evidence and counterexamples. Forge would connect the approved bounded test to its specification, artifacts, review and resource record. The responsible owner would decide whether the result justified an assisted pilot, a narrow product, internal use or a stop.
Before a commercial release, we would review the internal-tool-to-SaaS readiness delta: outside customer validation, tenancy, onboarding, terms, metering, support, exit and maintenance. We would test the relevant ordinary and denied cases under a defined first-release scope.
A proposed operating review would then inspect accepted outcomes, recurrence where appropriate, contribution, owner time and unresolved obligations. A growth decision would require evidence about the next increment and a maintained buyer route, not only a successful prototype or a rising revenue metric.
All of this remains a proposed Salars design. No execution of the engine, customer experiment or portfolio result is claimed. Its first success criterion would be better governed decisions than the manual baseline, with acceptable administration and preserved boundaries.
Make retained learning earn the next cycle
The broader AI capital flywheel describes how accepted work can improve retained capabilities. The venture engine gives that loop an operating lifecycle: observe, interpret, propose, authorize, execute, verify and decide the next commitment.
The retained improvement might be a clearer offer, a validated parser rule, a protected test or a better exit procedure. It should change future work under a stated scope. A growing archive of generated output does not establish compounding.
The engine can also produce a negative loop if it promotes every exception into a global rule, funds every prototype automatically or treats self-review as independent evidence. Revalidation and stopping are therefore part of its productive mechanism, not obstacles to speed.
A useful engine leaves the business with fewer ambiguous commitments, better evidence and more maintained customer value within its actual resources. That is the proposed destination of this series. The wider AI collection connects it to the business judgments software still cannot make on the owner’s behalf.
Sources
- GitHub: About issues — issues represent ideas, feedback, tasks and bugs; counts do not establish paying demand.
- OpenAI: Evaluation best practices — task-specific cases and human calibration; not a universal release threshold, business validation or recommendation to depend on the retiring hosted Evals platform.
Loading comments…