Three app ideas can all sound promising while competing for the same week of work. One addresses an expensive task but requires a difficult integration. Another has an easy prototype but no obvious buyer. The third solves a smaller problem for people the builder can reach today.
A score can make the comparison explicit. It can also turn guesses into numbers that look like facts. The useful question is how to rank opportunities while keeping the evidence, uncertainty, and reasons for rejection visible.
This cornerstone of The AI Software Factory, in the SalarsNet AI section, proposes a decision framework for choosing the next software investigation. It is not a validated predictor of startup success. No customer-outcome experiment has been executed for the framework presented here.
The aim is modest and valuable: make competing ideas comparable, expose the assumptions that determine their order, and direct the next research effort toward information that can change a decision.
Decide what the ranking is for
A list can rank which problem to investigate, which paid pilot to offer, which app to build, or which operating product deserves more capital. Those are different decisions.
An investigation ranking should reward reachable evidence and an important unresolved question. A build ranking should require much stronger evidence of buyer commitment, delivery feasibility, and scope. A capital-allocation ranking should use operating data such as retained customers, contribution, support burden, and reliability.
Using one score for all four stages creates confusion. An early idea with a dramatic market story can outrank an operating product with real but less glamorous results. A small discovery project can be rejected because it lacks data that only discovery could produce.
Name the decision at the top of the worksheet. For this chapter, the decision is: which candidate deserves the next bounded investigation? Passing the ranking does not authorize a full build or release.
The existing product opportunity score concerns ecommerce product selection. This framework concerns software candidates, including buyer access, adoption, integration, maintenance, and delivery obligations. The two may share decision principles without sharing weights or claiming identical economics.
Separate eligibility from preference
Some conditions should prevent an opportunity from advancing regardless of its average score. If the team lacks lawful access to essential data, cannot responsibly deliver the promise, or cannot identify a plausible buyer, a high total on other factors should not erase the issue.
Use eligibility gates before weighted comparison. A gate can be unresolved rather than permanently failed. For example, a proposed integration may require permission from a customer. The candidate can wait while that permission is investigated. It should not proceed as though access already exists.
A practical first gate asks whether the customer and task can be described without embedding the proposed product. Another asks whether at least one traceable episode supports the problem. A third asks whether the team can investigate the candidate within its authority and available capacity.
These gates are proposed operating rules. They should be tailored to the assignment rather than treated as universal scientific thresholds. Their purpose is to stop a spreadsheet from compensating for a fundamental constraint with optimistic scores elsewhere.
A medical decision product, for example, does not become a suitable first experiment merely because the market is large and the pain is severe. Competence, validation requirements, and consequences belong in the eligibility decision. A supplier-file tool may face fewer barriers while still needing responsible handling of confidential data.
Choose dimensions that describe different things
A score can double-count one attractive fact under several headings. “Large market,” “high revenue potential,” and “many customers” might all arise from the same uncertain population estimate. Their combined weight can dominate the worksheet without adding independent evidence.
For an investigation ranking, use distinct dimensions: problem consequence, recurrence, buyer reach, existing alternatives, adoption burden, delivery feasibility, and evidence quality. Each should have a defined question.
Problem consequence asks what happens when the task goes badly. Recurrence asks how often it occurs for the same customer. Buyer reach asks whether the team can access people who experience the task and people who can approve a purchase. Existing alternatives asks how satisfactorily the present methods solve it.
Adoption burden asks what customers must change to receive value. Delivery feasibility asks whether the team can supply the promised result within a bounded scope. Evidence quality asks how directly the current record supports the assessment.
The SBA’s guidance provides useful commercial context: demand, reach, saturation, alternatives, and prices deserve separate examination. The proposed score translates those questions into a software-investigation worksheet; it is an editorial design, not an SBA scoring method. Market research and competitive analysis.
Do not add a dimension merely because it sounds strategic. If nobody can explain how a score should change after observing new evidence, the dimension is probably too vague to help.
Write anchors before assigning numbers
A rating of four means little until the scale explains what four represents. Write anchors using observable conditions rather than adjectives such as excellent, huge, or disruptive.
For buyer reach, a low anchor might mean that the team cannot identify an authorized way to contact the proposed segment. A middle anchor might mean that relevant people can be reached but the purchasing role remains unclear. A high anchor might mean that the team has a legitimate path to both users and buyers for a bounded discovery conversation.
For adoption burden, a low favorable rating might mean that customers must replace a core system before seeing any benefit. A high favorable rating might mean that they can try the proposed result using a harmless copy of existing data, with a clear way to stop.
These anchors do not require a five-point scale. Three well-defined categories may work better than ten loosely defined numbers. Use enough resolution to distinguish meaningful conditions without pretending to measure differences the evidence cannot support.
If the scale uses numbers, state which direction is favorable. “Adoption burden” and “adoption ease” can reverse the sign accidentally. A worksheet should not reward a larger burden because the person entering values forgot the direction.
Attach an evidence note to every rating. The note should identify the record, observation, or explicit assumption behind it. A number without its reason cannot be challenged constructively.
Do not hide missing evidence inside a middle score
A common convention assigns three out of five when information is missing. That treats uncertainty as average quality. It can reward the least researched idea, because known difficulties reduce other candidates’ scores while unknown difficulties remain invisible.
Mark unknown values explicitly. A candidate with unknown buyer authority differs from one whose buyer authority is confirmed and moderately accessible. The investigation may be valuable precisely because the unknown can be resolved cheaply, but that is a separate judgment.
For each dimension, record a plausible range and a confidence note when a point estimate would be misleading. Buyer reach might be low until a partner confirms access, then high if the access is legitimate and practical. Delivery burden might range widely until a real input sample is examined.
Use the ranges to ask whether the ranking is robust. If one candidate leads under every reasonable value, more precision may not change the next decision. If a small change in one uncertain input reverses the order, that input deserves investigation.
This is a way to allocate research, not a statistical confidence interval unless the data and method actually support one. Label the ranges as judgment ranges. Their honesty is more useful than formal-looking uncertainty derived from guesses.
Work through a hypothetical comparison
Suppose a small builder is considering three hypothetical candidates. Candidate A converts irregular supplier files into a consistent format. Candidate B provides a general AI planning dashboard. Candidate C checks whether product photographs match intake records.
The builder can reach two operators for A and C through an authorized business relationship. The intended buyer for B is vague. A has a documented recurring task, but file diversity may create costly configuration. C has a narrower workflow, but the frequency and consequence of mismatches remain uncertain.
An illustrative worksheet might use a one-to-five favorable scale with equal weights at the beginning:
| Dimension | A: supplier converter | B: planning dashboard | C: photo-record check |
|---|---|---|---|
| Problem consequence | 4 | 2–4 | 2–4 |
| Recurrence | 4 | 2–4 | 2–3 |
| Buyer reach | 4 | 1–2 | 4 |
| Alternative gap | 3 | 1–3 | 3 |
| Adoption ease | 3 | 2–4 | 4 |
| Delivery feasibility | 2–4 | 3 | 3–4 |
| Evidence quality | 4 | 1 | 2–3 |
Every value in this table is hypothetical. The table illustrates how uncertainty can remain visible; it does not rate real products or establish the best business.
A simple sum gives A a range of 24–26, B a range of 12–21, and C a range of 20–25. Recompute rather than trusting the attractive table: A’s fixed values total 22 before adding its delivery range of 2–4. C’s fixed values total 11, with four uncertain dimensions contributing 9–14. B’s ranges and fixed values total 12–21.
Under these illustrative assumptions, B’s upper range remains below A’s lower range. A and C overlap. That suggests a focused next question about A’s configuration burden or C’s actual mismatch consequence. It does not justify coding A immediately.
The calculation makes one advantage of ranges visible: the worksheet can distinguish a clearly weak current case from a close comparison without inventing decimal precision.
Use weights to express a decision, not discover truth
Weights represent the decision-maker’s priorities. They do not become objective because they add to 100 percent.
A builder with two weeks of available time may place substantial weight on buyer reach and delivery feasibility. A team choosing among validated pilots may place more weight on recurring value and support economics. Both can be reasonable for their respective decisions.
Write why each weight differs from the baseline. If feasibility receives twice the weight of market breadth, explain that the current decision concerns a small investigation the team can complete. This prevents later readers from assuming the worksheet was intended to identify the largest possible company.
Avoid adjusting weights until the favored idea wins. That produces a justification machine. If weights change because the decision changes, retain the earlier version and state the new decision. If they change because evidence reveals a missing dimension, record that correction.
An equal-weight baseline is useful because it is easy to inspect. It is not automatically optimal. Compare the chosen weighting against it and ask what actually changes. A complicated weighting scheme that produces the same order may add maintenance without improving the decision.
Weights can make a disagreement specific. One operator may favor a reachable niche; another may favor a broader market. Seeing which weight drives the difference allows a direct conversation about goals and constraints.
Account for adoption and switching work
A product’s value depends on what the customer must do to obtain it. Importing historical records, configuring permissions, training employees, and adapting a process can consume the very time the tool promises to save.
Investigate the path to a first useful result. Can the customer try the output with a sample? Must they grant access to a live system? Does the product require a commitment from another department? Can they reverse the change if the trial fails?
A small application that fits an existing handoff can have an adoption advantage over a broad platform. The advantage is conditional: the narrow tool must still handle the important cases and avoid creating another disconnected record.
For the hypothetical supplier converter, a sample-based trial might be feasible. Yet the business may need to verify every transformed record before import. If that review takes as long as the current manual correction, the tool’s apparent benefit disappears.
Include that review burden in the rating. Do not score only the developer’s ability to make the mechanism run. The customer has to use the result, trust it, and handle exceptions.
The distribution-before-development chapter examines reaching customers. Adoption analysis begins after the customer arrives and asks what must happen before the promise becomes useful.
Put maintenance into feasibility
A prototype’s difficulty is a poor substitute for a product’s delivery burden. A quick script can require continual attention when customers supply changing formats, third-party APIs change, or ambiguous cases need judgment.
Estimate the obligation the team accepts after the first successful demo. Who handles failed jobs? What happens when an external service is unavailable? How does a customer report a wrong result? Can the team reproduce the issue without accessing unnecessary private information?
A candidate with a narrow and stable input contract may deserve a better feasibility rating than one that promises to accept anything. That does not mean the narrower candidate has more total market potential. It means its current promise is easier to deliver responsibly.
Separate required capabilities from attractive extras. If the first paid result needs only a reviewed file transformation, it need not include a general agent framework, an analytics dashboard, or automated publishing. A score should not reward an oversized scope merely because its imagined revenue is larger.
Maintenance also affects evidence needs. A single happy-path sample tells little about exceptions. Inspect representative difficult inputs before giving delivery feasibility a high rating. If no such inputs are available, retain the uncertainty.
Translate the next step into a small resource commitment
Two candidates with similar scores can require very different investigations. One might need an afternoon reviewing permitted sample files. Another might need several weeks arranging access to a buyer and learning a complex domain. The ranking should make that difference visible before the team commits.
Estimate the investigation’s time and direct expense separately from the product’s eventual build cost. The first estimate answers whether the next step is affordable. The second remains provisional because the investigation may change the scope. Mixing them can make a cheap discovery question appear to require a large investment, or make a quick prototype conceal expensive ongoing work.
For an illustrative planning comparison, suppose one investigation requires six hours and another requires twelve. At an assumed internal planning value of $30 per hour, their time commitments are $180 and $360 respectively. These are hypothetical resource valuations, not cash invoices, wage claims, or measured Salars costs. They help explain the opportunity cost of choosing one question over another.
Do not divide the score by those costs and call the result expected return. A point on an ordinal scale is not a dollar of benefit, and the intervals between rating categories need not be equal. The quotient can produce a number while losing the meaning of both inputs.
Instead, examine the actual decision. Does the twelve-hour investigation address an uncertainty that the six-hour one cannot? Could the larger investigation be broken into a cheaper first step? Would its result open or close a consequential option? If neither candidate has enough evidence to justify even a small commitment, keep both in the backlog.
Capacity also has a calendar dimension. A team can have enough total hours and still be unable to respond when the customer needs help. A pilot that requires immediate attention during the team’s busiest operating period may be infeasible. Record that constraint rather than treating every hour as interchangeable.
Separate commercial potential from moral preference
Builders sometimes choose a project because it helps people they care about, addresses a meaningful local problem, or improves their own operating independence. Those are legitimate priorities. They should be visible rather than hidden inside an inflated market score.
A small community tool may deserve development under a service or public-interest budget even when subscription economics are weak. Calling it a large commercial opportunity can damage the project by imposing unrealistic expectations. A profitable application can likewise be rejected because its incentives conflict with the operator’s values.
Keep a values and purpose note alongside the commercial comparison. State the relevant consideration in concrete terms: who benefits, what burden the team accepts, and what tradeoff it is willing to make. If the decision intentionally departs from the commercial ranking, explain that departure.
The worksheet should help the owner choose consciously. It should not convert every worthwhile activity into a claim about maximum return. Nor should a favorable purpose excuse inaccurate data, unreliable delivery, or misleading sales promises.
This separation also improves later evaluation. A community-service project should be assessed against its declared service outcome and budget. A commercial pilot should be assessed against buyer value and delivery economics. Comparing both only on revenue would misunderstand the first; comparing both only on social interest would evade the second.
Test the ranking against counterexamples
A ranking framework should survive more than the examples used to design it. Construct cases where a plausible score could lead to a bad decision.
One counterexample is severe pain with no purchasing authority. Another is easy implementation with a saturated alternative market. A third is recurring use with low value per occurrence and high support cost. A fourth is strong user enthusiasm that depends on unrealistic customization.
Check whether eligibility gates or dimensions expose each problem. If the worksheet ranks a candidate highly despite a missing essential condition, improve the decision process rather than lowering the candidate’s score informally.
Also construct an attractive low-volume case. A specialized task can be infrequent yet consequential enough to support a useful service. If the framework automatically rejects it because recurrence is low, the weights may be inappropriate for the decision.
These are conceptual checks under stated assumptions. They do not validate predictive accuracy. They help find internal contradictions and blind spots before real outcome data are available.
A framework that admits its scope is easier to improve. It can say that it ranks bounded investigations for a small team, rather than claiming to identify every valuable software business.
Design a bounded evaluation before trusting the score
The meaningful uncertainty is whether this framework improves the team’s investigation choices compared with a simpler baseline. A proposed evaluation would compare the scored ranking with an unweighted short judgment: customer, task, strongest evidence, biggest unresolved risk, and recommended next step.
Use no more than a few explicit hypotheses. One could be that the structured score identifies missing adoption information more consistently. Another could be that it reduces time spent investigating candidates with unresolved eligibility barriers. A competing explanation is that the worksheet adds paperwork without changing decisions.
Define success independently of the score itself. Do not call the method successful because selected candidates receive high later scores. Useful outcomes could include whether the investigation resolves a predeclared uncertainty, whether the next action is completed within its budget, and whether an independently reviewed decision changes appropriately when contrary evidence appears.
Reserve later candidates as protected evaluation cases before adjusting anchors or weights. If those cases influence revisions, they become development cases, and a fresh set is needed before making a transfer claim. Separate organizations or task families where shared examples would leak knowledge across the comparison.
Hold the investigation budget and decision stage steady. A method given twice the time can appear better because it collects more evidence. Record time, data identity, framework version, decisions, and actual outcomes. Costs that cannot be observed should be marked unavailable or estimated.
No such evaluation has been executed for this article. Real candidate records, independent reviewers, and subsequent investigation outcomes are missing. The framework remains proposed.
Keep the outcome measure fixed when results arrive
A builder can reinterpret almost any result as a success. A failed pilot becomes useful learning. A delayed product becomes careful discovery. A selected idea that never launches becomes a long-term option.
Learning can be valuable, but the evaluation needs a predeclared question. If the method is intended to choose the next investigation, assess whether that investigation produced the promised decision-relevant evidence within its limits. Do not silently switch the measure to eventual revenue when the early decision looks weak, or to learning when revenue disappoints.
Record rejected candidates as well as selected ones. Otherwise the team cannot examine false negatives: ideas the framework discarded that later evidence would have justified. A ranking system can be conservative and still miss useful opportunities.
Independent review should inspect the source evidence and decision rationale, not merely agree with the score. Two AI models reading the same summary can share the same omissions. Their agreement is not independent confirmation of demand.
Use outcome uncertainty honestly. A small exploratory comparison can support a local observation, such as more consistent recording of unknown buyer authority. It cannot establish that the method raises software-business success rates generally.
The framework should be revalidated after material changes in customer segment, team capability, product stage, data source, or scoring definitions. A useful result in one context should retain that context when reused.
Identify the next piece of information
A ranking becomes actionable when it points to a question that can reverse a choice. In the hypothetical comparison, A and C overlap. The next investigation should target the input that matters most to their order and can be clarified cheaply.
For A, inspect several representative file formats and estimate the recurring configuration burden. For C, examine recent mismatch episodes and their consequences. The goal is not to collect every possible fact. It is to resolve a decision bottleneck.
Suppose A’s configuration burden turns out to be high and unique for each customer. The next action may be a concierge service rather than automated software. Suppose C’s mismatch problem is rare and cheaply corrected. It may drop below A despite its easy prototype.
A score should make those updates straightforward. Change the relevant evidence note and value, preserve the previous state, and explain the resulting decision. Do not rewrite the earlier record to make the current judgment appear inevitable.
This is where ranking connects to customer discovery and payment validation. The worksheet organizes what to learn; those methods obtain stronger evidence.
Use a short decision record
For the selected candidate, record the decision date, stage, eligible alternatives, framework version, evidence links, uncertain inputs, ranking sensitivity, next action, budget, owner, and stopping condition.
The record can fit on a page. It does not need a new scoring platform. A spreadsheet or Markdown document can preserve the rationale as long as the evidence remains traceable.
Write the stopping condition in plain language. Close the candidate if the required data cannot be obtained legitimately, if an existing tool solves the intended customer’s problem adequately, or if the next investigation fails to identify a buyer who values the bounded result. The exact conditions should match the candidate, not become ritual language copied into every file.
A useful record also states what has not been decided. Choosing an investigation does not authorize a build. Choosing a pilot does not demonstrate retention. Choosing a product does not justify expanding its promise indefinitely.
The decision should be reviewable without asking its creator to reconstruct their memory. That is the practical payoff of explicit scoring: the assumptions survive long enough to be checked.
What Would We Do at Salars?
The proposed Salars approach would begin with a small candidate list drawn from traceable operating episodes. It would keep the software-investigation score separate from ecommerce product selection and from capital allocation among existing businesses.
The team would apply eligibility gates first, then a small anchored scale with evidence notes. Unknown buyer authority, data access, adoption work, and maintenance burden would remain explicit. The first comparison would use an equal-weight baseline before introducing weights justified by the current decision.
The top candidates would receive bounded investigations aimed at the uncertainties that could reverse their order. A file-conversion candidate might require representative redacted inputs. A customer-response candidate might require a recent record of unanswered requests. The team would not infer profit from a high score.
If enough real decisions accumulated, a proposed evaluation would compare this process with the simpler baseline, using independent outcomes and protected later cases. Until that work occurred, the method would remain a proposed operating aid with no measured effectiveness claim.
A score earns its place when it helps a team ask a better question, reject an unsupported promise, or choose a feasible next step. The numbers matter because their reasons can be examined.
Sources
- U.S. Small Business Administration: Market research and competitive analysis, official guidance informing the commercial dimensions.
- Government Digital Service: Learning about users and their needs, official guidance on evidence-based needs and separating needs from solutions; used as methodological context, not a commercial success model.
Sources inspected October 7, 2026. The score, anchors, hypothetical ratings, resource valuations, and evaluation design are proposals. No real candidate ranking, customer trial, or comparative effectiveness experiment was executed for this article.
Loading comments…