A public complaint is a lead, not a business. The person reporting it may lack purchasing authority, prefer a free workaround, or be describing a condition already fixed by a recent release. Ten reports may concern the same underlying incident. A discovery system that turns those reports directly into app recommendations can generate a persuasive queue of opportunities without establishing that any deserves investment.
Salars Opportunity Radar would organize evidence for bounded experiments, with analyst review between a signal and a commitment. It is a proposed system, not an operating service with verified leads, customers, or measured prediction accuracy. This article specifies ingestion, provenance, deduplication, evidence grading, scoring, and review within the AI Software Factory series and our AI guides. Its output would be a reviewable question and next test, rather than a declaration of validated demand.
Give the radar a decision it can support
The first decision would be whether a narrowly described buyer problem deserves a small investigation. Radar would not decide that a market is profitable based on public discussion alone. It could help an analyst find recurring conditions, compare alternatives, identify missing evidence, and prepare an experiment brief. The decision to fund implementation would require additional buyer, access, delivery, and economic evidence.
For a proposed Supplier Margin Guard candidate, the job might be comparing an eligible supplier quote with the corresponding invoice and identifying differences for review. Public comments about unexpected charges could supply leads. They would not establish how many purchasing managers experience the condition, whether documents can be obtained appropriately, or whether a software service would be preferred over an existing accounting process.
Define the first collection boundary: one buyer role, one job family, a few permitted sources, a limited period, and a review capacity. A broad system collecting every business complaint would face unclear classification and overwhelming noise. Narrowing makes false positives and missed relevant cases easier to understand. It also makes the analyst’s domain knowledge useful rather than asking a model to infer every market at once.
The existing profitable software problems guide teaches discovery methods. Radar would coordinate those methods and preserve evidence. The design should earn its operating cost by improving a real analyst decision relative to a manual baseline, not by producing more candidate names than a person can responsibly review.
Select sources by relevance and permitted use
A source policy would identify which places can supply the intended evidence, what access method is allowed, what content is relevant, and what should not be collected. Public visibility does not settle every question about access, reuse, confidentiality, personal information, or publication. Technical ability to retrieve a page is only one part of the collection assessment.
GitHub issues can be useful for understanding technical friction. The official issues documentation describes issues as a way to track bugs, features, ideas, and team work. A reporter may be a developer rather than a buyer. A maintainer’s priority may reflect volunteer capacity rather than market demand. The platform supplies a discussion mechanism, not a ready-made list of paying customers.
Other sources might include vendor documentation, permitted customer interviews, public workflow discussions, and appropriately obtained examples. Each source answers different questions. Documentation can establish a technical capability. An interview can reveal a reported workflow. A paid pilot can provide evidence of a specific transaction and delivery outcome. Radar should not combine them as though every record has the same evidentiary force.
Source eligibility would be reviewed before ingestion. A private community, customer document, or partner export would need its own handling and permission assessment. The first Radar version could exclude uncertain sources rather than attempt to solve all rights questions through automation. A smaller permitted collection is more useful than a large archive whose contents cannot responsibly support the intended next action.
Preserve provenance at the point of ingestion
Each observation would record a source identifier, URL or permitted reference, collection time, source date when available, retrieval method, relevant passage, and the narrow claim it supports. It would distinguish content actually read from a search snippet or inferred summary. An analyst should be able to return to the evidence and understand why the system classified the record.
Preserve the original meaning without storing unnecessary material. A short relevant excerpt or structured description may be sufficient for some purposes, subject to rights and handling requirements. Do not copy an entire discussion merely because it is easy. If the source later changes, a record of what was observed and when can explain the difference, while the current source remains important for present decisions.
Record uncertainty about dates. The day a crawler sees a page differs from the day the underlying problem occurred. A recently edited page might describe a years-old incident. A complaint can predate a vendor fix. Radar would avoid presenting collection freshness as proof that the pain remains current. Where occurrence time is unclear, the analyst would see that limitation.
An observation would link to a source policy and allowed-use category. A record permitted for internal review might not be suitable for public quotation or model training. Those categories would follow the actual arrangement rather than a universal rule invented by the control plane. Provenance would therefore include handling context as well as factual origin.
Keep observation separate from interpretation
A source might say that an export dropped a field. That observation supports a reported problem under described conditions. The model might interpret it as a compatibility issue. An analyst might propose a converter. The commercial hypothesis might be that an eligible operator would pay for reliable conversion. Each step adds assumptions, and Radar should retain the distinction.
A candidate record would contain separate fields for the source statement, interpretation, buyer hypothesis, alternative explanation, and proposed test. If the source does not identify a purchasing role, the buyer field would remain unknown. If the cost of the problem is unreported, Radar would not manufacture an annual loss estimate to complete the scorecard. Unknown information should change the next investigation.
Require the interpretation to cite its supporting observations. A model summary that says many merchants struggle needs identifiable evidence and a clear meaning of many. A group of repeated comments from one incident may support an incident summary, not a broad recurring problem. An analyst should be able to challenge the grouping and see what happens to the candidate claim.
Avoid converting emotion into financial urgency. Frustration can reveal a real problem without establishing willingness to pay. A calm technical report may describe a highly consequential failure. Classify the job, consequence, workaround, recurrence, and evidence rather than score only the intensity of language. The first system would need human review of ambiguous cases.
Deduplicate events without erasing corroboration
Duplicates can arise from copied posts, cross-linked issues, repeated reports by one person, bots, shared outages, and discussions of the same underlying defect. Radar would distinguish an observation, a reporter, a source thread, an incident, and a candidate problem. These identities prevent ten messages about one outage from becoming ten independent opportunities.
A proposed matching process could use source identifiers, quoted text overlap, linked references, time, affected component, and analyst-reviewed similarity. Exact duplicates can often be handled mechanically. Semantic similarity is less certain. Two reports about invoice discrepancies might concern tax handling and unauthorized substitutions, which require different workflows. Combining them too aggressively can hide important conditions.
Keep a record of grouping decisions and their reasons. An analyst could split a cluster, merge duplicate incidents, or mark a relationship uncertain. The original observations would remain traceable where retention and permissions allow. Do not delete inconvenient counterexamples to make a cluster coherent. A good grouping exposes the differences that determine the next test.
Corroboration would remain visible after deduplication. Five independent operators describing the same condition can be more useful than five posts from one operator, but independence itself needs assessment. They might share an employer, copied tutorial, or platform incident. The system would describe observed relationships rather than claim perfect independence from distinct usernames alone.
Define a taxonomy around the buyer’s work
A useful classification might include buyer role, job, current alternative, failure mode, recurrence, consequence, input access, supported action, and evidence maturity. It should help choose a next investigation. A large taxonomy of industry labels can look sophisticated while supplying little guidance about what the buyer actually needs.
For the invoice candidate, a reported price difference could be categorized as comparison friction, a potential contractual disagreement, an input-quality issue, or a legitimate approved change. Those categories imply different products and different authority. A tool that identifies differences for review should not imply that every difference is an improper charge. The taxonomy would preserve that distinction.
Include exclusions. A problem might be outside the intended buyer segment, already resolved, prohibited by a platform condition, or dependent on unavailable data. An excluded record can still be useful evidence about the boundary. It should not disappear from evaluation simply because including it lowers the apparent quality of the discovery system.
Review the taxonomy against real permitted examples before expanding it. If analysts repeatedly disagree, determine whether the labels are unclear or the source lacks information. The correct response may be an unknown category and a question, rather than forcing agreement. Classification quality matters because a score built on incorrect categories can be precise and useless.
Grade evidence by what it can establish
A proposed evidence ladder would distinguish an unverified public signal, a current corroborated condition, a direct buyer account, a permitted observed workflow, a paid bounded job, and a maintained outcome. These are management categories for this design, not a universal scientific scale. A record might supply strong technical evidence while providing weak demand evidence.
Grade dimensions separately. Technical feasibility, buyer role, consequence, willingness to pay, data access, delivery cost, and repeat need each require suitable evidence. A vendor document can support feasibility while leaving all commercial dimensions unknown. A payment can show that one buyer accepted particular terms without proving a broad market or a sustainable recurring model.
The SBA’s business planning guidance discusses market questions and direct research alongside secondary information. Radar would use public signals to prepare direct investigation, not replace it. The guidance does not provide a formula that guarantees software profitability. The system’s role would be to identify the most decision-relevant gap.
Evidence grades would include date, scope, limitations, and revalidation triggers. A compatibility finding can expire when an API changes. A buyer’s stated preference may differ from behavior at a real price. A successful assisted pilot may not validate self-service. The interface would show those limits where the analyst chooses the next stage, rather than hide them in a separate research folder.
Use scoring as a review aid
A score could help prioritize a queue when it summarizes explicit dimensions and missing information. It should not convert uncertain estimates into authority. The software opportunity score guide develops that distinction. Radar would display the underlying evidence and allow the reviewer to inspect how the result changes when an assumption changes.
A proposed scorecard might consider severity under the eligible workflow, recurrence, current alternative quality, buyer access, data feasibility, delivery burden, and fit with available capabilities. Some dimensions could be categorical until evidence supports quantification. Unknown would not automatically receive a neutral midpoint. Depending on the decision, an unknown data-access condition could block implementation while still permit an interview.
Avoid false precision. A score of seventy-eight does not establish that a candidate is better than one scored seventy-six when both depend on unverified assumptions. Use bands, reasons, and sensitivity where appropriate. The purpose is to direct limited analyst attention toward questions worth investigating, not rank businesses with a simulated investment return.
The reviewer would record why a candidate advances, waits, narrows, or stops. An override could be justified by evidence the score misses, such as a temporary access constraint or a support burden beyond current capacity. Capture the reason and revisit it when conditions change. A score becomes harmful when the team obeys it without understanding the decision it was designed to support.
Include counterevidence in every candidate brief
A credible brief would ask what could make the opportunity unattractive. The existing alternative may be adequate. The problem may occur rarely. The buyer may lack authority or budget. The data may be inaccessible under the intended terms. A platform may prohibit the proposed action. A specialist may already solve the job cheaply. Those possibilities deserve a visible place before development.
Search for current resolutions and substitutes as well as complaints. A closed issue may contain a fix or a supported workaround. A vendor release note may change the compatibility assessment. A community answer may show that the supposed product is a short configuration step. The GitHub issues market research guide examines how to interpret those signals in context.
Do not treat every competitor as proof of demand or every absent competitor as proof of opportunity. Competition can reveal a viable job while leaving the newcomer without differentiation. Absence can reflect neglected demand, but it can also reflect weak economics or difficult rights. Record alternative explanations and choose a bounded test that can discriminate among them.
A candidate that fails a key condition can remain in the archive with the reason and date. That prevents the next collection cycle from rediscovering the same idea as fresh evidence. Reopening would require a relevant change, such as new access or a different buyer scope. The archive would retain learning, not an endless list of dormant launch plans.
Turn the brief into one proposed experiment
The output would name the next question and the smallest appropriate test. If buyer willingness is unknown, a compatibility benchmark alone will not answer it. If data access is unknown, an enthusiastic interview will not establish it. Match the test to the uncertainty that controls the next commitment. The validate before coding guide provides the broader sequence.
A proposed invoice experiment could ask whether an eligible purchasing operator values a traceable comparison enough to accept a paid, review-required service. It would need a defined offer, permitted inputs, independent correctness criteria, budget, stop date, and clear exclusions. It would not market identified differences as recovered money. No such experiment is reported as executed here.
Another proposed experiment could inspect whether required line identifiers exist in a documented input family using independently created or appropriately permitted examples. A favorable result would support compatibility under tested conditions, not demand. A negative result could justify narrowing the offer or stopping before a costly build. Both outcomes would improve the decision if the test was designed around a material uncertainty.
The experiment brief would include protected cases and counterexamples where appropriate, including legitimate changes that should remain unresolved or accepted. An analyst would approve the plan’s scope before a delivery system receives it. Candidate status would change to experiment approved, rather than validated opportunity. The wording matters because it prevents a plan from being mistaken for evidence.
Hand off to Forge with clear acceptance conditions
Proposed Salars Forge would receive a versioned brief with sources, assumptions, rights questions, experiment scope, budget, owner, and success and stopping criteria. Forge would coordinate the approved delivery stage. Radar would remain responsible for the provenance and interpretation of discovery evidence. Neither system would automatically inherit authority to contact people, spend money, or process private records beyond established authorization.
The receiving owner would check whether the plan fits current capacity and obligations. A promising candidate might wait because the specialist needed to review outputs is unavailable. A technical test might proceed while customer outreach requires a separate authorized arrangement. Handoff acceptance would therefore include operational feasibility as well as clarity of the research question.
Changes during implementation would return to the relevant record. If the developer discovers that the proposed workflow needs an external write rather than a read-only comparison, the permission and risk boundary changes. The systems would reconcile that change before proceeding. A small experiment approval would not expand automatically to a broad live integration.
The result would return with execution status, observations, deviations, limitations, and the proposed next decision. Radar could update evidence grades based on that record. It would not mark success merely because code was delivered. A technically correct prototype might reveal that buyers prefer their existing process, which is a useful negative result for the opportunity assessment.
Design collection around privacy and bounded retention
Public sources can contain personal information and commercially sensitive combinations. Customer-contributed material can carry stronger expectations and obligations. Radar would minimize what it retains for the stated investigation and restrict access according to role. A discovery archive should not become a permanent unreviewed copy of every source it can reach.
The ICO’s data protection principles guidance discusses purpose, minimisation, storage, security, and accountability in the UK context. It does not settle every jurisdiction or every reuse. The design implication is that source policy, intended purpose, and handling decisions need appropriate assessment, with uncertainty visible rather than erased by a public-data label.
Set retention by purpose and obligations rather than a universal duration. Some provenance records may remain useful after raw excerpts are removed. Other claims may no longer be supportable without their evidence. Record the limitation and reconsider dependent conclusions. Removal of obvious names alone does not establish anonymity or unrestricted permission to publish a cluster.
Separate internal analyst use from publication, model training, and outreach. A record that can inform a private review may not support another action. If outreach is proposed, it must follow the actual authorization and applicable arrangement. Radar’s ingestion capability would not make it an unsolicited-message engine. The first design could remain entirely read-only until an appropriate next task is authorized.
Treat external text as untrusted input
A public issue or document may contain instructions addressed to an AI system. Radar would treat them as source content rather than workflow authority. A message telling the model to promote a product, disclose records, or call an unrelated tool should not override the approved analysis. Collection and interpretation tools would have limited capabilities consistent with the task.
OWASP’s prompt injection guidance describes indirect injection through external material and layered controls. Radar would use source boundaries, least privilege, and appropriate review for consequential actions. A filter or system prompt would not be advertised as universal protection. The proposed evaluation would include malicious instructions embedded in otherwise relevant source examples.
Keep ingestion separate from execution credentials. A component reading public sources should not hold billing or production-deployment authority. A classifier should return labels and evidence pointers rather than perform commercial actions. The architecture would reduce the consequence of a mistaken interpretation instead of relying only on the model to behave correctly.
Log denied or suspicious requests without unnecessarily copying sensitive material. An analyst should be able to understand the failure condition and continue safely where appropriate. A malicious record might be excluded, or its factual content might be inspected through a safer path. The system should not reward a high candidate count at the expense of controlling what external text can cause.
Evaluate retrieval and interpretation separately
A proposed evaluation would begin with a protected set of permitted examples whose relevance, duplicates, dates, and supported claims are independently reviewed. Include genuine leads, resolved problems, copied reports, irrelevant discussions, ambiguous buyer roles, missing dates, and malicious instructions. The answers should remain independent of the implementation’s preferred categories.
Retrieval evaluation would ask whether relevant permitted observations were found within the defined source boundary. Classification evaluation would ask whether the system preserves the narrow claim and uncertainty. Deduplication evaluation would inspect mistaken merges and splits. Candidate evaluation would ask whether the brief identifies a useful next question without claiming demand that the evidence cannot support.
The OpenAI evaluation best practices emphasize task-specific evaluation and human calibration. Radar’s thresholds would be defined for its decision and risk. A model-based grader could help inspect some records, but its agreement with qualified analysts and its biases would need assessment. Another model’s agreement is not independent proof of correctness.
Compare against a maintained manual baseline on the same eligible cases. Measure analyst time, missed relevant evidence, unsupported conclusions, duplicate handling, and decision clarity. Faster summaries can be harmful if they erase caveats. A slower but more traceable process can be preferable for consequential investment decisions. The result would support the tested local use, not universal opportunity prediction.
Follow one illustrative signal through the system
Imagine a hypothetical public report that a supplier invoice includes a charge absent from a quote. Ingestion would record the source and date, along with the narrow statement. The interpreter would identify a possible comparison problem, but buyer role, contractual legitimacy, recurrence, and willingness to pay would remain unknown. No estimate of recovered margin would be generated from that report alone.
A second report could concern the same incident, a different supplier, or an approved surcharge. Deduplication would preserve uncertainty until relevant evidence distinguishes those possibilities. The analyst might create separate clusters if the resolution differs. The candidate brief would include the possibility that the invoice is correct and the quote incomplete. That counterexample would influence the proposed comparison tool’s refusal behavior.
A current vendor feature might already perform the comparison adequately. The analyst would inspect its documented scope and relevant limitations before proposing a new app. If the alternative fits the job, Radar might recommend no build or a narrower service. That negative recommendation would be a valid outcome, even though it reduces the candidate queue.
If the remaining question is whether operators want a traceable review service, Radar would prepare a bounded test plan. Only an executed, appropriately documented test could update the commercial evidence. The lead would remain a lead while the plan is pending. This walkthrough is synthetic and proposed; it establishes no actual source event, customer interest, or Salars experiment result.
Budget review capacity before ingestion volume
Suppose an illustrative weekly collection returns sixty candidate observations. Automated grouping reduces them to twenty review items. If an analyst spends twelve minutes on each, review requires four hours: twenty times twelve minutes is 240 minutes. Add one hour of source-policy and quality checks and another hour of experiment preparation, and the planned workload becomes six hours. These are hypothetical quantities, not Radar performance measurements.
If the analyst has only three hours, doubling ingestion makes the queue worse. Narrow the source boundary, improve eligibility filters without hiding relevant cases, or obtain suitable review capacity. A growing backlog can make evidence stale and postpone the very decision the system is meant to help. Count reviewed decision value rather than raw records acquired.
Budget implementation and maintenance separately. A new source connector can break, source rules can change, and a taxonomy can need revision. The system would need an owner for those tasks. A low-cost crawler is not necessarily a low-cost discovery operation when interpretation and rights review dominate the workload.
Set a stop condition for Radar itself. If it produces no better decisions than the manual baseline or consumes more resources than its supported value, simplify or pause it. The first success could be a useful analyst worksheet rather than an autonomous discovery platform. Radar should earn continued funding through scoped evidence, just like the candidates it evaluates.
What Would We Do at Salars?
We would propose one narrow Radar pilot focused on a defined operator job, with a small permitted source set and manual analyst review. Supplier Margin Guard and Merchant Revenue Guard would remain candidate hypotheses, not established businesses. The first output would be a provenance-rich brief containing known facts, unresolved assumptions, counterevidence, and one proposed next test.
We would establish a manual baseline and protected evaluation cases before building a broad ingestion system. We would test stale reports, duplicate incidents, inadequate buyer evidence, available substitutes, and malicious source instructions alongside useful leads. A favorable local evaluation would support only the tested source and decision boundary. It would not establish prediction accuracy across markets.
We would preserve authority boundaries during the Forge handoff. An approved research step would not automatically authorize outreach, private-data processing, spending, or deployment. Actual execution records would update evidence maturity, while unexecuted plans remained visibly proposed. We would fund review capacity before expanding collection volume.
The proposed aim would be fewer unsupported commitments and sharper experiments. A candidate rejected because an existing alternative already solves the job could be an excellent Radar result. A candidate advanced with a clear question and modest budget could be useful without being called validated. The system would serve the investment decision, rather than persuade the business that every collected complaint deserves an app.
Sources
- GitHub about issues: issue purposes and access mechanisms, not a paying-buyer count.
- SBA Plan your business: market questions and direct research alongside secondary information.
- ICO data protection principles: UK-context purpose and handling principles.
- OWASP prompt injection guidance: indirect injection and layered controls.
- OpenAI evaluation best practices: task-specific evaluation and human calibration.
Sources checked October 7, 2026. Radar architecture, evidence categories, counts, workflows, and experiments are proposed or hypothetical. No deployed discovery service, validated demand, executed experiment, or predictive business result is asserted.
Loading comments…