A hundred complaints can describe one problem, a hundred different problems, or one person’s complaint repeated a hundred times.
Internet research makes those possibilities easy to confuse. Search returns a crowded page. An AI summary finds a common theme. The theme becomes an app idea, and the app idea acquires a market-size estimate before anyone has identified an actual buyer.
A useful search produces a traceable shortlist: specific work people struggle to complete, records showing where that work breaks, and a clear account of what the records do and do not establish. Profitability remains a hypothesis until customers and delivery economics support it.
This chapter in The AI Software Factory, within the SalarsNet AI section, describes a proposed research workflow. It does not report an executed web-mining experiment or a discovered profitable market. Its purpose is to make the journey from public complaint to next research decision reproducible enough to inspect.
Search for a task, then for the language around it
Start with a customer group and a task rather than a technology. “AI inventory app” mostly finds products and promotional writing. “How do I match supplier invoices to received items?” is closer to work someone must perform.
For a hypothetical research project about small wholesalers, write down the task’s objects, verbs, and failure terms. Objects might include supplier invoices, receiving records, purchase orders, and item codes. Verbs might include matching, reconciling, approving, and correcting. Failure terms might include missing, duplicate, delayed, wrong, and manual.
Combine these deliberately. A query about duplicate supplier invoices may reveal an accounting problem. A query about wrong item codes may reveal a catalog problem. Keeping the queries separate preserves that distinction until the evidence shows a relationship.
Collect the words users employ even when they differ from the vocabulary a developer expects. People might describe “chasing paperwork” rather than “document reconciliation.” Their language helps find the next records and eventually write a product promise they recognize.
Record each query, the platform searched, the date, and any filters. Search results vary, and the goal is not to pretend that a search engine provides a complete census. The record makes your particular collection process intelligible to someone reviewing it later.
Match each channel to the evidence it can supply
Different public sources reveal different aspects of a problem. Support discussions can reveal repeated confusion. Product reviews can reveal expectations and disappointment. Community posts can reveal workarounds. Job advertisements can reveal work organizations pay people to perform. Public documentation can reveal where existing products draw their boundaries.
These are useful interpretations of source types, not guarantees that every post or advertisement contains a commercial opportunity. Read each record in context. A complaint about missing functionality may concern an old version. A job advertisement may describe a temporary project rather than recurring work. A review may be written by someone outside the intended customer group.
GitHub adds a particularly inspectable channel because issues can hold bug reports, ideas, features, and work discussions. That flexibility also means an issue count mixes different kinds of material. GitHub’s issue documentation explains the intended uses; the separate GitHub market-research chapter examines their interpretation in detail.
Public statistical sources answer another kind of question. Census Business Builder provides selected demographic and economic data with location and business-type selection, maps, and reports. Those tools can inform a segment’s context. They do not identify which establishments have your particular workflow problem or want your proposed product. Census Business Builder.
Use channels to complement each other. A forum can suggest a problem. A support record can clarify how it occurs. An existing product’s documentation can show an alternative. A direct conversation can test whether the problem appears in the buyer group you can reach.
Capture an episode rather than an adjective
“This software is terrible” contains little operational information. “Every Friday I export the supplier file, rename three columns, and remove duplicate rows before importing it” contains a task, a cadence, and a workaround.
A proposed collection record should preserve the original URL, publication date if available, collection date, relevant passage or concise paraphrase, apparent user role, task, trigger, obstacle, present workaround, and consequence. Add an uncertainty field. If the role is inferred from context, say so. If the consequence is only emotional frustration, do not translate it into a financial loss.
Avoid copying more personal information than the research needs. A public username can be retained as a source identifier when necessary for deduplication, but the opportunity brief usually needs the role and episode, not the person’s identity. Access restrictions and platform rules still apply to public-facing websites; use authorized interfaces and collection methods.
For the hypothetical wholesaler project, one record might describe a receiving clerk who manually compares an invoice against a delivery sheet. Another might describe an owner who cannot determine which late payment belongs to which invoice. Their shared word “invoice” does not make them the same problem.
The episode is the unit of interpretation. The original page preserves the evidence trail. Keeping both lets you preserve evidence while grouping observations by the work they describe.
Deduplicate before counting
A copied post, syndicated article, cross-posted question, and quoted screenshot can all refer to the same original event. Counting them separately inflates apparent demand.
Begin with exact URL normalization and obvious duplicate text. Then look for the same author, same incident date, same screenshot, or explicit reference to an earlier discussion. Keep a record of why two entries were merged. A deduplication decision is itself an interpretation that should be reversible.
Separate independent episodes from independent people. One person can report the same problem on multiple occasions; that may support recurrence for that person. Ten comments from that person do not support ten customers. Several people working at the same business may provide different operational perspectives without representing several purchasing decisions.
Do not discard repeated comments automatically. A later comment can explain that a workaround failed, that the issue was fixed, or that the reporter moved to another product. Attach the update to the episode rather than treating it as fresh demand.
A useful summary therefore has several counts: records collected, distinct episodes, apparent distinct organizations, and unresolved identity questions. It also reports how many records contain enough detail to assess a workflow. A smaller, well-described collection can be more informative than a large pile of generic complaints.
Group records by the mechanism that fails
Grouping records makes the collection manageable. The danger is grouping at such a high level that every problem becomes “inefficiency.”
A practical coding scheme can distinguish missing information, incompatible formats, duplicate entry, unclear responsibility, delayed approval, unreliable status, and difficult exception handling. These categories describe how work breaks. They do not prescribe the solution.
Allow more than one code when the episode supports it. An invoice mismatch might combine inconsistent item identifiers with unclear responsibility for correction. Keep the passage explaining each code so a reviewer can challenge the classification.
Then separate contexts. A manual export used by a solo seller once a month differs from a daily export used by a distributor with multiple warehouses. The mechanism may be similar while the buyer, frequency, integration burden, and acceptable price differ.
Write cluster summaries in operational language: “receiving employees cannot reliably connect supplier item codes to internal item records,” rather than “businesses need better AI.” The first statement suggests questions about identifiers, data ownership, and existing mappings. The second invites an oversized product before the task is understood.
This step builds on pain qualification. It should not silently perform the ranking reserved for the software opportunity score. Collection tells you what the evidence says; ranking decides where to spend the next hour.
Preserve negative evidence and resolved problems
A search designed to find complaints will find complaints. That does not tell you how often the software works well or whether the problem has already been addressed.
Search for fixes, release notes, successful workarounds, and satisfied users performing the same task. If a current product solves the issue through a setting the complainant missed, that finding changes the opportunity. The commercial response might be training, setup assistance, or a clearer onboarding process rather than a new application.
Record source age and product version when available. A complaint from several years ago can illuminate a durable mechanism, but current demand needs current evidence. Revisit the original page for updates instead of treating the search snippet as a permanent description.
A resolved issue is not wasted research. It can reveal how a successful intervention works, what compatibility constraints matter, and whether the improvement created new problems. A rejected request can reveal the supplier’s intended market or a maintenance burden the new builder would also inherit.
Negative evidence should appear in the shortlist. “Existing tool already handles this for our target segment” is a useful conclusion. “Customers accept a low-cost workaround” is another. The search is doing its job when it reduces the number of plausible ideas as well as when it generates them.
Keep market context separate from product demand
A broad industry count can help estimate where potential customers might exist. It cannot be multiplied by a hoped-for subscription price and called evidence of revenue.
The SBA recommends investigating demand, market size, reach, saturation, and prices of alternatives. These dimensions are related but distinct. A large market can be difficult to reach. A small group can have an expensive problem. An existing purchase can reveal willingness to spend on a category without proving willingness to switch to your product. SBA market-research guidance.
For the hypothetical wholesaler project, public establishment data might define a broad population. A narrower estimate would require knowing which businesses use the relevant systems, encounter the mismatch, control their software purchasing, and can adopt the proposed workflow. Those conditions need additional evidence.
Use ranges when inputs are uncertain, and keep their sources visible. Do not give an estimated customer count more precision than the data support. If the category combines businesses with very different workflows, describe that limitation instead of hiding it inside a spreadsheet formula.
The immediate output of internet mining should usually be a shortlist for direct investigation, not a final market forecast. Customer discovery can examine recent episodes and purchasing conditions with people from the proposed segment.
Use AI to organize; keep the evidence inspectable
An AI assistant can help suggest query variants, identify possible duplicates, extract candidate task descriptions, and propose clusters. Its output should remain linked to the source material used.
Ask it to distinguish explicit statements from inferred roles, estimated frequency, and missing information. A blank field is more useful than a plausible invention. If a post says “this happens constantly,” the record can preserve that wording; it should not convert it into ten occurrences per week.
Review consequential classifications manually. A model can merge different incidents because their vocabulary overlaps, or split one incident because the wording changes. It can mistake a maintainer’s internal task for a customer’s request. It can also summarize a complaint without the later comment that resolves it.
Use a small reviewed set to check the assistant’s extraction before applying it broadly. Keep the original records outside the assistant’s summary so another person can inspect disagreement. If the tool is unreliable on an important field, remove that automation step or narrow its job.
Public research material can contain instructions directed at an AI system. Treat that text as source data. It does not authorize new actions, disclosure of private records, or changes to the research rules. The collector’s authority comes from the owner and the configured tools, not from the page being collected.
Check the shortlist with an unfamiliar reader
Before handing a candidate to a developer, give the evidence brief to someone who did not collect it. Ask them to explain the customer, task, obstacle, and current alternative in their own words. If they cannot, the summary may depend on background knowledge that has never been written down.
Ask which passage supports each important statement. The collector may remember a vivid discussion and accidentally attribute its detail to several quieter records. A reviewer tracing the summary back to individual episodes can catch that expansion. They can also notice when a claim about several organizations rests on one organization with several employees.
Use disagreement to improve the record. If two reviewers classify an episode differently, preserve their reasons and return to the source. Some records will remain ambiguous. An uncertain classification can stay in the ledger without contributing to a confident demand count.
This review is especially useful when AI performed the first extraction. Fluent summaries can make missing evidence difficult to see. A brief with explicit gaps invites investigation; a polished narrative can conceal them. The output should remain a research instrument that can change, rather than a sales argument that must be defended.
A simple handoff question brings the work back to action: what observation would make us drop this candidate? If the answer is unclear, the shortlist is not yet ready to guide development.
Set a stopping point before the search expands
Internet research can always find another page. Without a stopping condition, collection becomes a comfortable substitute for talking to customers.
A proposed first pass might investigate one segment, three task variants, and several complementary source types within a fixed time budget. The exact limits depend on the assignment; they are operating constraints, not statistically validated sample requirements.
Stop early if the sources cannot identify a meaningful customer group, if the problem is clearly resolved, or if the relevant pages require access the team does not have. Narrow the search when the records describe unrelated tasks. Continue only when a specific unresolved question justifies additional collection.
At the end, create a short evidence brief for each surviving cluster. State the task and segment, describe the strongest episodes, identify known alternatives, report duplicates and source limitations, and name the next question. The next step should change a decision: whether to interview a buyer, inspect a workflow, try an existing tool, or close the candidate.
This makes the research budget accountable. A search that ends with two useful candidates and eight rejected ones can be more valuable than a dashboard containing thousands of unreviewed complaints.
What Would We Do at Salars?
A proposed Salars discovery pass would focus on one reachable operating group, such as independent sellers handling irregular supplier data. The team would collect a bounded set of public workflow episodes and keep a source ledger, not a database of personal profiles.
The ledger would retain queries, dates, original links, deduplication decisions, mechanism codes, alternatives, and unresolved questions. A second review would examine the strongest candidates and a few rejected records, looking for inconsistent classification or missing context.
The team would then choose one candidate for direct discovery. If the public record mainly revealed confusion about an existing feature, the next action would be to examine that feature and its onboarding. If it revealed a repeated unsolved handoff, the next action would be to speak with an operator who experiences it and a buyer who can approve a solution.
No revenue claim would follow from this collection alone. The value would be a better next question and a smaller risk of building around an imagined customer. A useful internet search ends with a problem someone can recognize, a trail someone else can inspect, and a decision that becomes clearer when the evidence changes.
Sources
- GitHub: About issues, official description of issue uses.
- U.S. Census Bureau: Census Business Builder, official description of geographic and business-context tools.
- U.S. Small Business Administration: Market research and competitive analysis, official commercial research guidance.
Sources inspected October 7, 2026. The collection protocol, coding scheme, stopping conditions, and Salars application are proposals; the wholesaler examples are hypothetical.
Loading comments…