A support archive contains many problems. It does not necessarily contain solved problems. Someone may have accepted a workaround without confirming its result. A ticket may have closed because the customer stopped replying. A successful output may have depended on a manual correction that never reached the record. If a founder treats every closed ticket as a trustworthy example, the product can learn the wrong lesson with considerable confidence.
A useful corpus connects a permitted input, the conditions of the task, a verified resolution, and a clear statement of what may be reused. Its value comes from the decisions it improves, not from the number of documents stored. This article explains how to build and challenge that asset within the AI Software Factory series and our broader AI guides. The previous discussion of software moats asks whether an advantage survives a plausible rival. Here the narrower question is whether a collection of cases deserves to inform the next product decision at all.
Define solved before collecting more
Consider a proposed catalog-cleanup tool. A customer submits a file with duplicate product identifiers, the tool recommends a merge, and a reviewer approves it. Has the problem been solved? Approval establishes one event, but the final import might still fail. Two records might represent distinct variants that should remain separate. The apparent resolution could create a later fulfillment error. A label that says success must specify which result was confirmed and which remains unknown.
Define resolution at the level of the buyer’s job. For a file-preparation tool, the narrow outcome might be a valid import accepted by the destination system without losing designated fields. That is different from improved revenue, reduced returns, or a more accurate entire catalog. The definition should name the eligible input family, the required output properties, the acceptance check, and the follow-up interval where relevant. Avoid expanding the claim beyond the evidence just because a downstream outcome is commercially attractive.
Include an unresolved status as a legitimate result. A case can be useful because it demonstrates that an ambiguous identifier requires human review. The corpus should distinguish correct refusal, incorrect refusal, supported completion, unsupported completion, and unknown outcome. Collapsing these into good and bad labels prevents useful reasoning about coverage. A cautious system can appear worse on completion rate while being preferable on harmful error rate.
Write down who establishes the answer. The operator who generated a recommendation may have an incentive to call it correct. A customer may approve something without checking every detail. A domain reviewer may judge a rule correctly but lack evidence of the real-world outcome. Record those differences. The answer can remain provisional until a specified independent check or later observation supports it.
Preserve the chain from input to answer
A case needs provenance: where it came from, when it was observed, what was supplied, what transformations occurred, and which evidence supports the result. This does not require preserving every raw document indefinitely. It requires enough information to understand what the example means and whether its use is allowed. The design must reconcile reproducibility with the purpose and retention constraints of the actual service.
For the catalog example, retain a versioned schema description, the relevant field relationships, the proposed merge rule, the output check, and the review record when permitted. Identify the product version and configuration. A later developer should be able to tell whether an error came from parsing, a business rule, an unclear instruction, or an external import change. A screenshot of a success message cannot usually establish all of those facts.
Treat transformations as part of the record. If an employee removed a confidential field, normalized a date, or replaced a name with a synthetic value, note what changed and why. A sanitized case may behave differently from the original. That can be acceptable if the difference is understood. It becomes misleading when the team reports a synthetic reconstruction as an untouched customer observation.
Maintain an evidence pointer rather than an unsupported conclusion. “Reviewer accepted the recommended merge on this date under this rule” is narrower and more useful than “AI was correct.” If the original evidence has been deleted under a retention requirement, record that limitation and decide whether the remaining case can still support its intended use. A case without adequate evidence may belong in a hypothesis notebook rather than the trusted evaluation set.
Establish allowed uses at collection time
Customer data arrives for a reason. The customer may authorize processing a file to deliver one service without authorizing unrelated training, publication, or reuse across tenants. A corpus plan therefore needs an allowed-use decision for each category of information. Do not make the decision implicitly by copying everything into a shared development folder. Access to a record and permission to use it for a new purpose are separate questions.
Separate at least the operational record, the improvement example, and the published illustration in your design. An operational record might be required to provide support or explain a completed action. An improvement example might be used to test a parser or rule. A published illustration might reveal patterns even after obvious names are removed. Each use needs appropriate justification, access, retention, and review under the applicable arrangement.
The UK Information Commissioner’s Office explains principles including purpose limitation, data minimisation, storage limitation, security, and accountability in its data protection guidance. The guidance is specific to its legal context and is not a global legal opinion for every software business. Its relevance here is practical: a founder should be able to state why particular information is needed and how its handling supports that purpose. Other jurisdictions and contractual obligations may change the assessment.
Ask a qualified reviewer when the intended reuse is consequential or unclear. The software privacy design guide connects purpose and data handling to product architecture. The corpus design should preserve that discipline rather than treat research as an exemption. If a reuse cannot be supported, obtain appropriate permission where possible, use independently created examples, or decline that use. Losing an interesting record is preferable to building a claimed asset on an unsupported right.
Collect the minimum useful representation
A raw invoice or catalog can carry far more information than a rule test needs. It may contain names, addresses, commercial terms, unpublished product plans, credentials accidentally placed in a field, or identifying combinations. Start by defining the smallest representation that preserves the mechanism under evaluation. For a duplicate-key rule, the relationship among identifiers and variant attributes might be enough; the customer’s full inventory might not be necessary.
Minimal representation is a design exercise, not a promise that all risk disappears. A rare product combination can identify a business even when its name is absent. A date, supplier, and quantity can expose a transaction. Predictable values may be recoverable from simple hashes. Removing obvious identifiers does not establish that a dataset is anonymous or safe for unrestricted circulation. Document the remaining risk and the intended access boundary.
The OpenTelemetry documentation on handling sensitive data describes how instrumentation can capture sensitive information and discusses minimization and processing approaches. That source concerns telemetry rather than a complete corpus governance system. It supports the narrower caution that automatic collection is not neutral: the operator must understand the data it emits and receives. A diagnostic pipeline should not quietly become the source of an unreviewed customer-document archive.
Synthetic examples can be valuable when they preserve a known rule without exposing an actual customer record. Label their origin clearly. They can test whether the software follows the specified rule, but they do not establish how frequently that situation occurs in the market. Keep a separate source field so a later analyst cannot confuse ten thousand generated variations with ten thousand independently observed customer problems.
Design labels that preserve disagreement
A label should describe an observable judgment under a stated rule. “Bad output” leaves too much hidden. Was a required field dropped? Did the parser choose the wrong date convention? Was the output technically valid but unacceptable to the customer’s workflow? Were two reviewers applying different rules? Specific categories let developers repair the right mechanism and let management understand whether the product promise needs to narrow.
Use a structured label with room for explanation. For the catalog tool, a case could record parsing status, duplicate classification, permitted action, output validation, reviewer agreement, and later import result. Avoid a single score that obscures incompatible outcomes. A case can pass parsing and fail business interpretation. Another can contain an ambiguous source that no reasonable implementation should resolve automatically.
Disagreement is evidence about the task. If qualified reviewers reach different answers, investigate whether the instructions are incomplete, the source lacks information, or the customer’s policy permits several resolutions. Do not automatically choose the majority answer and erase the ambiguity. The product might need to ask which customer policy applies instead of returning a confident result. That refusal can become a valuable protected evaluation case.
Record label revisions and their reasons. A changed customer policy may make an older label obsolete without making the original reviewer incompetent. A discovered mistake may require rerunning prior evaluations. Versioning supports both distinctions. Without it, a team can improve its reported score merely by changing the answer key to match the software, while presenting the result as progress in the product.
Map coverage instead of counting rows
A corpus of five thousand cases might contain four thousand near-identical examples from one customer. That is not the same as coverage across supplier formats, language conventions, missing fields, product types, and customer policies. Build a coverage map around the conditions that can change the correct answer. Count distinct conditions and important gaps before celebrating volume.
For the catalog tool, a useful map might include file encoding, delimiter, identifier style, variant representation, expected destination schema, incomplete rows, and ambiguous duplicate relationships. Choose categories grounded in the actual product boundary. A broad taxonomy copied from another domain can create administrative work without improving evaluation. Add a category when it explains a consequential difference in behavior or buyer outcome.
Frequency and consequence deserve separate attention. A common harmless formatting variation may justify automation. A rare destructive merge may justify a hard stop. A corpus sampled only from successful jobs can miss both difficult cases and customers who abandoned the service before completion. Record the collection method and who is absent. Do not interpret an observed success percentage as a market-wide estimate without appropriate sampling evidence.
Hold back protected cases that development work cannot casually inspect or alter. Protecting them reduces the chance of adapting the implementation to the answers rather than improving general behavior. The AI evaluation suites guide develops task-specific tests. The corpus supplies material for those tests, but a training or development collection and an independent evaluation set should remain distinct in purpose and access.
Make access reflect the job
Not everyone who improves the software needs the same information. A developer might need sanitized structure and an established rule. A support specialist might need the original customer record for a permitted investigation. A reviewer might need outcome evidence. A publisher may need only an approved aggregate. Design access around those tasks rather than granting a whole team permanent access to every case.
Separate customer records from reusable test fixtures where feasible. Use clear tenant boundaries, controlled export, and a record of who accessed sensitive material. A case identifier should not accidentally permit someone to fetch another customer’s raw input. Development convenience is not a sufficient reason to put production documents into broadly shared repositories. Repositories are excellent for versioned code and carefully reviewed fixtures; they may be inappropriate for confidential source material.
Deletion and correction need a path through copies. If a record is removed from operational storage but survives in a test fixture, notebook, backup, or analyst export, the apparent action may not fulfill the actual requirement. Inventory the destinations before promising a deletion process. Some retention requirements may differ across records; understand them and communicate the limits accurately. This article does not prescribe a universal retention period.
Include incident handling in corpus operations. If a fixture exposes information, suspend access, assess the scope, correct the distribution path, and follow the applicable response obligations. The founder should know who owns that work before collecting sensitive cases. A valuable corpus needs an operating budget for permissions, review, secure handling, and maintenance as well as an extraction script.
Demonstrate that the corpus improves a decision
Run a bounded comparison that connects the collection to a useful change. Suppose the catalog tool struggles with a particular representation of variants. Build a proposed update using permitted development cases, then evaluate it against protected examples including legitimate separate variants. Measure correct resolution, unresolved cases, harmful merges, and review effort. A higher overall completion rate is insufficient if it hides a rise in destructive decisions.
Compare with a credible baseline. That might be the current product version, a deterministic rule, or the customer’s established manual process. Use the same eligibility conditions and disclose differences in assistance. If a new system receives extra human review, its apparent advantage may come from the reviewer rather than the corpus. That can still be useful commercially, but it should be described as an assisted process with its actual cost.
The OpenAI evaluation best practices recommend task-specific evaluation, representative difficult cases, and calibration with human judgment. That guidance does not supply a universal success threshold for this product. Set independent criteria for the buyer’s risk and workflow. A model-based judge can help inspect some outputs, but its own biases and agreement with qualified reviewers require assessment.
Preserve counterexamples after an improvement. If the update fixes one format and harms another, the result supports a scoped change, not a global claim. Expand testing only when the observed uncertainty warrants it. A corpus earns value when it lets the team discover that tradeoff before customers do. The advantage is a more informed decision, which might be to refuse a release as readily as to ship it.
Account for the maintenance burden
Consider a hypothetical monthly collection of eighty candidate cases. If review takes fifteen minutes each, that is twenty hours before permission checks, sanitization, disagreements, and evaluation maintenance. At an illustrative internal cost of $50 an hour, first-pass review alone costs $1,000: eighty times fifteen minutes is 1,200 minutes, or twenty hours. These assumptions are not Salars measurements or a market wage estimate. They show why a growing archive needs a budget rather than a congratulatory row count.
Suppose only twelve of those cases add a materially new condition or correct a consequential misunderstanding. It might be sensible to improve the intake filter while retaining enough ordinary cases for a fair view of performance. Selecting only unusual failures would distort frequency estimates; selecting only duplicates would waste review effort. State the use of each sample. A coverage set and an operational monitoring sample can follow different collection rules.
A retained case also creates future work. Rules change, destination systems change, permission changes, and new information can invalidate an answer. Assign an owner and a review trigger. A case may remain valid under an old version while no longer describing the current product. Archive or relabel it instead of silently letting old examples determine current decisions.
Calculate the return through avoided review, prevented errors, improved delivery, or better release decisions where evidence supports those mechanisms. Do not assume each new case adds commercial value. If the corpus consumes specialist time without changing product choices, pause expansion and investigate the bottleneck. The purpose is a maintained body of trustworthy evidence, not a collection that grows because storage is cheap.
What Would We Do at Salars?
We would propose a small governed case library before attempting a large shared learning system. No operating Salars corpus, customer permissions, or measured performance improvement is established by this article. For a candidate such as Supplier Margin Guard, the initial library would use independently created fixtures and any customer examples whose handling and intended reuse can be supported. Public availability alone would not settle all relevant rights.
Each admitted case would have an origin, allowed-use record, defined task, established answer or unresolved status, version information, coverage tags, owner, and retention decision. We would keep raw customer delivery material separate from reusable fixtures and publication examples. We would require a second look at high-consequence labels before using them to judge a release. If two reviewers disagreed, the ambiguity would stay visible.
A proposed first experiment would ask whether a narrow rule update improves matching without increasing unsupported conclusions on protected cases. We would establish the baseline, independent acceptance criteria, eligible inputs, budget, and stop date before running it. A favorable result would support that local rule change within tested conditions. It would not prove a general data moat, market demand, or lawful reuse in every jurisdiction.
We would connect supported findings to the software feedback loop, where product decisions require evidence of real customer outcomes. A corpus can help diagnose and evaluate an improvement. It cannot replace knowing whether the improved product solves a job buyers actually value. If the cases are unreliable, the permissions unclear, or the maintenance cost greater than their decision value, we would narrow or stop the collection.
Sources
- ICO data protection principles: purpose, minimisation, retention, security, and accountability in the UK legal context.
- OpenTelemetry handling sensitive data: instrumentation exposure and implementer responsibility for sensitive information; telemetry guidance is not comprehensive corpus governance.
- OpenAI evaluation best practices: task-specific evaluation and human calibration rather than a universal product success threshold.
Sources checked October 7, 2026. The catalog cases, cost calculation, and Salars library design are illustrative or proposed. No customer dataset, executed experiment, legal approval, or commercial advantage is claimed.
Loading comments…