A document app deletes an upload from the customer’s account. The same document remains in a failed-job log, a model request archive and a support attachment. The visible delete button worked exactly as implemented. The privacy promise did not match the system’s data lifecycle.
Privacy design begins by identifying what data the product needs, where it goes, who can access it and how its retained copies end. Those decisions shape architecture before launch. A policy page can describe the design, but it cannot make an undocumented collection path disappear.
This chapter of The AI Software Factory develops a hypothetical document-review app. It concerns product architecture and operating decisions, rather than the user-level question in what to share with an AI tool. Applicable legal duties depend on jurisdiction, customer role and processing; the article does not present one checklist as complete global compliance.
Start with the job and the minimum useful data
The hypothetical app checks whether a supplier document contains required commercial fields and prepares a report. It may need product identifiers, amounts and terms. It may not need employee names, private notes or unrelated account details embedded in the file.
Describe the job before designing the upload form. Which fields determine the result? Which are needed only for display? Which can be omitted, replaced or extracted without retaining the original document? A broad upload route is convenient, but it can collect more than the product can justify.
Minimization should preserve usefulness. Removing every context field can make the result wrong or force repeated clarification. The goal is the smallest information set that supports the promised task, with a clear rule for exceptions.
The UK’s ICO data-protection principles guide describes purpose limitation, minimization, storage limitation, security and accountability within the UK GDPR framework. Its current page notes that guidance is under review after legislative changes. Those jurisdiction-specific principles are a useful reference, not a complete legal assessment of this hypothetical app.
Draw the actual data path
Follow one safe example from upload to result. The browser sends a file, the server stores it, a worker extracts fields, a model receives selected content and the app stores a report. Logs, analytics, backups and support tools may create additional copies.
For each location, identify the data, purpose, owner, access and retention. Include failed paths. A failed parse may retain the original upload in an exception queue even when the successful path deletes it promptly.
The map should distinguish authoritative records from temporary processing state. An extracted amount can be needed for the report while the full original document is unnecessary after validation. The product may still need a source reference to explain the result; that requirement can be satisfied in different ways depending on the task.
Use the map to inspect the privacy promise. If the product says uploads are deleted after processing, verify which copies that statement covers. Ambiguous wording can make a technically accurate narrow behavior misleading to a customer who reasonably interprets it more broadly.
Tenancy is an authorization rule throughout the path
A customer identifier in a database row does not automatically isolate the data. Every read and write route needs to establish the authorized scope. Background workers, exports, search indexes and support tools can bypass a check that exists only in the page interface.
The hypothetical app should associate a document and report with the appropriate customer organization and permissions. A job reference alone should not grant access. The application needs a rule for members, invited reviewers and revoked users.
Test denied paths with safe fixtures. One account should not retrieve another account’s upload, report or trace. A support role may need limited evidence without receiving unrestricted file access. A deleted membership should affect pending jobs according to the accepted policy.
Agent security addresses tool authority. Product privacy adds the data-access invariant: the agent and every supporting service should receive only the tenant information required for the current authorized task.
Treat model processing as a deliberate disclosure
Sending content to a model provider is a processing decision. The developer should understand the actual product, account settings, contractual terms and retention behavior relevant to that request. A familiar provider name or paid plan does not settle every condition.
Select the smallest useful input and avoid sending fields the model does not need. Deterministic extraction or calculation can keep some work outside the model. If the model needs context, preserve the purpose and source relationship rather than attach the whole customer archive by default.
Provider behavior can change, and products within one provider can differ. Verify current official documentation and terms for the intended integration before implementation. This article deliberately makes no blanket assertion that a named provider never retains data or never uses it for another purpose.
The customer-facing explanation should correspond to the actual route. If a report uses an external processor, the product should not imply that all processing occurs solely inside the customer’s account. The architecture and applicable disclosure duties need to be assessed together.
Logs should preserve evidence without reproducing the file
Debugging often begins by printing inputs. That can turn a short-lived upload into a long-lived record accessible to more people. Error reports can be especially revealing because they include the exact data that broke a parser.
Define structured diagnostic fields. A job identifier, failure category, parser version and affected field can help support without copying the whole document. When a snippet is necessary, limit and protect it under the product’s policy.
Inspect both ordinary and error paths. A successful request may omit sensitive content while an exception handler includes it. A model response can repeat private input in a log even when the request itself was redacted.
Logging controls need tests and an owner. A promise that developers should avoid sensitive logs is weaker than a reviewable logging interface and restricted retention. The evidence retained for reliability should remain proportionate to its purpose.
Give temporary state an expiry and an owner
Uploads, extraction results, previews and conversation state can outlive their usefulness. Without a lifecycle, “temporary” becomes a label rather than a property.
Define when the state is created, what event permits removal and what happens if processing fails. A file awaiting human review may need longer retention than one already processed. The reason should be tied to the job rather than an arbitrary indefinite default.
A cleanup task should be repeatable and inspectable. It should identify which records were removed and which remain because of an unresolved obligation. A failed deletion should not disappear behind a successful scheduler invocation.
The durable workflow chapter explains retained progress. Privacy design needs that progress to include access revocation and expiry conditions. A persistent workflow should not continue publishing data after the customer’s permission for that operation ends.
Deletion is a system operation
The visible delete action should have a defined scope: original upload, extracted fields, generated report, search entries, pending jobs and relevant processor copies where the integration permits or requires it. The product should know which are affected.
Some records may need retention for legitimate obligations under applicable rules. Identify those conditions precisely rather than make an unconditional promise that cannot be met. The legal basis and customer role require appropriate jurisdiction-specific analysis; architecture should expose the actual data categories and retention needs for that analysis.
Deletion can be asynchronous. If so, the interface should distinguish requested, in progress, completed and failed states. A button click does not establish that every copy has disappeared. Retain enough evidence to verify completion without recreating the deleted content in the audit record.
Test a deletion request against safe records across the mapped services. Confirm that pending jobs cannot regenerate removed output and that revoked access remains revoked. A cleanup job that removes one table while another worker restores the content has not fulfilled the intended operation.
Backups require a recovery-aware policy
A backup may retain information that the live system removed. Restoration can reintroduce it. The recovery process should account for deletion and access changes that occurred after the backup point.
Keep a defined way to apply relevant lifecycle events during restoration. The exact mechanism depends on the storage design and applicable obligations. The important question is whether a recovered environment recreates data or permissions the product had already retired.
Rollback design treats restoration as a customer-obligation problem. Privacy is one of those obligations. A successful restore that brings back another customer’s deleted report can create a new failure while resolving availability.
Protect backup access and avoid casual downloads during support incidents. Copies created for a recovery exercise should have their own safe fixtures or controlled retention. Testing recovery should not distribute production customer data to environments that otherwise would not have it.
Separate product improvement from the original purpose
A customer uploads a document to receive a report. Using it later to train a model, build a solved-case corpus or publish benchmarks is a different design question. It needs an appropriate authority and purpose assessment rather than being assumed from the original upload.
The app can often learn from less sensitive signals: error categories, aggregate completion rates or carefully authorized sanitized cases. Those signals still need a defined collection and retention policy. “Anonymous” should not be used casually when combinations of fields can identify a customer.
The solved-problem corpus chapter examines the value and rights required for retained cases. This chapter establishes the boundary that lets that proposal be evaluated honestly. A data moat cannot depend on quietly repurposing customer information.
Keep improvement datasets traceable enough to honor their permitted use and lifecycle. If a source case must be removed or its permission changes, the team needs to know which derived records depend on it. A folder of unlabeled examples makes that responsibility hard to carry out.
Support access should be a product capability
Support staff may need to inspect a failure without accessing every customer file. Design a narrow diagnostic view with appropriate role checks, purpose and retained access evidence.
The customer can provide an authorized support reference or sanitized example. If broader access is necessary, establish the applicable permission and scope. A shared administrator account makes it difficult to know who accessed which data and why.
Support tools should preserve the same tenant boundary as customer routes. Searching by a familiar name should not reveal unrelated reports. Export features and attachments deserve review because they can move information outside the application’s usual access controls.
The self-diagnostics chapter shows how product instrumentation can reduce unnecessary support exposure. A clear failure category and actionable correction can let the customer solve a problem without sending the whole document to a human.
Privacy promises should survive ordinary change
A new model provider, analytics integration or support tool can create a new data route. Review the lifecycle map when architecture changes rather than only when the privacy page is rewritten.
The release checklist should identify new collection, disclosure, retention or access behavior. A small code change can have a large privacy consequence if it adds full request logging or broadens a search index. File size is not a measure of impact.
Keep the accepted privacy assumptions near the relevant implementation. A developer should be able to discover why a field is excluded from model context or why a trace expires. Otherwise an apparently helpful refactor can remove the control as unnecessary complexity.
The policy page should describe the implemented state accurately. If the product cannot meet a proposed promise, revise the design or the promise before launch. Publishing attractive wording first and planning to implement it later creates a mismatch customers cannot inspect.
Test the lifecycle with protected cases
A proposed privacy suite should follow one safe document through successful processing, failed processing, support inspection, deletion and restoration. It should also test unauthorized tenant access and a revoked user with a pending job.
Define expected outcomes independently of the current implementation. Inspect retained copies, not merely the interface message. If the suite cannot access a provider’s actual retained state, report that limit and verify the applicable contractual and product behavior through the available authorized evidence.
Retain configuration and results so a future change can reproduce the case. A passing local test supports the tested lifecycle under that setup. It does not certify legal compliance across every market or establish that a provider never retains another copy.
A real compliance assessment should connect this technical evidence with current applicable duties and contracts. The architecture map gives that assessment concrete facts; it is not a substitute for the assessment.
Make exports useful without broadening disclosure
A customer may need to leave the app with their records and results. Design a portable export that preserves meaning: identifiers, units, versions and the relationship between inputs and reports. A spreadsheet of unexplained numbers can be technically downloadable and practically unusable.
The export route needs authorization and scope checks. An organization administrator may be allowed to export all organization records, while an invited reviewer can inspect only one report. The application should enforce that distinction even if both users see an export button somewhere in the interface.
Exports create copies outside the app’s direct control. The product should state what it can and cannot delete after a customer downloads a file, while protecting the download route and avoiding unnecessary exposure in URLs or logs. A temporary export artifact needs its own expiry and access policy.
A safe test can create two tenant fixtures, request an authorized export and confirm that the other tenant’s records remain absent. It can also inspect whether units and source relationships survive the format. This tests both privacy and usefulness without using real customer information.
Portable records can reduce dependence on the app and support an honest customer exit. That is compatible with a commercial product: retention should come from continuing value rather than making customer data difficult to retrieve. The retention chapter examines that relationship; privacy design supplies the actual export behavior it relies on.
What Would We Do at Salars?
A proposed Forge template would start with safe fixtures, narrow tenant access and a documented data path. Supplier Margin Guard would request only the information its diagnostic needs. Full customer archives would not become default model context or generic logs.
Before a customer pilot, the team would inspect collection, external processing, support access, retention and deletion. It would test a safe deletion across all mapped copies and a restore after deletion. Any unverified processor behavior would remain an explicit limitation rather than a blanket privacy claim.
Original merchant records would not automatically enter Opportunity Radar, a solved-case corpus or a public benchmark. Those uses would require their own purpose, authority and lifecycle design. The initial customer job would remain the boundary for the data it supplies.
Privacy becomes maintainable when the product can explain where a record goes and what ends its use. That answer should remain true after the upload succeeds, after it fails and after the system recovers.
Sources
- ICO: guide to data protection principles, read October 7, 2026, for UK GDPR principles and the current under-review notice. Jurisdiction-specific guidance does not establish global compliance.
- Salars: what can I share with an AI tool?, for the related user-input decision; this chapter develops app lifecycle architecture.
- Salars AI library, for related security and reliability guides. Document app designs, deletion tests and Forge practices are proposed or hypothetical.
Loading comments…