AI · Article 45 of 72 · Part 10

AI Software Has Real Cost of Goods Sold

Build a cost ledger covering inference, tools, runtime, storage, human exceptions and vendor minimums.

Variable costs must be deducted from collected revenue. Collected revenue. Then Less runtime and usage. Then Less payment and variable support. Then Contribution before fixed costs.
Define each cost consistently; revenue growth can coexist with declining contribution.

The app charges $40 for a monthly plan. Its model call appears to cost only a few cents, so the founder expects a comfortable margin. Then a large input triggers several calls, a timeout triggers retries, a support request requires manual cleanup and stored files continue accumulating after processing.

The original estimate priced one operation inside delivery, not the service the customer bought. Build a delivery-cost ledger around accepted customer work, including attempts, tools, runtime, storage and attributable human exceptions. Keep the operating ledger’s classification explicit; formal accounting treatment requires the business’s actual policies and professional review.

This chapter of The AI Software Factory owns cost attribution. Pricing architecture defines the commercial unit. Contribution margin combines retained revenue with attributable costs. The cost ledger supplies the evidence those decisions require.

Define the cost object before collecting bills

A cost object is the unit whose delivery the team wants to understand: a job, account, cohort or product. A vendor invoice describes purchased resources, which may span many such units.

For a hypothetical report app, record each job under an account and product identity. The ledger can then aggregate per accepted report, per account-month and per source format. Those views answer different questions without pretending the invoice itself is a customer-level measurement.

Define an accepted report separately from an attempt. A failed attempt consumes resources even if the customer is not billed. Removing failures from the ledger makes cost per successful result look artificially low.

Choose a period that matches the comparison. Revenue from one month and costs from a different usage period can create misleading margins. Record event time, invoice period and any delayed adjustment so reconciliation remains possible.

Separate variable, attributable and shared costs

A model call tied to one job is directly traceable. Account-level storage may be attributable over time. A shared monitoring subscription may support every product. These costs need different treatment.

Start with traceable facts and state any allocation rule. Do not force every shared expense into a precise per-job number merely because a spreadsheet needs a cell. An honest shared-cost pool is better than false precision.

An allocation can still aid decisions if its purpose is explicit. Dividing a shared service by active accounts may help evaluate a portfolio, but it may not describe the marginal cost of adding one small account. Keep that distinction visible.

Formal cost of goods sold and operating-expense classifications can differ from a founder’s decision ledger. Use a clearly named operating boundary and reconcile it with the accounting view rather than claiming one universal treatment for every software business.

Record inference beyond the happy path

Inference cost can include input processing, generated output, repeated reasoning, classification, verification and corrections. The cost of a single successful call may represent only a small part of the accepted job.

Record the model, relevant usage quantities and provider cost attribution available for each attempt. Keep current vendor schedules separately dated; they can change. This article does not quote current model prices or claim a particular vendor is cheapest.

Large inputs and long outputs deserve separate observation. A template that encourages verbose explanations can increase delivery cost without improving the accepted result. A smaller output can be better only if it preserves the customer promise.

Model routing addresses choosing models for development work. Product inference needs its own acceptance evidence: replacing a model in a live service can change result quality, retries and human review as well as the unit rate.

Attribute tools and external services

An AI job may query a search provider, extract a document, call a database service or use another paid API. These costs should be recorded under the operation that created them when the evidence permits.

Separate exploratory tool activity from necessary delivery. An agent that repeatedly searches without a stopping rule can create a long tail of cost. A task budget and explicit completion criteria help bound that behavior.

Tool failures also matter. A paid lookup that returns unusable information still contributes to delivery cost. Whether the customer should be charged is a separate commercial policy, not a reason to erase the expense.

Keep the external service’s billing unit distinct from the product’s customer unit. A merchant buys an accepted report; the provider may buy queries and extraction pages. The ledger should explain the mapping rather than expose all internal units on the price page.

Include runtime and coordination resources

Compute, storage operations, queues, durable state and network transfer can contribute to delivery. Their relative importance depends on architecture and workload; do not assume inference dominates every product.

Cloudflare’s agent platform illustrates a runtime architecture with durable coordination. Its availability is not a cost estimate for the specific app. Collect actual usage under the chosen architecture and account configuration.

Distinguish active processing from waiting and persistent state. A long workflow may spend most elapsed time awaiting input while retaining data. A short burst may consume substantial resources. Wall-clock time alone is not a reliable cost proxy.

Preserve limits and resource budgets as part of the operating design. If a malformed job can create unbounded work, the product has both a reliability and an economics problem. A cost ledger can reveal the pattern, while the runtime boundary needs to prevent recurrence.

Treat retries as a delivery mechanism with a cost

Retries can recover transient failures, but each attempt can consume resources. Record the cause, attempt count and final outcome rather than treating retry work as invisible overhead.

A failed acknowledgment can create a particularly expensive ambiguity. Repeating a completed operation may duplicate both cost and effect. Idempotent agents addresses safety; the ledger should also distinguish a legitimate recovery attempt from a duplicate job.

Set a bounded retry policy appropriate to the operation. A dozen futile retries against a permanent validation error do not improve reliability. The system should recognize the failure class and request correction or stop.

Review cost per accepted result alongside failure rate. A cheaper attempt that fails more often may create a more expensive service. The decision concerns completed useful work, not an isolated provider rate.

Trace storage through the data lifecycle

Raw files, normalized records, results, logs and backups can have different retention periods. Cost continues after the initial job when those artifacts remain stored or accessed.

Record which data must persist to deliver the promised service. Privacy design asks the same lifecycle question from a data boundary. Keeping every file forever can create both cost and privacy obligations without improving customer value.

Account-level storage can be measured over time or allocated using an explicit rule. Large dormant accounts may create a different pattern from small active accounts. Avoid attributing every stored byte only to the month in which it arrived if the service retains it for much longer.

Deletion and export can also consume resources. Include those operations when they are meaningful to delivery. They are part of the customer relationship, not exotic events that can be ignored because most demos never reach them.

Human exceptions belong in the economics

A founder may manually repair a file, interpret a result or answer an integration question. If that work is required to complete the promised job, it affects delivery even when no separate wage payment occurs.

Record time and task category rather than relying on memory. Distinguish product-error repair, supported setup help and bespoke consulting. These categories suggest different responses: fix the product, improve onboarding or create a separate service boundary.

Do not automatically convert owner time into a precise cash expense without explaining the assumption. A decision ledger can show minutes separately and use an explicit valuation for sensitivity. Profit per human hour gives the time denominator its own treatment.

An average can conceal concentration. If a small cohort produces most manual exceptions, inspect its source formats and expectations. A narrow supported scope may solve the problem more effectively than a general price increase across all accounts.

Work through a hypothetical delivery ledger

Suppose one accepted report requires $0.60 of inference, $0.30 of extraction, $0.20 of runtime and $0.10 of attributable storage and monitoring. Direct measured-resource cost in this illustrative example is $1.20.

Now suppose one in ten reports needs six minutes of human correction, valued for planning at $30 per hour. Each such exception represents $3 of time value, or $0.30 per report averaged across ten reports. The illustrative combined delivery figure becomes $1.50.

These numbers are hypothetical and are not vendor rates or Salars observations. The example shows why the exception rate and its valuation should remain visible. If exceptions rise to one in three, the same time assumption adds about $1 per report instead of $0.30.

Keep cash resource cost and imputed time value as separate columns. Combining them can aid a decision, but separating them prevents a reader from mistaking an owner-time valuation for an invoice payment.

Reconcile vendor minimums and included allowances

A vendor may charge a base amount, include usage or provide credits. The marginal cost of one more job can differ from the amount paid that month. Both views can be useful if labeled.

Record the purchased package and actual usage. If a service minimum is paid regardless of current volume, show it as a shared or product-level obligation under the chosen operating boundary. Do not make unused capacity disappear from total cost.

An included allowance can temporarily make measured marginal usage look free. That does not establish that growth remains free after the allowance is exhausted. Model the next relevant threshold and recheck actual vendor terms before implementation.

Credits and promotional rates also need dates. A pilot supported by temporary credits should not be presented as a steady-state cost model. Preserve the undiscounted scenario separately when evaluating the ongoing offer.

Match billing records to resource records

The customer meter and resource ledger should reconcile through stable job identities. They may count different events because a provider retry can incur cost without creating a new customer charge.

Stripe’s usage-billing overview describes current product options for metered billing. It does not determine which internal resource events your business should bill. Decide that commercial policy first, then implement the appropriate meter.

Inspect duplicates, corrections and delayed events. A resource expense may arrive after the customer billing period closes. A credited job may remain in the cost ledger. An account merge can change identity relationships without changing the underlying delivery history.

Reconciliation should explain the difference rather than force artificial equality. Revenue events and cost events describe related but distinct obligations. Keeping their mapping explicit helps investigate unusual margins and billing disputes.

Use distributions instead of one comforting average

Average delivery cost is useful but incomplete. Record median, expensive cases and the account patterns that create them when the sample supports such comparisons. A few extreme jobs may determine the exposure of an unlimited plan.

Group by relevant mechanisms such as file size, source format, retry count or human exception category. Avoid slicing tiny samples into impressive-looking statistics. The purpose is a decision about scope or architecture, not a dashboard filled with unstable percentages.

Include failed and unsupported attempts. Their cost may be reduced through earlier validation. A report that never reaches acceptance still consumes the service’s resources and the customer’s time.

Revalidate after material changes. A new model, parser, integration or retention policy can alter the distribution. A cost figure collected under the previous version should retain its date and scope rather than silently become the current truth.

Choose a response that fits the cost mechanism

If verbose output drives inference cost, test a shorter accepted result. If repeated extraction drives expense, inspect caching within the data and freshness boundary. If unsupported files drive human cleanup, narrow the offer or improve readiness validation.

Preserve result quality and authority as independent requirements. A cheaper path that produces an unreliable recommendation or exposes private data is not a successful optimization. Use protected cases from the acceptance suite.

Set a bounded experiment and stopping condition. Compare the chosen cost object under comparable work and record any quality or exception change. A proposed saving remains a scenario until the evidence is collected.

The useful result is a scoped finding: this change reduced a defined cost under tested conditions while preserving the acceptance criteria. It is not a universal claim that the architecture is cheaper for every workload.

Account for idle obligations between customer jobs

An app may keep a connection, scheduled check or retained workflow active even when the customer submits no new job. Attribute those obligations at the account or product level instead of forcing them into nonexistent reports.

A hypothetical monitoring account can create regular source checks, storage access and alert evaluation while producing few visible exceptions. Low output volume does not necessarily mean low delivery cost. The product may be doing useful background work whose evidence belongs in a separate activity record.

Conversely, idle capacity purchased for future growth is not evidence that each current customer directly caused its full cost. Label the allocation and inspect the marginal change separately. That distinction helps the founder decide whether to add accounts, change the architecture or reduce an unnecessary fixed obligation.

Keep development expense separate from recurring delivery

Building a parser, researching a model and writing onboarding instructions can require substantial work before the first customer. Those investments matter to the business, but they answer a different question from the cost of delivering the next accepted report.

Maintain a development and maintenance record alongside the delivery ledger. A product may have favorable recurring contribution while requiring so much maintenance that it remains unattractive overall. The contribution calculation should not conceal that larger obligation.

When a maintenance task is caused by a particular customer customization, identify that relationship where evidence permits. It may belong to a separate service boundary rather than an assumed shared product investment. Preserve uncertainty instead of assigning every engineering hour to the nearest account merely to complete the spreadsheet.

What Would We Do at Salars?

For a proposed merchant report app, we would give each job an account, source-format and version identity. The ledger would record attempts, inference, extraction, runtime, storage and any manual correction separately from customer billing.

We would start with hypothetical planning assumptions, then replace them with measured pilot evidence. Cash costs and owner-time valuation would remain distinct. We would reconcile vendor invoices and note minimums, included allowances and temporary credits rather than treating them as permanent free delivery.

A proposed weekly review would inspect ordinary accepted reports, failed jobs and the expensive tail. It would select one mechanism to improve and preserve acceptance quality, privacy and safe action boundaries. This article does not report an executed Salars cost experiment or an existing app margin.

If most cost came from custom unsupported work, we would reconsider the product scope. If a recurring resource dominated, we would investigate its actual usage before changing price. The ledger would make the decision explainable rather than simply justify a number already chosen.

AI software has a delivery system behind every useful customer result. Making that system’s cost visible is a prerequisite for a sustainable offer. Explore the wider AI collection for the pricing and operating decisions built on that evidence.

Sources

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home