An AI project can look successful in a demonstration and disappointing in the monthly accounts. The demonstration measures how quickly an answer appears. The business pays for the accepted work, including the time spent checking, correcting, coordinating, and recovering from mistakes.
The gap between those measurements is where many claims about leverage become unreliable. A tool may genuinely make one step faster while making the complete process only slightly better—or worse.
How can we estimate net AI leverage after review, failure, and coordination costs? We need a calculation that starts with the outcome, distinguishes different kinds of benefit, and includes the whole process. The equation in this article is a proposed operating framework. It is not an established empirical law, an investment-return guarantee, or a replacement for financial records.
Begin with the unit of accepted work
A useful unit might be an accepted estimate, a resolved support case, a reconciled record, or a product description that accurately represents the goods. The unit should correspond to something a person or customer actually needs.
A draft is a different unit. It may help create the accepted estimate, but it does not necessarily complete the work. Counting drafts as completed outcomes hides the remaining review. Counting generated words as productivity can reward output that makes the process more burdensome.
Define the quality threshold before measuring the new workflow. For an estimate, the threshold could include current costs, feasible capacity, clear exclusions, and no unsupported promise. For a support case, it could include actual resolution rather than a polite reply that leaves the problem unanswered.
The old and new processes should be compared against the same threshold. Otherwise a faster process may appear better because it quietly performs a smaller or lower-quality task. Changing the threshold can be legitimate, but it should be explicit and assessed as a different service.
This discipline makes the calculation understandable. We know what was produced, why it counted, and whether the comparison is fair.
A time equation that includes the missing work
For a defined set of cases, net time released equals the old process’s total time minus the new process’s total time. The new process includes preparation, prompting or setup during use, review, correction, coordination, maintenance allocated to the period, and incident handling.
Written compactly:
Net time released = baseline task time − assisted task time − added review and correction − added coordination − allocated maintenance − incident time.
The terms must not overlap. If assisted task time already includes review, do not subtract review a second time. If maintenance is measured monthly, allocate it consistently rather than charging the whole month to every case.
The baseline also needs complete time. A manual process may already involve review and correction. The calculation should compare complete processes rather than assume the old one was perfect. The purpose is accuracy, not a presumption for or against the new tool.
A positive result means capacity has been released under the measured conditions. It does not by itself establish additional cash or revenue. The next question is what the organization can and wants to do with that capacity.
A cash equation answers a different question
Net cash contribution from the change equals additional realized contribution plus actual avoided cash costs, minus software, integration, additional paid work, correction costs, and other attributable expenditures. Contribution means revenue after the relevant variable costs, not the full value of an order.
A useful compact version is:
Net cash contribution = additional realized contribution + actual avoided expenditure − attributable cash costs.
This is deliberately separate from the time equation. An owner’s saved hour is not a cash saving unless a payment actually changes. An additional order is not pure benefit if fulfilling it requires goods, shipping, fees, or labor.
Attribution remains difficult. If sales increase while the business changes its offer, season, and channel, the entire increase cannot simply be credited to AI. A credible estimate explains the comparison and retains uncertainty where the causes cannot be separated.
The calculation can use a range rather than a single number. A conservative case may assume no additional sales, a central case may use observed accepted work, and an optimistic scenario may explore possible capacity use. The scenario must remain labelled as a scenario.
Worked example: a weekly estimate workflow
Consider an explicitly hypothetical business preparing twenty estimates each week. The baseline takes thirty minutes per accepted estimate, or ten hours. The assisted process takes eight minutes to prepare and ten minutes to review each estimate, or six hours in total.
The initial comparison releases four hours. Now add forty-five minutes of weekly maintenance, thirty minutes spent resolving exceptions, and fifteen minutes of additional coordination. Net time released falls to two and a half hours.
Suppose the software and related usage cost thirty dollars per week. The owner values the released capacity, but the business has not reduced a payment or received additional revenue. The measured result is two and a half hours of capacity for thirty dollars, under the stated assumptions. Reporting a cash profit would go beyond the evidence.
Now suppose the business uses that capacity to complete one additional job that contributes eighty dollars after its relevant variable costs. If the additional job would otherwise have been declined, and if the attribution is credible, the weekly cash contribution from the change could be fifty dollars after the software cost. The capacity and the contribution describe related aspects of the same change; they should not be added as independent gains.
Change the assumption again. If the business was already able to fulfill that job, the additional contribution may not belong to the AI workflow. The time saving could still be useful, but the cash claim needs a different basis.
The arithmetic is simple because the difficult work is defining the terms honestly.
Quality belongs beside the equation
Time and cash do not fully describe whether a process improved. An estimate can be faster and less accurate. A support response can be cheaper and less helpful. The quality threshold protects the calculation from these substitutions.
Track measures tied to the task: correction rate, unresolved cases, inaccurate promises, customer reopening, or another observable failure. Avoid using only a satisfaction score if it can reward reassuring language that conceals an unresolved problem.
Some quality changes can be valued in money when the evidence supports it. A reduction in actual refunds may be a financial benefit. A lower risk of a rare incident is harder to estimate. It may be better represented as a constraint or scenario than as a precise expected dollar return.
This is especially important for severe consequences. A business may reject a workflow even if its average return looks positive because one unsupported action could create an unacceptable obligation. The decision includes a boundary about what the organization is willing to risk.
The Permission Architecture addresses how that boundary can be enforced. Here it is enough to recognize that a positive average does not automatically justify every level of authority.
Why published productivity results cannot fill in your equation
The revised author version of Generative AI at Work reports a 15% average increase in issues resolved per hour among 5,172 support workers in its setting, with substantial differences across workers. That is a measured result for a particular deployment. It is not a plug-in estimate of your net cash return.
Your task, records, review requirements, and bottlenecks may differ. The published result can justify investigating a workflow. It cannot supply the local baseline or the cost of maintaining your integration.
The distinction is particularly important when an advertised benchmark concerns a step rather than the complete process. Faster drafting may coexist with longer review. More resolved cases per hour may not create revenue if demand is fixed, though it may improve service or reduce workload.
A useful local estimate therefore begins with your own accepted unit of work. Published evidence remains context, including its population, version, and limitations. The equation does not become more credible by borrowing a precise percentage from a different environment.
Failure costs need an explicit treatment
Frequent small failures can be measured directly. If a workflow misclassifies several cases each week, record the correction effort and any downstream effects. They belong in the recurring cost rather than an invisible category called “edge cases.”
Rare large failures require more care. A short trial may observe none without establishing that the risk is negligible. A process with permission to make an expensive purchase has a different consequence profile from a process that prepares an internal draft.
One practical response is to limit the exposure before trying to estimate it precisely. Spending caps, scoped actions, staged approvals, and an immediate stop mechanism can make the worst supported consequence smaller. The system can then be evaluated within that bounded scope.
A second response is scenario analysis. Ask what happens if a source is stale, a tool repeats an action, or a record belongs to the wrong customer. Estimate the recovery work where possible and identify consequences that remain uncertain. The scenario is not a predicted event rate; it is a way to inspect the design.
The later article on The Dangers of AI Leverage follows these effects at larger scale. In the equation, they remind us that throughput and exposure can grow together.
Coordination is often the hidden term
A workflow can reduce one person’s task time while increasing another person’s interruptions. If the new system produces many escalations, draft approvals, or ambiguous requests, a manager may become the new bottleneck.
Record the handoff work. How many decisions does the system ask for? Are they well prepared? Does the recipient have the evidence needed to answer? Are several requests about the same unresolved rule? These questions make coordination cost visible.
An escalation that clearly states the missing information can be efficient. An escalation that says only “please review” may force the person to reconstruct the whole task. The same number of escalations can therefore represent very different burdens.
Recurring coordination problems may be repairable. A clearer rule can resolve a class of ordinary cases. A better source record can remove repeated questions. An unnecessary approval can be eliminated where standing authorization already supports the action. The goal is appropriate control with useful handoffs.
Do not treat every minute of supervision as waste. Some attention is the service the customer actually needs. The equation should reveal the work, allowing the owner to decide which costs are productive and which are avoidable.
Setup should have a recovery plan, not a story
A project may require upfront work before it produces recurring benefit. Record that investment separately: implementation, source cleanup, evaluation, training, and the effort of moving from the old process.
A simple recovery estimate compares setup cost with a conservative recurring net benefit. If the recurring benefit is uncertain, show the range. If the estimated recovery depends on an unobserved sales increase, state that dependency plainly.
The estimate should also have a stopping condition. For example, the business may stop after a defined trial if accepted work has not improved, correction costs exceed the baseline, or the process requires authority that the owner cannot safely grant.
This prevents setup effort from creating an obligation to continue. Money or time already spent cannot establish that the next month is worth funding. The decision should depend on the expected future benefit and the evidence now available.
AI Can Make Bad Businesses Fail Faster examines the wider version of this problem. A faster workflow does not repair an offer that lacks demand or a business whose basic unit economics are negative.
Compare alternatives, including doing less
The relevant alternative is not always the existing manual process. A simpler template, a clearer form, or eliminating an unnecessary task may create more value than an AI integration.
Suppose staff spend hours summarizing meetings nobody uses. An assistant might reduce that effort. But the best change may be a short decision log instead of a lengthy summary. The organization should compare the AI workflow with that simpler alternative.
Similarly, a classification task might be solved by a required field at intake. A scheduling issue might be solved by an authoritative availability record. A model can help design these improvements without becoming a permanent part of the solution.
This is why the baseline should include a credible low-complexity option. The new technology must earn its place against the best practical alternative, not merely against an unnecessarily cumbersome process.
The leverage equation becomes a decision aid when it permits that comparison. It becomes a sales pitch when every term is arranged to make the new tool win.
Match the observation period to the claim
A short trial can establish that a draft process works on the cases it encountered. It may not establish that maintenance remains inexpensive through a supplier change, that customer benefits persist, or that a rare exception is handled correctly. The observation period should match the claim being made.
For a seasonal workflow, compare similar periods or state the limitation. For a newly introduced process, distinguish initial learning from steady operation. A team may become faster as it learns, while another process may become slower as accumulated exceptions appear. Neither pattern is visible in a single demonstration.
Keep the evaluation cases separate from the examples used to tune the instructions. Repeatedly adjusting a system until it handles the same familiar cases can create confidence without showing that it handles new work. A small set of protected cases offers a more useful check on whether the improvement transfers within the supported scope.
The point is proportional evidence. A reversible internal draft needs less than a workflow that makes external commitments, but each claim should remain no broader than the observation supports.
AI Leverage in Practice
Choose one workflow and define its accepted output. Measure several representative baseline cases, including correction and coordination. Record the quality threshold and the conditions under which the task should stop or escalate.
Run a bounded assisted version on comparable cases. Track preparation, review, correction, handoffs, maintenance, and incidents. Keep time and cash results in separate columns. Mark hypothetical capacity use separately from realized revenue or avoided expenditure.
Use a short summary that shows the assumptions. A reader should be able to change review time, case volume, or maintenance cost and understand how the result changes. If one assumption controls the conclusion, make it prominent.
Today’s tools can help collect records, prepare comparisons, and explain the calculation. Use deterministic arithmetic for the calculation itself and authoritative records for actual financial results. A future system might automate more of the measurement, but its attribution and source quality would still require inspection.
Keep the workflow only if the complete result supports the intended benefit. A smaller useful scope is a success. So is discovering that a simpler process wins before a large integration creates unnecessary cost.
A calculation that disciplines the promise
The AI leverage equation is valuable because it makes the promise concrete. It asks what improved, which costs changed, who did the remaining work, and whether the benefit was capacity, cash, quality, or resilience.
It also makes uncertainty visible. A measured time saving can coexist with an unproven revenue scenario. A positive average can coexist with an unacceptable failure mode. A useful tool can coexist with a poor integration.
Those distinctions are the foundation for responsible investment in The Age of AI Leverage. The next article asks what happens after time is actually released: when does that time become investable capital, rather than simply a larger space for more work?
Sources
- Erik Brynjolfsson, Danielle Li, and Lindsey Raymond, Generative AI at Work, revised author version, November 2024, subsequently published in 2025; a setting-specific productivity result, distinct from earlier versions and from net business profit.
Loading comments…