AI · Article

Measure Whether Digital Intelligence Improves Decisions

A practical scorecard for a knowledge and AI workflow: track decision quality, correction cost, time, exceptions, and who benefits before expanding it.

A system can answer in seconds and still make an organization slower. Someone checks every answer against an old spreadsheet, corrects a third of the drafts, and calls customers after a promise turns out to be wrong. The dashboard records “50 AI responses.” The people doing the work remember the cleanup.

Measure whether a knowledge system improves the decision and the complete path to its outcome. Count the time it takes to gather, review, act, and correct. Check the accuracy and usefulness of decisions, not only the fluency of answers. Watch exceptions and effects on the people who receive the result. Then use the evidence to keep, change, narrow, or stop the workflow.

This is the final article in the Building Digital Intelligence series. The opening model defined an intelligence loop; the previous article showed how to put knowledge into a reviewed workflow. Measurement closes that loop by asking what actually happened.

Define the result before choosing a metric

If the job is to give customers reliable repair dates, a faster draft is useful only if the date is supportable and customers receive a clear update. If the job is to publish trustworthy articles, producing more pages does not show that claims are accurate or readers can use them. Begin with one sentence about the intended outcome:

Customers receive a truthful repair status and a defensible completion date when the required evidence is available; uncertain cases receive a timely update instead of a made-up promise.

Now decide what would count as success and failure. A successful routine case might include a correct order or job match, current evidence, an approved message, and no later correction due to the team’s error. An exception may also be a success if the workflow recognizes uncertainty and routes the case to the right person. A system that refuses to guess can look slower on a superficial speed chart while protecting the relationship that matters.

Name whose outcome you are measuring. The business may save staff time, while a customer waits longer for a useful answer. A manager may see fewer escalations because staff stopped reporting them. Include the recipient’s experience and the reviewer’s work, not only the operator’s preferred number.

Build a baseline from actual cases

Before changing the process, record a small sample of recent cases. For each, note the type of request, evidence needed, time spent finding it, time spent deciding and communicating, corrections, exceptions, and known outcome. Avoid collecting private details that are unnecessary for measurement. If the cases differ greatly, group them by type rather than averaging unlike work together.

A baseline might use 20 repair-status requests from one month. The sample can reveal where staff lose time and which facts are often missing. It is not a reliable estimate of a future year’s performance, and it may not represent holiday demand, unusual repairs, or new suppliers. Keep the period, sample size, and selection method alongside every reported number.

Do not reconstruct a precise baseline from memory after launching a tool. People remember unusual failures and recent successes unevenly. If past records are insufficient, start measuring the current process for a short period before changing it. That delay can save the team from celebrating an improvement it cannot substantiate.

Use a balanced scorecard

One metric invites a shortcut. If you reward fast answers alone, the process may produce fast guesses. If you reward zero errors alone, it may escalate everything. Combine a small number of measures that capture the task, cost, and limits.

Measure What it reveals How it can mislead
Decision correctness on reviewed cases Whether the answer matched the applicable evidence and rule A convenient sample may exclude hard cases
Complete-case time Time from trigger through review, action, and correction A faster average may hide severe outliers
Correction and recontact rate How often work needed repair after the first action Some errors are never reported back
Appropriate escalation Whether unusual cases reach a person Fewer escalations may mean missed exceptions
Source freshness and coverage Whether needed records exist and are current A filled field can still be wrong
Recipient outcome Whether the customer or reader got a usable answer Satisfaction surveys can be biased toward respondents
Total operating cost Tool fees plus setup, review, upkeep, and incident work Estimates omit work done outside the tracked system

The UK Government Data Quality Framework distinguishes completeness, accuracy, consistency, validity, uniqueness, and timeliness. Use the dimensions that matter to the decision. An inventory system may need a very current count; a historical research page needs accurate source attribution and clear dating. A single “data quality” score can obscure the problem that is actually harming decisions.

The NIST AI RMF Measure guidance encourages measuring risks, human oversight, policy exceptions, and accountability for go/no-go choices. Its playbook is voluntary; this scorecard is a small-team adaptation, not a claim of NIST endorsement or compliance.

Count the complete cost of a case

Separate the model’s response time from the time required to finish the job. For each case, count locating source material, preparing input, waiting for the tool, reading output, checking claims, making corrections, obtaining approval, taking the action, and handling later questions or errors. Add setup and ongoing maintenance over a sensible period.

Imagine a fictional shop handling 40 routine messages a week. Before a new workflow, a typical message takes about eight minutes from opening to recorded reply, or roughly 320 minutes for the 40 cases. Afterward, the tool drafts in one minute, but staff spend four minutes finding and checking sources, two minutes correcting or approving, and an average of one minute per case on system upkeep and later repairs. The complete process is about eight minutes per case again. The draft became faster; the job did not.

That example is arithmetic, not a benchmark or forecast. A real shop would need actual case observations, tool costs, and error records. Different messages may take different times, so a mean alone is weak evidence. Report the number of cases and the spread: How many were fast, how many were unusually slow, and why? The site’s full-cost guide explains the comparison in more detail.

Check quality with cases a person can inspect

For a workflow that produces an answer or recommendation, collect a small review set with expected outcomes. Include normal cases, missing data, conflicting sources, changed policies, and an unanswerable question. A knowledgeable reviewer should check both the retrieved evidence and the final answer. A model may use the wrong source; it may also receive the right source and still omit a critical exception.

Microsoft’s RAG evaluation documentation distinguishes retrieval quality, groundedness, relevance, and completeness. That separation is useful beyond RAG products. If an answer fails, identify whether the source was missing, the search missed it, the model misread it, or the decision rule was ambiguous. Each cause needs a different fix.

Keep the exact source version and date with a failed case. Otherwise, a later retest may pass only because the underlying document changed, leaving the team unsure what fixed the original failure. Make the review repeatable enough that another person can inspect it, while avoiding an elaborate benchmark that costs more than the decision it serves.

Watch exceptions and quiet harms

A good workflow should not pretend every case is routine. Count how often it escalates and inspect whether those escalations were appropriate. Review some cases it did not escalate; a system can miss an exception and make its dashboard look efficient. Ask users and recipients whether the result was understandable, respectful, and correctable.

Different mistakes have different costs. A typo in an internal draft is not equivalent to exposing private information, making a false safety claim, or sending an irreversible commitment. Set explicit stop conditions for severe incidents. Record what happened, who was affected, whether an action needs reversal, and what source or process must change. If a tool’s use shifts more burden onto customers or staff, the outcome is not captured by a narrow time-saving claim.

Check whether the system is helpful for the range of people it serves. A workflow trained on tidy English emails may fail on incomplete messages, translated text, or accessibility needs. A knowledge search built around management vocabulary may miss the words frontline staff use. Test representative language and formats with consent and appropriate privacy controls. Do not declare the system broadly effective after a few easy cases.

Compare a change without fooling yourself

When you revise a workflow, write down the change and the expected effect. Perhaps you add an effective-date label to policies. The hypothesis is that staff will choose the applicable version more often for older orders. Test a set of such orders before and after the change, keeping the cases and review criteria as comparable as practical. Also check whether routine current orders became harder to answer.

Avoid changing the source library, prompt, approval screen, and tool at the same time if the goal is to learn which change helped. Sometimes operational needs force several fixes at once; record that limitation instead of assigning all improvement to one feature. If case volume is small, use careful case review and staff feedback rather than presenting a statistically precise percentage from a handful of observations.

A simple log can include: date, workflow version, case type, expected decision, actual decision, reviewer correction, full time, and outcome. Aggregate only what the organization needs. Protect the underlying records and set a retention period appropriate to their sensitivity and purpose.

Make a keep, revise, or stop decision

At a scheduled review, look at the balanced evidence rather than a single headline number:

Record who made the choice and what evidence they used. If the workflow expands to a new customer group, action, or source collection, treat that as a new use case requiring its own review. A successful internal draft assistant does not automatically justify an unsupervised public answer service.

Digital intelligence is not a growing count of documents, prompts, or automations. It is a repeatable ability to make a decision from evidence, act within authority, notice when reality disagrees, and improve the next decision. The useful final test is plain: can the team show one meaningful choice that is now more accurate, explainable, timely, or humane at a cost it understands? If it cannot, the next step is to investigate the work, not to scale the system.