A worker used to prepare every customer quote personally. Now a system gathers requirements, estimates costs, drafts the quote, and flags unusual requests. The worker reviews a smaller number of cases and sends the approved offers.
The visible work has changed. The obligation has not disappeared. Someone must still know whether the estimate is sound, whether the business can honor the promise, and what happens when an error reaches a customer.
What changes when a person becomes responsible for directing automated work? They need to govern a process rather than merely complete its individual tasks. That means choosing standards, assigning decision rights, maintaining competence, detecting failures, and being able to intervene.
The transition can increase useful capacity. It can also leave responsibility attached to a person who lacks the information, time, or authority needed to exercise it. Calling that person a reviewer does not solve the problem.
Execution and governance are different jobs
Execution asks how to complete a task. Governance asks whether the task should happen, what counts as acceptable, who may decide, and how exceptions should be handled.
Consider an automated purchasing assistant. Execution includes finding suppliers, comparing prices, and preparing an order. Governance includes setting the budget, determining acceptable materials, deciding which suppliers meet obligations, and assigning authority to commit money.
A cheaper recommendation can be a bad purchase if it fails a requirement the system did not understand. A technically correct order can still be unauthorized. A low error rate can hide a rare mistake with consequences large enough to threaten the business.
These differences explain why moving a person out of direct execution does not automatically reduce the skill needed. The person may have to understand a broader process and judge situations that the automated system cannot settle within its assigned rules.
The technical controls belong in the agent permission architecture. This article concerns the human role around those controls: who has the knowledge, authority, and obligation to make them meaningful.
Start with the decisions, not the job title
Titles such as supervisor, governor, operator, or human in the loop can conceal different arrangements. A useful design begins by listing the decisions that actually occur.
For each consequential decision, identify who proposes it, who approves it, who can challenge it, and who deals with its effects. The same person may hold several roles in a small operation, but the roles should remain distinguishable.
A model might propose a refund. A policy might permit automatic refunds below a narrow limit. The owner might handle unusual circumstances. A customer should have a way to contest an incorrect result. The accounting record should show what happened rather than merely what the model intended.
That arrangement is more informative than saying that a human supervises the system. It shows where discretion exists and where it has been reserved. It also reveals gaps, such as a customer appeal with no responsible recipient or a financial action that no one can reconcile.
The map should include routine decisions as well as dramatic ones. Small recurring choices can shape customer treatment, supplier relationships, and cumulative spending. Governing only the exceptional case may leave most of the system’s actual behavior unexamined.
Responsibility needs matching authority
Someone held responsible for a process must be able to change or stop it. Otherwise the role becomes a place to assign blame after failure.
Imagine an employee asked to review generated lending explanations. They cannot see the information used, cannot alter the decision, and are expected to approve each case in seconds. A review checkbox exists, but meaningful judgment has little room to operate.
A better arrangement would make the relevant information accessible, define what the reviewer can question, provide enough time for the consequential cases, and offer a route to someone with authority to resolve the problem. The exact requirements depend on the setting; a small commercial workflow and a regulated decision system require different controls.
The general principle is simple. Responsibility, information, resources, and authority should fit together. If one is missing, the organization should acknowledge the gap rather than claim that the mere presence of a person has repaired it.
This is an institutional design argument. The OECD’s work on governing with AI provides public-sector context for accountability and oversight. It should not be read as evidence that a private business can copy one universal supervisory arrangement and become safe.
The governor must understand the work
Direct task experience can help someone recognize implausible outputs. A person who has prepared quotes understands that a missing specification can change the estimate. A person who has reconciled invoices notices when a summary mixes billed revenue with cash received.
Automation can reduce opportunities to maintain that experience. If every ordinary case is handled elsewhere, the human may eventually see only rare exceptions. Those exceptions can be the most difficult cases to judge.
A practical response is to preserve selected contact with ordinary work. Review samples from the full population, occasionally perform a task manually, and compare the system’s process with an independent method where the consequence justifies the effort.
The point is not to keep humans busy for appearances. It is to maintain the competence required for oversight. The appropriate amount depends on the work. A person supervising a familiar administrative routine needs a different depth of continuing practice from someone responsible for engineering decisions.
Competence also includes knowing when judgment must move to a specialist. A business owner can govern the use of a tool without becoming qualified to approve legal, medical, or technical conclusions beyond their expertise.
A review should be capable of changing the outcome
A useful review has access to evidence, a clear standard, and the ability to produce an action. It can approve, request correction, narrow the scope, seek expert judgment, or stop the process.
If every case is approved because the queue is too large, review has become a throughput ritual. If rejected cases return unchanged until someone approves them, review has become pressure. If the reviewer never hears what happened after approval, they lose information needed to improve judgment.
The organization should inspect review outcomes as well as generated outputs. How often do reviewers request changes? Are the same problems recurring? Can reviewers explain their concerns without being penalized for slowing a process that genuinely needs to stop?
A very low rejection rate is ambiguous. It may reflect excellent outputs, weak standards, missing evidence, or a reviewer who has learned that objections are unwelcome. The number needs interpretation.
A governor therefore needs a feedback path from consequences to standards. Customer complaints, delivery failures, reconciliations, and corrected records should inform the review process. A clean queue alone is not evidence that the underlying work is sound.
Design escalation before the exception arrives
Exceptions are easier to manage when there is an agreed route for handling them. Waiting until an incident occurs forces people to invent authority under pressure.
A simple escalation rule specifies the trigger, the person or role contacted, the evidence included, and what the system does while waiting. The waiting state matters. Some workflows should pause; others can continue with an explicitly safe default. Continuing to improvise is not a neutral option.
For a hypothetical purchasing workflow, escalation might occur when a supplier changes bank details, a product specification is incomplete, or total commitments exceed a budget. The packet should include the proposed action and its supporting records, not merely a model’s assertion that something seems unusual.
The recipient must know what decision is being requested. “Please review” is weaker than “Approve this exception to the supplier rule, decline the order, or request verification.” Clear options reduce unnecessary back-and-forth without concealing the judgment involved.
If the designated person is unavailable, a fallback should be defined. A narrow, reversible administrative task may wait. An urgent customer problem may need another responsible person. An automated system should not silently expand its own authority because the preferred approver did not respond.
Govern capacity as well as quality
Automation can expand the number of proposals faster than it expands the capacity to review them. That creates a human bottleneck even when each generated result is useful.
Suppose a system produces fifty supplier recommendations each morning. If evaluating one recommendation carefully takes ten minutes, the apparent automation has created more than eight hours of review. The arithmetic is hypothetical, but the capacity problem is real enough to design for.
The response might be to narrow the search, filter by explicit requirements, batch comparable cases, or stop producing proposals that cannot be evaluated. It should not be to pretend that superficial approval is equivalent to the review the task requires.
A governor also manages attention across workflows. Ten small systems can produce a combined burden that no single dashboard makes obvious. Shared review hours, interruption frequency, and unresolved commitments matter alongside each system’s local success rate.
The one-person AI company explores the broader organizational consequences. A person can coordinate more work without acquiring unlimited judgment, availability, or emotional capacity.
Standards should distinguish kinds of error
Not every mistake deserves the same response. A formatting problem, an unsupported claim, an unauthorized payment, and a breach of confidence differ in consequence and reversibility.
A practical standard identifies the errors that are unacceptable, those that require correction before release, and those that can be tolerated within a narrow process. The boundaries should be explicit enough to guide a reviewer and modest enough to fit actual knowledge.
Some standards are deterministic. An invoice total should equal its component amounts. A proposed action should remain inside a known spending limit. Other standards require interpretation, such as whether a customer explanation is misleading or whether an exception is fair.
Separating these kinds of judgment protects human attention. Mechanical checks can remove routine errors before a person evaluates meaning and consequence. A person should not have to spend scarce review time repeating arithmetic that ordinary software can perform reliably.
Standards also need examples of failure. A clear counterexample can reveal a gap that an abstract instruction misses. The review process should include difficult cases relevant to its scope, rather than only the cases that make the system look competent.
Keep records that support answerability
A governor needs to be able to explain a consequential decision after it occurs. That does not require retaining every conversation forever. It requires preserving the evidence and authorization appropriate to the decision and relevant obligations.
Useful records distinguish the proposal, the permission, the action attempted, and the outcome verified. If a system prepared an order but never submitted it, the record should not imply that a purchase happened. If a person approved one version and the system sent another, that discrepancy matters.
The provenance and auditability chapter develops the record design. Here the human question is whether the person responsible can reconstruct enough of the process to correct a mistake and explain what will change.
Records should also protect privacy and avoid unnecessary surveillance. More logging is not automatically better governance. Sensitive information creates obligations of its own. A smaller, well-defined record may be more useful than an enormous archive nobody can interpret or safely maintain.
The purpose of an audit trail is correction and accountability, not the construction of a persuasive story that every decision was inevitable.
The transition can improve work, or hollow it out
A worker may welcome relief from repetitive administration and use the recovered time for relationships, craft, or more difficult judgment. Another may find that automation removes the part of work through which they developed competence and experienced contribution.
Those possibilities should be taken seriously when designing a role. A job made entirely of reviewing mistakes can become stressful, fragmented, and difficult to understand as a meaningful contribution. A job with real authority, learning, and contact with beneficiaries can develop in a different direction.
This is why the worker-to-governor transition should include a conversation about the work itself, not merely training in a dashboard. Which responsibilities remain? What competence will the person develop? What do they have permission to improve? How does their judgment affect the outcome?
The scarcity of purpose concerns choosing worthwhile commitments. Organizational governance must make space for that human question while meeting concrete delivery obligations. Efficiency does not excuse a role that assigns accountability without a workable way to exercise it.
A staged transition
Start by mapping one workflow the person already understands. Identify consequential decisions, ordinary checks, exception triggers, and the people affected by mistakes. Define the permitted automated actions narrowly.
Next, run the process in a preparation role. Let it assemble evidence and recommend actions while the existing decision process remains authoritative. Compare its proposals with actual judgments and record disagreements. This is a proposed evaluation approach, not a claim that a particular system has passed it.
If the evidence supports a limited expansion, authorize only the class of actions that was actually evaluated. Keep a return path to the prior process and a way to stop new commitments. A successful trial of invoice sorting does not establish authority to change payment instructions.
Then review the human role. Is the person better informed? Is the review burden manageable? Can they explain difficult cases? Do they receive consequence feedback? Have they gained a real ability to govern, or merely inherited responsibility for a system they cannot inspect?
An expansion should depend on those answers as well as task accuracy. A productive system with an unworkable supervisory arrangement has not completed the transition.
The transition also needs an exit criterion. If reviewers consistently lack the evidence needed to decide, if the queue exceeds available attention, or if stopping authority does not work in practice, return the affected actions to the previous process. That return is not proof that automation can never help. It identifies an arrangement that is currently inadequate. Record the missing condition so that a later proposal addresses the actual problem instead of repeating the same rollout with more optimistic language.
AI Leverage in Practice
What changed: systems can prepare and sometimes execute more work, shifting some human effort toward standards, exception handling, and consequential decisions.
What you can do today: choose one workflow and list who proposes, approves, challenges, and answers for each important decision. Check that responsible people have information, time, competence, and stopping authority. Test the escalation path before relying on it.
What may come later: more capable systems may handle broader tasks. That can increase the importance of governing scope and shared consequences. Technical capability alone does not determine legitimate decision rights.
Govern something you can answer for
Moving from worker to governor is a change in the shape of work. It is successful when a person can direct a process, understand its limits, hear challenges, and intervene effectively.
A review label cannot supply missing competence or authority. A useful governing role connects responsibility with the practical means to exercise it.
The next question is what increased leverage should achieve. Find the full Age of AI Leverage series and the wider AI section.
Sources
- OECD, Governing With Artificial Intelligence, 2025, public-sector governance context, not a universal private-business control prescription.
- NIST, AI Risk Management Framework, voluntary risk-management context. Using its language does not establish certification or successful implementation.
- The workflow examples and staged transition are proposed practices. No completed organizational experiment or guaranteed performance improvement is claimed.
Loading comments…