A hospital wants to know whether a screening program can hand its first read of a mammogram to a model. A court wants to know whether a risk score may inform a sentence. A defence ministry wants to know whether a targeting decision may be made at machine speed. A person wants to know whether to let a model decide how to respond to a message from a dying parent. In each case someone asks the same question in almost the same words: if the system is more accurate, why not let it decide?
The question is reasonable, and the answer that treats accuracy as the only currency is wrong in a way that is easy to miss. Delegation is not one act. Handing a decision to a system can mean handing over a calculation, the authority to choose, the answerability for the outcome, or the work of forming the judgment in the first place. Those four transfers come apart, and most of the useful thinking about delegation boundaries lives in the gaps between them.
Accuracy is one input among several
It helps to start by being clear about what a system can be better at. A model can be more accurate than a person at estimating a probability, detecting a pattern in an image, or ranking a list of risk factors. Those are real advantages, and sometimes enormous ones. What a model cannot do is be the party whose decision it is in any sense beyond authorship of the output.
An example from a well-studied domain makes the point. In the United States, the Wisconsin Supreme Court reviewed the use of a commercial risk assessment called COMPAS in the sentencing of Eric Loomis in 2016. The presentence report included the tool’s risk scores, and the sentencing court referred to them while ruling out probation. The court upheld the practice, but only with unusual caution, requiring that presentence reports carry written warnings: that the tool’s methodology is proprietary and not disclosed, that its scores describe groups rather than individuals, that it had not been validated on the Wisconsin population, that studies had raised questions about disproportionate classification of minority offenders, and that the tool was designed for post-sentencing decisions rather than sentencing itself. The report accompanying the assessment said plainly that the scores “are not intended to determine the severity of the sentence or whether an offender is incarcerated.” A commentator in the Harvard Law Review argued a year later that a written advisement of this kind is a weak instrument, because it tells judges to be sceptical without telling them how much to discount, and because judges generally cannot evaluate a closed model’s reasoning even when they want to.
The interesting part of that case is not the judicial outcome. It is that the law, which is not usually shy about adopting efficient tools, drew a distinction between a decision being informed by a probabilistic estimate and the decision being made by it. The distinction is not sentimental. It rests on reasons that hold regardless of how accurate the model becomes.
Four reasons a decision may need to stay human
The four reasons are distinct and can pull in different directions. Naming them separately prevents a common argumentative error, which is to refute the strongest one and assume the others have been answered.
The epistemic reason. A human overseer adds value only if they can genuinely evaluate what they are overseeing. If they cannot understand the basis of an output, cannot tell a confident wrong answer from a confident right one, and would defer anyway, then placing a person in the loop is a formality that launders a machine decision into an apparently human one. This is a reason about whether oversight is real, and it is the reason most often ignored, because a nominal human in the loop looks reassuring in a schematic.
The accountability reason. Someone must be answerable when a decision causes harm. Responsibility can, in principle, be assigned to the person who deployed a system, the organisation that operates it, or the state that authorised it. It cannot be assigned to the model itself, because a model cannot bear consequences, testify, or be punished. Where the point of a decision is partly that some identifiable party owns the outcome, the decision has to rest with someone who can own it.
The consent and relationship reason. Some decisions are not conclusions reached about a person. They are acts the person is a party to: agreeing to treatment, entering a relationship, forgiving, committing, blessing, saying goodbye. A model can inform such decisions with analysis and context. It cannot be the one who consents, and its participation does not substitute for the person’s participation, because the participation is the thing that matters.
The legitimacy reason. Some decisions have to be made by a body with standing, through a process a person can contest, because the value of the decision depends on its provenance rather than on its content. A verdict reached by the correct court following the correct procedure is a verdict; an identical conclusion reached privately and announced afterwards is not. This is the reason that complicates the otherwise appealing idea that the most accurate decision-maker should always be the one who decides.
The oversight constraint, in regulation and in practice
The epistemic reason is not abstract, and it has been written into law in a way that is unusually concrete. Article 14 of the European Union’s AI Act requires that high-risk systems be designed so that people can effectively oversee them during use, and it spells out what that oversight has to enable: understanding the system’s real capacities and limits, monitoring its operation for anomalies, remaining aware of the tendency to over-rely on outputs, correctly interpreting those outputs, deciding not to use the system or to disregard, override, or reverse its output, and interrupting it safely. For certain biometric identification systems, a decision based on the identification requires separate verification by at least two competent people. The regulation’s language is a useful antidote to the assumption that “human in the loop” is a binary property. Oversight is a set of capabilities, and it can be absent even when a human is nominally present.
The empirical literature explains why. Automation bias — the tendency to treat an automated cue as a substitute for active checking rather than an input to it — has been documented in aviation, medicine, and elsewhere since the 1990s. Its most uncomfortable property is that it does not depend on the human being careless. In a 2023 study in Radiology, Thomas Dratsch and colleagues asked 27 radiologists of varying experience to read 50 mammograms while a simulated AI system offered BI-RADS suggestions. When the AI’s suggestions were correct, performance improved. When they were incorrect, performance fell significantly across every experience level, and the least experienced readers were the most likely to follow a wrong suggestion upward. The radiologists were competent professionals doing the task they were trained for, and the presence of an authoritative-looking hint pulled their judgment off course.
That finding should change how the oversight requirement is read. Placing a person in the loop does not automatically create independent judgment; it can create a person who endorses. Genuine oversight requires conditions — time, competence, evidence, and the social permission to disagree — and where those are absent, an override button that nobody uses is decoration.
Decisions that use force, and what the international debate settled
The most fully developed deliberation about delegation boundaries comes from an unexpected place: the international debate over autonomous weapon systems. It is worth studying less because its conclusions transfer wholesale than because it is the longest-running serious attempt to reason about what must stay human.
The International Committee of the Red Cross has argued, in a 2018 report on ethics and autonomous weapon systems, that human agency and intent must be retained in decisions to use force, and that responsibility for such decisions cannot be transferred to machines. A 2020 SIPRI and ICRC study tried to make that principle operational by describing practical elements of human control — the kinds of knowledge, time, and authority that a person must have for control to be meaningful rather than nominal. Later SIPRI analysis identified four reasons that human agency must be retained: some rules of international humanitarian law are addressed to human beings; accountability cannot be transferred to a machine; some rules require context-dependent judgements that do not translate into technical indicators; and the Martens Clause, which keeps conduct subject to the principles of humanity and the dictates of public conscience, cannot be satisfied by an algorithm.
Notice what that list contains. The first and third reasons are epistemic, the second is about accountability, and the fourth is about legitimacy. The weapons debate arrived, through a route quite different from the one this article has been following, at the same four-part structure. That convergence is a reason to take the structure seriously rather than as an invention of the moment.
It is also worth noticing what the debate has not settled, because honesty about the boundary requires it. There is broad agreement that a human must retain control over the decision to kill specific people. There is much less agreement about where that control has to sit, how far in advance it must operate, or what it looks like when a system is screening thousands of sensor readings to propose targets. Everyone can agree that a person must decide; the hard engineering and legal questions begin immediately after.
Where delegation clearly works, and why the cases differ
An account that treats delegation as inherently suspect would be as wrong as one that treats it as inherently efficient. The mammography literature shows both halves at once.
In a prospective, population-based study published in Lancet Digital Health in 2023, Karin Dembrower and colleagues screened 55,581 women in Stockholm using a setup in which one radiologist’s read was supported by an AI system, with a second radiologist reading as well. The AI-supported double reading was non-inferior to the standard two-radiologist double reading on the primary measure, and detected about four per cent more cancers. The study’s authors disclosed that the AI vendor had funded the work, which is a reason to read the result with the appropriate scepticism rather than to dismiss it. A 2024 study of the Danish screening program reported similarly encouraging operational results, with substantial reductions in reading workload alongside improvements in several screening metrics.
Those are genuine gains, and they show that the relevant question is not whether to delegate but which part of a process to delegate and how to keep the residual human judgment real. Replacing a first reader who is subsequently checked by a second is a different act from removing the second reader, and the literature distinguishes them. Simulations have found that replacing the first reader is broadly feasible, that replacing the second reader lowers sensitivity, and that triaging which cases get human attention can improve results. The same technology, deployed in slightly different places in the workflow, produces benefits on one side and a safety regression on the other.
The distinction that matters is not domain-by-domain. It is structural. Delegation works when the task is high-volume, statistically regular, reversible, and checkable by a party who has an independent basis for disagreeing. It becomes dangerous when the task is rare, consequential, irreversible, and checked by someone whose only basis for confidence is the system’s own output. Most of the contested cases are contested precisely because they sit between those poles.
The moral-skilling argument
There is a further cost to delegation that does not appear in accuracy statistics at all: what happens to the person who stops doing the deciding.
Shannon Vallor’s work on moral deskilling and upskilling makes the case carefully. In a 2015 paper in Philosophy & Technology, she argues that moral capacities are like other skilled capacities: they are acquired and maintained through practice, and they can atrophy when the practice is removed. Technologies that take over moral work can deskill the people who use them, in the way that a tool can take over any other practised skill. This is a claim about the formation of character rather than about the correctness of any particular decision, and it applies even when the machine’s judgments are better than the person’s would have been.
The argument’s force depends on a premise that should be stated rather than smuggled in: that a person’s moral capacities are worth maintaining for reasons independent of the decisions those capacities produce. Someone who holds that view will find moral deskilling a serious cost. Someone who holds that only outcomes matter will not. This is, in the end, a disagreement about what a person is for, and no amount of evidence about model accuracy resolves it. That is a reason to be explicit about the premise rather than to pretend the argument is technical.
Religious and theological traditions go further, treating some decisions as acts of the person that cannot be performed by proxy — worship, repentance, forgiveness, the acceptance of a calling. These are framework claims about what a human life is, not empirical predictions about model performance, and they should be presented and weighed as such. What they share with the secular deskilling argument is a structure: the value of the act resides partly in its being the agent’s own, which is a property no improvement in accuracy can supply.
Sentencing, consent, and the law’s existing floors
Some of these boundaries are already written into law, which is useful evidence that they are not merely philosophical.
Article 22 of the General Data Protection Regulation gives people the right not to be subject to a decision based solely on automated processing, including profiling, that produces legal effects or similarly significantly affects them, subject to narrow exceptions and with safeguards that include the right to obtain human intervention, to express a point of view, and to contest the decision. The provision does not say that automated decisions are inaccurate. It says that a person is entitled to a decision-making process in which a human is genuinely involved and can be argued with. That is the legitimacy reason in statutory form.
Notice how carefully the regulation’s own guidance treats the requirement. The human involvement has to be more than a rubber stamp: the overseer needs the competence to evaluate the case independently, the authority to change the outcome, and the information to consider circumstances the model did not weigh. A review process meeting that description is doing real work. One that consists of a staff member clicking approve is not, and calling it human oversight does not make it so.
The delegation boundary, then, is rarely a clean line between “AI may never decide X” and “AI may decide X.” It is usually a line about which functions must be preserved in the human role: the ability to disagree, the authority to change the outcome, the standing to be answerable, and the person’s own participation where the decision is theirs to make.
A test that survives contact with real systems
Abstractions about dignity and accountability are easy to affirm and hard to apply. A short set of questions, asked of any proposed delegation, tends to separate the cases that can be settled from the ones that should worry us.
Who bears the consequence if the decision is wrong, and can that party actually bear it? If the answer is “the model” or “no one,” the accountability requirement has not been met, whatever the accuracy figure.
Can the person who is supposed to be in charge disagree in practice? This requires more than a button. It requires that they understand the basis of the output, have time and standing to challenge it, and will not be penalised for doing so. Where those conditions fail, the oversight is fictional and should be described as such.
Is the decision one the affected person is party to, rather than the subject of? Consent, disclosure, and agreements made in the course of a relationship belong in this category. A model may inform them thoroughly. It cannot be the one who agrees.
Does the decision require a legitimate authority, and is that authority’s standing part of what makes the decision binding? Adjudication, discipline, the allocation of public benefits, and the use of force all fall here.
What does the person lose by not doing this themselves, and is that loss acceptable to them once it is pointed out? This is where the moral-skilling cost gets weighed, honestly and with the premise exposed.
Would you be comfortable describing the arrangement accurately? A formulation that survives this test sounds like “a model produced a risk estimate and a human being, who could see the underlying data and had authority to reject it, decided.” A formulation that fails sounds like “the AI decided and a human confirmed it.”
Designing for contestability
If these boundaries are worth keeping, they create obligations for the systems built near them. The requirements are not exotic; they overlap substantially with the capabilities Article 14 already names.
Systems used where the epistemic reason applies must expose enough reasoning to be argued with. Sources, uncertainty, the factors that drove an estimate, and the cases where the model is least reliable matter more than a polished recommendation. A system that produces a confident conclusion with no inspectable basis has not enabled oversight; it has replaced it.
Systems used where the accountability reason applies must leave a record that supports the assignment of responsibility to a person or institution, including who decided, on what basis, and with what opportunity to dissent. A decision that cannot be reconstructed afterwards cannot be owned afterwards.
Systems used where consent or legitimacy applies must be built to route the decision to a person at the points where the person’s participation is the point, and to make that routing a matter of design rather than of the user’s diligence.
And all of them should be built on the assumption that a human overseer will drift toward agreement, because that is what the evidence shows people do. Designing against automation bias is a different exercise from designing against error. It requires slowing decisions down in the places that matter, varying recommendations so that agreement is not automatic, and deliberately preserving the cases where the human’s independent judgment is exercised, in the same way a training program preserves a skill by requiring it to be used.
Where the boundary actually lies
The question that opened this article has no single answer, and the attempt to produce one usually fails in the same way: it treats accuracy as the only dimension along which delegation can be assessed. That framing makes every boundary look arbitrary, because against accuracy alone any boundary is arbitrary.
Add the other dimensions and the picture settles. Decisions that use force against persons, impose serious sanctions, determine whether someone receives a public benefit they are entitled to, or constitute a person’s own commitments are poor candidates for full delegation, not because machines are bad at them but because part of what makes those decisions legitimate is who makes them and how they can be contested. Decisions that are high-volume, reversible, statistically regular, and open to independent checking are strong candidates, and the evidence from screening programs shows that the benefits can be substantial.
The hard part is what sits between: the large middle of professional work where a model is better than a person at the estimate, the person is still accountable for the outcome, and nobody has decided who is really in charge. That middle is where the boundary is currently being drawn by accident, one deployment at a time. Drawing it deliberately requires asking the four questions — epistemic, accountability, consent, legitimacy — before the system ships rather than after it fails. The machinery to do so already exists in law, in the human-factors literature, and in a decade of international argument about weapons. What is missing is the habit of using it.
Sources and further reading
- Vallor, “Moral Deskilling and Upskilling in a New Machine Age: Reflections on the Ambiguous Future of Character,” Philosophy & Technology, 2015
- ICRC, “Ethics and autonomous weapon systems: An ethical basis for human control?”, 2018
- SIPRI & ICRC, “Limits on Autonomy in Weapon Systems: Identifying Practical Elements of Human Control,” 2020
- European Union, “Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 14: Human oversight,” 2024
- European Union, “Regulation (EU) 2016/679 (General Data Protection Regulation), Article 22: Automated individual decision-making, including profiling,” 2016
- State of Wisconsin v. Eric L. Loomis, 2016 WI 68, Wisconsin Supreme Court, 2016
- Dratsch et al., “Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance,” Radiology, 2023
- Dembrower et al., “Artificial intelligence for breast cancer detection in screening mammography in Sweden (ScreenTrustCAD): a prospective, population-based study,” The Lancet Digital Health, 2023
Loading comments…