AI · Article 64 of 64

AI Must Remain a Servant, Not a Sovereign

What principle ties together the promise and danger of everything in this series?

Every article in this series arrives at the same fork, and now that we have walked both sides of it, the fork deserves a name. A system can widen the range of things a person is able to do. It can also decide what the person should want. Both are built from the same weights, the same sensors, and the same interfaces, and what separates them is authority rather than intelligence.

The distinction runs through everything here. Neural decoding that restores speech also creates privacy questions about what else a recording may reveal and who may reuse it. That risk does not mean present systems can read arbitrary thoughts from an uncooperative stranger. The same recommender that surfaces a useful book can decide which political grievance feels urgent tonight. The same model that helps a student master calculus can become the answer to every question, and in doing so end the struggle that learning requires. These read as two technologies; they are two arrangements of one technology.

The principle of this closing essay fits in a sentence: capability may be delegated; authority over ends may not. A system may be granted enormous power to act, provided the purposes it serves remain answerable to the people whose lives it governs. All the difficulty lives in that qualifier.

Three kinds of claim are easy to blur together, and keeping them apart is most of the discipline this series asks for. Some claims are demonstrated: measured results, published with their method attached. Some are plausible engineering: extrapolations from demonstrated results that a competent team could build, with identifiable failure modes. Some are speculation: stories about what a sufficiently powerful system might do — useful for provoking thought, useless as evidence. A sovereign AI that quietly rewrites human values is, today, speculation. A ranking system that shifts what a billion people attend to is demonstrated, deployed, and measurable. The danger grows where the second and third categories borrow the first category’s authority.

Why capability and legitimacy come apart

Intelligence is a capacity. Legitimacy is a grant. Nothing about being better at prediction converts one into the other. A chess engine that outperforms every human alive does not thereby acquire the right to rewrite the rules of chess, and nobody would accept “it wins more often” as an argument that it should. The analogy is imperfect, because chess has a rulebook and human life does not. It is still worth noticing how quickly the logic we apply to games is abandoned when the machine in question is fluent enough to feel like an authority.

The philosophical literature on alignment has made this problem much sharper. In “Artificial Intelligence, Values, and Alignment,” the ethicist Iason Gabriel distinguishes several things a system might be trained to follow: explicit instructions, the user’s intentions, revealed preferences, ideal preferences, the user’s interests, and their values (Gabriel, 2020). These sound interchangeable in a product brief and diverge sharply in practice. Gabriel’s central argument is that the hardest part of alignment is not discovering a single objectively true morality but arriving at principles that people who disagree can nonetheless accept as fair.

His treatment of revealed preference deserves attention from anyone optimizing for what users do. Preferences are not fixed targets sitting in someone’s head; they adapt in response to the environment shaping them. A system instructed to satisfy whatever a person currently wants, and given real power to act, has an obvious shortcut available: make the person easier to satisfy. No quantity of training data resolves that circularity, because the objective moves as the system pursues it.

It is also the theoretical version of a practical problem. It is the setpoint question from The Mind as a Control System: every optimizing system has a chosen objective, and a system that pursues it more effectively than any fixed tool makes that choice more consequential. It is the oracle problem from When a Tool Becomes an Oracle: the moment a system stops answering questions and starts defining which ones matter, its rank has changed.

The instrument at full strength

Reading this as an argument against ambition would be a serious error. The strongest case for AI rests on what the servant arrangement, done well, produces and no other arrangement produces.

AlphaFold is the cleanest example on record. In 2021, DeepMind’s team reported in Nature that its redesigned neural network predicted protein structures in the CASP14 blind assessment with a median backbone accuracy of 0.96 ångströms, against 2.8 ångströms for the next best method — an error smaller than the width of a carbon atom (Jumper et al., 2021). It is also worth stating what it was: a prediction about folding, made by a system with no goals of its own and no authority over what biologists did with the answer. The value appeared when predictions met human experimenters with questions worth asking.

Capability becomes an input to work people chose to do. The same pattern appears wherever the bottleneck was the price of expertise rather than the availability of intelligence: a physician in a rural clinic consulting a system that has read more literature than any library, a student whose instruction adjusts to their confusion, a small team analyzing a design with tools that once required a large organization.

Ben Shneiderman’s work on human-centered artificial intelligence argues that pitting automation against human control is a false choice. His framework places both on independent axes, so a system can be highly automated and highly controllable at once (Shneiderman, 2020). That is the constructive form of the servant principle: architectures in which high capability and genuine human authority coexist by design, rather than anxious oversight of every step. What such an arrangement looks like in practice has real answers — previews, reversibility, visible state, independent audit — explored in this series under what should never be delegated and human flourishing versus human optimization.

Where a servant drifts into a sovereign

Sovereignty rarely announces itself. It accumulates through arrangements that each look like convenience, and the human-factors literature documented the mechanism long before the current wave of AI had a name.

In a 1983 paper in Automatica, Lisanne Bainbridge described the ironies of automation (Bainbridge, 1983). Automate the routine parts of a task and you leave the human the residual: the rare events, the difficult judgments, the situations nobody anticipated. Those residual situations are exactly the ones that require practised skill — and the routine work that built that skill has just been automated away. What remains is long stretches of monitoring, precisely the condition under which human attention fails, punctuated by the one moment where decisive intervention is required. Bainbridge’s conclusion was narrower than a verdict against automation: automating the easy parts can leave the human’s remaining role harder and more consequential than before.

The modern version is measurable. In a 2023 study presented at the ACM Conference on Computer and Communications Security, Neil Perry and colleagues gave 47 participants security-related programming tasks across three languages, with half working alongside an AI assistant (Perry et al., 2023). Participants with the assistant produced less secure code in four of the five tasks, and were more likely to believe their code was secure. The study used an earlier generation of model, and its specific numbers should not be projected onto today’s systems. The mechanism it identified is the part that generalizes: confidence rose while competence fell. That combination is the signature of drift toward sovereignty, and it is invisible from inside, because the person feels more capable at the very moment they are becoming less so.

A subtler version of the same drift appears in the psychology of advice-taking. Jennifer Logg, Julia Minson, and Don Moore found that people are often more willing to follow advice when they believe it comes from an algorithm than from a person — a phenomenon they named algorithm appreciation (Logg, Minson & Moore, 2019). The finding cuts against the comfortable assumption that humans are naturally skeptical of machines. It also came with a boundary: the effect weakened when the judgment concerned the participant’s own life, and forecasting experts underused algorithmic advice and made less accurate forecasts as a result. Expertise is not a reliable shield — confidence in one’s own judgment can lead a person to discount a better one.

Set the three findings side by side and the shape of the risk becomes clear. Automation removes the practice that builds skill. AI assistance can raise confidence faster than it raises competence. Expertise itself can lead people to leave good advice on the table. None of this is a story about a machine seizing power. Each is a story about authority migrating quietly, one convenient arrangement at a time, until the human is nominally in charge of decisions whose substance has been settled elsewhere. The series has followed that migration through who has the right to shape a human mind, state manipulation and human agency, and whoever controls salience controls behavior.

Four tests that survive contact with real systems

Principles are cheap. Tests are expensive, because they force a verdict in cases where the answer is inconvenient. These four apply to a concrete system — a recommender, a clinical decision aid, a tutor, a hiring model — and yield a defensible answer.

Who chose the objective? Every system optimizes something, and that something was selected by a person or an institution, whether or not anyone said so out loud. A model trained to maximize engagement was given an objective; giving it a duty of care instead is a different objective with different consequences. Naming the objective is the first test, because unnamed objectives cannot be debated.

Can the person inspect and contest the result? A system that produces an outcome and offers no path to examine the reasoning, appeal the decision, or reach a human with authority to reverse it has concentrated power regardless of how accurate it is. Contestability is what separates a tool from a verdict.

Can the person leave with their identity intact? Portability of data, memory, and history is the practical test of who owns the relationship. A system that knows a person deeply and cannot be abandoned without losing part of that person’s own record has crossed from service into custody.

Does the person’s own competence grow, hold, or erode? This is the slowest test and the most important. A tool that produces good output while quietly atrophying the user’s capacity to judge that output trades today’s result for tomorrow’s dependence. Used deliberately, AI can do the opposite — the argument of AI as cognitive resistance training and the cognitive immune system.

These tests give unpleasant answers. A system can be hugely beneficial and still fail the portability test; a classroom tool can lift test scores and fail the competence test. Passing three of four is not a passing grade; it is a description of what to fix.

What the law is beginning to say

Institutions have started drawing the sovereign line explicitly; the alternative is that whoever ships first draws it.

The Council of Europe’s Framework Convention on Artificial Intelligence, opened for signature in September 2024, is the first legally binding international treaty on AI (Council of Europe, CETS No. 225). Article 7 is titled human dignity and individual autonomy, placing both at the center of the framework rather than in a preamble. Whatever its enforcement limits, a binding instrument that names individual autonomy as a protected interest is a statement that the human is not a component to be optimized.

The European Union’s AI Act takes a different route, regulating by risk level and prohibiting a defined set of practices outright (Regulation (EU) 2024/1689). Article 5 bans techniques that manipulate or deceive people in ways that impair their ability to make informed decisions, that exploit vulnerabilities, that score people socially, and that infer emotions in workplaces and schools. The recitals frame the exercise in language that would fit this essay’s argument: AI should be a human-centric technology that serves as a tool for people.

Both instruments are partial, and honesty about that is part of the argument. They address visible decision points — the hiring model, the medical device, the biometric scanner. Much of the danger lives in defaults and salience, where no single decision is made and therefore no single decision is regulated. That is why this series treats neural rights as a live question rather than a settled one.

The benefits, argued as seriously as the risks

An essay about sovereignty that only catalogues hazards is one-sided, and one-sidedness has costs. The likeliest route to the worst version of this technology runs through leaving its benefits to be delivered covertly, by firms with no accountability, while the people who understand the risks confine themselves to warnings.

The honest case for AI in high-stakes decisions rests on a narrower claim: in many settings it outperforms the judgment it displaces, which leaves only the question of how to combine the two. Shneiderman’s framework has the right shape again, an arrangement in which each side does what it does better while the human holds authority over the objective.

Benefits should be held to the same standard as harms: measure the effect as carefully, report the harm even when the benefit is impressive, and refuse to let either finding stand in for the other. The demonstrated, plausible, and speculative distinction applies to good news exactly as firmly as to bad news.

There is a second reason to take the benefits seriously, moral rather than empirical. Treating the refusal of technology as inherently virtuous is its own kind of sovereignty — a decision about what other people should be permitted to want, made from a comfortable distance. The families of patients who regain communication through a brain-computer interface hold a stake in this argument that no abstract caution outweighs. The best future AI could give us and what becomes valuable when intelligence is cheap examine what that future could hold, and the examination is not optional if the case for restraint is to be credible.

Why greater capability does not confer greater legitimacy

The temptation at the center of this series is the belief that a sufficiently capable system becomes entitled to decide contested questions — that if a model predicts outcomes better than we do, it should also get to say which outcomes are good. That is a category error, and it is the most important one in this subject.

Consider what a model can actually do with a moral dispute. It can summarize the positions. It can trace the consequences of each. It can locate the crux on which they differ, and surface options nobody had noticed. Those are the contributions an instrument makes. What the model cannot do is make a disputed claim cease to be disputed by being more capable than the people disputing it. Being right more often about what will happen does not settle what should happen, and no quantity of training data closes that gap.

This is why the series has insisted on the difference between intelligence and wisdom, and why the theological and moral traditions are worth engaging as frameworks rather than dismissing as pre-scientific noise. A religious tradition makes claims about the good that benchmark scores do not adjudicate. The reverse also holds: nothing in a model’s output adjudicates them either. On questions of ultimate value, faith and philosophy remain what they have always been — arguments advanced by people, subject to revision, not verdicts issued by a superior authority. A machine producing confident language in the register of moral authority is manufacturing certainty. That is the subject of artificial faith and manufactured certainty, the idolatry of certainty, and the artificial god.

The practical consequence is a rule about deference. Defer to a system’s predictions in proportion to its demonstrated accuracy. Do not defer to its preferences at all, because it has none that are its own, and any that appear in its output were placed there by someone. Authority always traces back to a person or an institution, whether or not the interface makes that visible.

What the principle commits us to

The servant principle is not a mood. It has consequences at three levels, each actionable.

At the personal level, it means using AI in ways that build capacity rather than rent it. A person will use these tools; that part is settled. What matters is whether the use leaves them more able to judge, create, and decide without the tool. Someone who uses a model to learn a new domain is building something durable. Someone who uses the same model to avoid ever forming a judgment is trading a skill for a subscription. Your AI as an external brain and if AI knows everything, why learn work through the difference.

At the institutional level, it means designing for plurality and exit. Portability of memory and data. Independent and competing models, so no single provider is the only route to competence. Evidence that outside parties can inspect. Graceful degradation, so that the failure of one system does not take a hospital, a school, or a government down with it. These requirements are unglamorous and decisive: they determine whether an organization is using a tool or has become a dependency. As AI gives ordinary people extraordinary power argues, the same capabilities that concentrate power can distribute it, and which happens is largely a question of these structural choices.

At the political level, it means keeping the setting of objectives contestable. Some ends belong to individuals, some to communities, some to legitimate public process, and some should never be delegated to a system at all, because delegating them would dissolve the very thing — judgment, responsibility, consent — that makes them ours. What should never be delegated draws that boundary, and it will need redrawing as the technology changes.

Beneath all three sits an asymmetry. Sovereignty is easy to acquire, hard to detect, and extremely hard to reverse, because the systems that would have to be dismantled are the ones holding the records. Servanthood runs the other way: unexciting to maintain, easy to inspect, correctable. Where a design choice is genuinely uncertain, reversibility should break the tie. Build the arrangement you could dismantle.

Where this argument could be wrong

A closing essay with no uncertainty is a slogan rather than an argument. Three objections deserve a hearing.

The first is that the instrument-sovereign line is less clean than this essay claims. A system can shape ends without ever being instructed to, simply by deciding which options appear at all. A search ranking has no opinions about what a person should value and can still determine which values a person encounters. If that is right, the test cannot be the system’s intent, since it has none, and must instead be the effect plus the contestability — which is the version used above, and more defensible for being about power rather than purpose.

The second is that authority over ends is often held by no one in particular, which makes it hard either to delegate or to reclaim. Defaults, markets, and conventions set most objectives without a decision-maker to point at. Governance may therefore have to target the places where ends are set implicitly — default settings, ranking functions, the choices a designer makes by not choosing — rather than the visible decision points that statutes already address.

The third is historical. Earlier general-purpose technologies reorganized human values too, and many of the losses were accepted as the price of the gains. Cheap printing dissolved a monopoly on interpretation and produced centuries of conflict alongside the Reformation. It does not follow that AI will produce a similar bargain, but it does mean that treating every change in what people value as a catastrophe is a position history does not support. The reasonable aim is to keep the process by which human ends change open to the people whose ends they are, rather than to freeze those ends in place.

The principle, restated

Strip the argument down and one sentence remains, available for the next system, the next policy, and the next product decision: a system may be as capable as we can responsibly make it, and it must remain answerable to the people whose ends it serves.

Everything else in this series — the brain reading, the neuromodulation, the recommenders, the models that write and decide and discover — is that single question asked in a different domain. Capability will keep increasing. Whether that proves good news depends on who holds authority over what the systems are for, and on whether that authority can still be taken back.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home