A student sits with an AI tutor that has spent an hour learning which explanations make her curious and which make her shut down. It adjusts, and she works longer than she meant to. A voter answers a chatbot’s questions in a debate and finds his position has moved slightly, then more. A patient with treatment-resistant depression has an electrode in her brain that detects the signature of her low mood and stimulates in response, and for the first time in years she has preferences she can act on. In each case, a system changed a state from which a decision later emerged. In each case, the person made the final choice. The question is what that word “made” is doing.
That question is the subject of this article: if a system alters the state from which a decision emerges, whose decision is it? It is not a question about whether machines can be conscious or whether free will exists in some metaphysical sense. It is a practical question about authorship. When we say a decision is someone’s own, we do not mean it was uncaused. We mean it issued from the person in a way that person could recognize and endorse, rather than being steered past their capacity to weigh it. The interesting cases are not the ones where someone is replaced by a machine. They are the ones where the person acts, sincerely and voluntarily, from a state that has been shaped.
A distinction with more than rhetorical force
The philosophical literature has a well-worked vocabulary for this, and it earns its keep by drawing lines the law and everyday speech blur together. Persuasion addresses a person’s capacity to judge: it offers reasons, arguments, or images and leaves the assessment to the chooser. Coercion removes or threatens options, so the person’s range of acceptable choices shrinks. Manipulation is the subtler category. Daniel Susser, Beate Roessler, and Helen Nissenbaum define it as influence that covertly subverts a person’s decision-making power, usually by exploiting a vulnerability the person has not chosen to expose. Their account, developed across a 2019 Georgetown Law Technology Review article and a companion piece in Internet Policy Review, treats manipulation as distinct from both persuasion and coercion, and — importantly for the digital case — as possible without any false belief. A system can manipulate without lying. It can simply arrange the encounter so that a person’s reasoning is bent around a fact about themselves they do not see.
Allen Wood, in a chapter in the 2014 Oxford volume Manipulation: Theory and Practice, offers a related taxonomy: manipulation works through deception, through pressure, or through the exploitation of emotional and character vulnerabilities. All three share a feature: the manipulator aims at a result while concealing the aim or the mechanism. That concealment is what makes manipulation corrosive to autonomy in a way that openly stated persuasion is not. If you know someone is trying to sell you something, you can discount for it. If the attempt is invisible, the discount cannot be applied.
Joseph Raz’s The Morality of Freedom supplies the broader frame: autonomy requires not only mental capacities but an adequate range of options and independence from coercion and manipulation. On this view, autonomy is not a switch that is either on or off. It is a condition with several ingredients, and a system can degrade one without destroying the others.
Layered agency
The classical account of layered agency comes from Harry Frankfurt’s 1971 essay “Freedom of the Will and the Concept of a Person.” Frankfurt distinguished first-order desires, the wants we simply have, from second-order desires, wants about our wants, and from second-order volitions, the desires that a particular first-order want be the one that moves us. The willing addict wants to take the drug and wants his wanting to be his own. The unwilling addict wants the drug and wants not to want it. The difference is not in the strength of the craving but in the relation of the person to it. Frankfurt’s “wanton” is the figure who has first-order desires but no second-order volitions at all — who simply acts on whatever pull is strongest, with no evaluative stance toward it.
This layered picture is the reason the question of authorship is not dissolved by the fact that every decision has causes. A craving that arrives unbidden is still mine if I can take a stance toward it, endorse it or refuse it, and have that stance count. What threatens agency is not influence as such but influence that bypasses the layer at which the person could take a stance. The willing addict’s agency is intact; the unwilling addict’s is compromised not because he is caused to want the drug but because his second-order stance has no purchase on his behavior.
Now apply the machinery. If a system changes the strength of a first-order pull, and the person retains the capacity to evaluate that pull and act against it, the layered account says the decision is still substantially theirs, even if the pull was engineered. If the system changes the state at which the evaluation itself happens — if it makes the evaluation come out differently than it would have under other conditions, in a way the person cannot detect and would not endorse — then the second-order layer has been co-opted. The decision still issues from the person. It is no longer clearly authored by them.
Reuben Sass’s 2024 Philosophy and Technology article adds a useful complication. Sass argues that autonomy should be understood as supervening on three distinct elements — agency, authenticity, and individual control over decision-making — and that these can move in opposite directions. An algorithmic feed that reduces a user’s moment-to-moment control might, in some contexts, increase the authenticity of what they encounter by steering them toward more mainstream sources. The implication is that “manipulation” and “reduced autonomy” are not synonyms. A manipulative intervention might, net, leave a person more able to act on their own values than the alternative. This is uncomfortable for critics and worth taking seriously, because it forces the argument away from the word “manipulation” and onto the specific question of which element of autonomy has been affected and how.
What the experiments show
The empirical literature is younger but moving quickly. In a preregistered study published in Nature Human Behaviour in 2025, Francesco Salvi and colleagues pitted GPT-4 against human debaters across 900 conversations. GPT-4 equipped with personal information about its opponent was more persuasive, achieving higher agreement rates. An author correction later clarified that the direct comparison between personalized and non-personalized AI fell short of significance, so the headline should be stated carefully: the AI was persuasive, and its personalization advantage is suggestive rather than established. That correction is itself a small lesson in how quickly claims about AI persuasion outrun their evidence.
The more striking result for the agency question comes from Kevin Costello, David Rand, and Gordon Pennycook’s 2024 Science study of conspiracy beliefs. Across 2,190 participants, conversations with a GPT-4-based assistant reduced belief in conspiracy theories by roughly 20 percent, an effect that persisted for at least two months and spilled over to conspiracies the conversation had not addressed. Crucially, the researchers identified the mechanism: the AI worked by supplying accurate facts and alternative explanations, not by flattery or emotional pressure, and a fact-checker rated 99.2 percent of the AI’s claims as true. The effect was absent for conspiracy theories that were actually true.
That study complicates the framing that these systems are mainly engines of manipulation. Changing someone’s mind through accurate information and better arguments is, on most accounts, the paradigm case of respecting their agency. It also raises a disquieting question: if the mechanism is indistinguishable from good teaching, then the same machinery that improves a person’s grip on reality could, with a different objective, degrade it. The tool does not know the difference between correcting a false belief and installing one. The objective does.
The deference paradox
A separate strand of the problem does not involve manipulation at all. It involves the rational erosion of competence through relying on a system that is usually right. If an AI’s judgment is better calibrated than yours across a domain, deferring to it is not irrational; it is the same reason you defer to a physician or a pilot. But the capacity to catch the system’s errors is itself a skill, and it atrophies without use. The consequence is a paradox: the better the adviser becomes, the less equipped the person is to know when it has failed.
This is not a hypothetical. The history of automation is full of it. Pilots who hand-fly rarely are slower to recognize an autopilot failure; radiologists who lean on computer-aided detection can develop weaker independent reading. The relevant question for AI is not whether deference is warranted — often it is — but whether the design preserves the friction required for the human to remain a genuine check rather than a rubber stamp.
A healthy design, on this account, does not maximize the smoothness of the handoff. It shows evidence, exposes uncertainty, invites challenge, and keeps a named person responsible for consequential calls. Sass’s distinction between control and autonomy cuts here too. Reducing the user’s moment-to-moment control to make a system faster can be a fair trade; the point is that it trades a specific thing against a specific gain, to be named and weighed rather than waved away.
The positive case is real
It would be a mistake to read this as an argument that state-changing technology is inherently suspect. The clearest cases run the other way. The 2021 case report by Khambhati and colleagues, on a patient with severe, treatment-resistant depression, described closed-loop stimulation that detected a neural signature of low mood and intervened; the patient, who had been largely unable to act on her own preferences, regained the ability to do so. By the layered account, this is not a compromise of agency. It is a restoration of it. Her second-order stance finally had purchase on her first-order states, because the machinery that had drowned it out was being regulated. The control-theoretic details of that loop are the subject of The Mind as a Control System.
The same logic extends to less dramatic tools. Meditation aids, sleep interventions, attention scaffolds, and behavioral programs all change the state from which a decision is made. People seek them out precisely because they want their momentary impulses to align with their long-term commitments — to drink less, sleep more, be less reactive with their children. The right comparison is not “manipulated versus untouched.” It is whether the combined human-plus-system arrangement leaves the person more able to recognize and pursue their own values.
The distinction between helping a person act on reflective commitments and making a person easier to direct is doing genuine work here. It is the same distinction, applied at different scales. A tool that helps an alcoholic refuse a drink by surfacing his own stated reasons is on one side of it. A tool that raises his craving while leaving his reasons untouched is on the other. Both change state. Only one respects the layering.
Where consent breaks down
Consent is the standard safeguard and the weakest one at the margin. A worker consents to algorithmic management because refusing costs the job. A user in an altered state consents to continue the intervention that produced the altered state. A platform discloses its data practices fully while designing a service whose economics reward dependence. In each case the form of consent is satisfied and the substance is not, because the person is not positioned to weigh the choice they are ostensibly making.
The hardest case is self-authorizing escalation. Suppose an intervention changes a preference. The changed preference then endorses more of the same intervention. The formal record — the checkbox, the consent form, the accepted terms — becomes a mechanism for ratcheting past the point the person would have agreed to at baseline. This is the structure of addiction, and it is also the structure of any system that lets a modified state authorize its own intensification. The safeguard that answers it is a commitment to baseline control: consequential changes should be reversible, inspectable, and capable of being reaffirmed when the person is not under the state-changing intervention. Medicine already practices a version of this when it revisits, once a patient is stable, a plan set during an acute episode.
A related trap is treating discomfort as evidence that the person needs modifying. Sometimes anxiety is excessive, and treating it restores function. Sometimes the anxiety is a correct response to a real danger, and treating it is a way of adapting a person to a bad situation. Sometimes the fatigue is a disorder; sometimes the body needs sleep. A system optimized only for the absence of distress cannot tell those apart, and neither can an institution that rewards it for reducing complaints. The information the optimizer lacks is not more data about the person. It is an account of what the person’s states are for.
Reversibility, inspection, and the right to say no
If authorship is the value at stake, the operational test is surprisingly concrete. Can the user disconnect? Can they inspect the record of what was inferred and done? Can they choose a different system, recover their data, and return to a baseline state without penalty? A system can be formally voluntary and functionally sovereign at the same time. The test distinguishes them.
This is where the layered account meets institutional design. A right to explanation, a right to contest an inference, a right to exit — these are procedural, but they are the procedures that keep the second-order layer from being co-opted without the person’s knowledge. They are what make the difference between a person deciding from within a shaped state and a person being decided through. The 2024 EU AI Act’s prohibitions on subliminal and exploitative manipulation, gated on significant harm, represent one legal attempt to draw this line; how far it reaches into ordinary persuasive systems remains an open question.
What would change the argument
For empirical claims, the demand is replication and transfer: findings that survive beyond a single platform, population, or model version. The personalized-persuasion literature is already teaching this lesson, with contradictory findings and corrections arriving faster than consensus. Neural Advertising works through the same evidentiary problem in the commercial context.
For the moral and theological questions, the demand is different. Neuroscience can describe the correlates of conviction, decision, and craving without settling whether a belief is true or whether a choice was right. An AI can summarize a religious tradition’s teaching on free will without becoming a source of its authority. The frameworks that answer “whose decision is it” in the deepest sense — theological accounts of responsibility, moral philosophy’s accounts of respect — are not confirmed or refuted by a scanner. They are the terms in which the empirical findings get their human meaning. Treating them as frameworks rather than as hypotheses is not a concession to vagueness. It is the correct description of what kind of claim they are.
The boundary to defend
What is at stake is not whether machines will replace human choice but whether the conditions for meaningful choice will be quietly hollowed out from below. The endpoint worth guarding against is not a world of puppets. It is a world of people who act, sincerely believe they are choosing, and cannot quite locate the source of their own preferences — a world where the layered structure of agency has collapsed without anyone noticing. The dystopian novelists saw the first version of this coming, and the version that actually arrives is likely to be quieter; that contrast is the argument of Brave New World Was Closer Than 1984.
The defense is not a blanket refusal of influence. Influence is how humans have always formed each other, and it is what makes teaching, friendship, and love possible. The defense is a set of commitments about transparency, reversibility, and the preservation of a stance: that a person can always, in principle, step back from the state they are in and look at it, and have that look count. Where a system preserves that, it can be extraordinarily useful — a tutor, a scaffold, a prosthetic for a capacity that has failed. Where it removes it, no amount of voluntariness in the final click repairs the loss. The decision may still be signed by a person. Whether it is authored by one is the question that the design of these systems will settle, piece by piece, long before anyone declares free choice over.
Sources and further reading
- Frankfurt, H. G. “Freedom of the Will and the Concept of a Person.” The Journal of Philosophy 68(1), 1971. https://www.jstor.org/stable/2024717
- Susser, D., Roessler, B., & Nissenbaum, H. “Online Manipulation: Hidden Influences in a Digital World.” Georgetown Law Technology Review 4(1), 2019. https://georgetownlawtechreview.org/online-manipulation-hidden-influences-in-a-digital-world/GLTR-01-2020/
- Susser, D., Roessler, B., & Nissenbaum, H. “Technology, autonomy, and manipulation.” Internet Policy Review 8(2), 2019. https://policyreview.info/articles/analysis/technology-autonomy-and-manipulation
- Coons, C., & Weber, M. (eds.). Manipulation: Theory and Practice. Oxford University Press, 2014. https://academic.oup.com/book/4870
- Raz, J. The Morality of Freedom. Oxford University Press, 1986. https://philpapers.org/rec/RAZAAP-2
- Sass, R. “Manipulation, Algorithm Design, and the Multiple Dimensions of Autonomy.” Philosophy & Technology 37(3), 2024. https://doi.org/10.1007/s13347-024-00796-y
- Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. “On the conversational persuasiveness of GPT-4.” Nature Human Behaviour 9(8), 2025. https://www.nature.com/articles/s41562-025-02194-6
- Costello, T. H., Pennycook, G., & Rand, D. G. “Durably reducing conspiracy beliefs through dialogues with AI.” Science 385, 2024. https://www.science.org/doi/10.1126/science.adq1814
Loading comments…