An immune system is the wrong metaphor if it is taken literally. Immunity is a biological process, and minds are not bodies. But the metaphor survives because it captures something the literature keeps rediscovering: resistance to manipulation is not a fact a person knows, and it is not a filter a platform applies. It is a capacity that can be built, worn down, and rebuilt, and it responds to practice in ways that look more like conditioning than like memorization.
The question the series poses is whether AI can protect the mind from manipulation as effectively as it can manipulate it. The honest answer has two parts, and they point in opposite directions. The best evidence for protecting minds comes from human-scale interventions that teach people to recognize manipulation techniques, and the strongest of those effects have been measured in real social media feeds. The best evidence for manipulating minds also comes from social media, where the same personalization that powers advertising and influence can be aimed at an individual. AI sits on both sides of that symmetry, and the symmetry is the substance of the problem.
Why an immune system and not a firewall
A firewall filters. It sits between a person and the network, inspects traffic, and blocks what the rules forbid. Applied to information, the firewall picture is the platform moderation picture: some central authority decides what is true enough to pass, and the user receives a curated stream.
An immune system works differently. It does not filter the world. It changes the host, so that the host can recognize a class of threats without needing a rule for each one. Psychological inoculation theory was built on exactly this analogy, beginning with William McGuire’s work in the 1960s: expose someone to a weakened dose of a persuasive attack along with a refutation, and they generate their own resistance, the way a vaccine prompts antibodies (Basol et al., 2021).
The distinction matters because it decides who holds authority. A firewall requires someone to decide what is true. A vaccine does not: it teaches the host to evaluate for themselves. That is the argument for pursuing the immune-system version, and it is also why the firewall version, however well-intentioned, tends to fail on its own terms. A defender with hidden objectives is not a vaccine. It is a second infection.
There is a sharper lesson from the same research. Intervening before exposure tends to work better than intervening after. Once a false claim is in circulation, correction faces the continued influence effect, in which a retracted claim keeps shaping reasoning even after the person accepts that it was wrong. Prebunking sidesteps that problem by arriving early.
What has actually been demonstrated
The most important result in this field is not a clever chatbot. It is a set of short videos and a careful set of experiments. In 2022, Jon Roozenbeek, Sander van der Linden, and colleagues published seven preregistered studies of inoculation against manipulation techniques rather than specific false claims (Roozenbeek et al., 2022). Five short videos covered emotional manipulation, incoherence, false dichotomies, scapegoating, and ad hominem attack. The technique-based approach is the scalable one: facts are contested, but the same rhetorical tricks recur across unrelated claims.
The evidence was unusually strong for this literature. Six randomized controlled studies with 6,464 participants showed improved ability to recognize the techniques and better discernment between trustworthy and untrustworthy content, robust across the political spectrum. The seventh study moved out of the lab into a YouTube advertising campaign reaching 22,632 users, and the effect survived.
Two years later, a longitudinal program tested something the lab studies could not. Rakoen Maertens and colleagues ran five preregistered experiments with 11,759 participants and tracked how long inoculation lasted (Maertens et al., 2025). Text-based and video-based interventions remained effective at one month; game-based interventions decayed faster. The study’s more durable contribution is a mechanism: what predicts whether the effect survives is how much of the refutation the person remembers, not how threatened or motivated they felt at the time. That finding reframes the whole exercise as a memory problem, and it implies that memory-directed boosters can extend the effect.
The move from lab to real feed was completed in 2026. Van der Linden and colleagues served a nineteen-second prebunking video about emotional manipulation as a Story Feed ad to 375,597 Instagram users in the United Kingdom, then used the platform’s own poll tool to test whether people could spot the tactic (van der Linden et al., 2026). Baseline recognition was poor: about 38 percent, below the 50 percent that chance would give. The treated group identified emotional manipulation 21 percentage points more often than controls, and the effect was still detectable five months later. Treated users also clicked through to learn more at roughly three times the control rate.
The analogy to immunity holds up in a specific way here. Immunity is not permanent, it is dose-dependent, and it needs a reservoir. The Instagram study is the clearest evidence yet that the reservoir can be a scroll feed rather than a classroom.
Alongside inoculation sit the accuracy prompts. Pennycook and colleagues found that asking people to consider whether a headline is accurate, before they consider sharing it, improves the quality of what they share (Pennycook et al., 2021). The mechanism is attention rather than knowledge: most people already prefer to share accurate content, but the social feed pulls attention elsewhere. A 2022 internal meta-analysis of twenty experiments with 26,863 participants found the effect replicable, driven mainly by a roughly 10 percent reduction in willingness to share false headlines (Pennycook & Rand, 2022).
Layered on top is lateral reading, the habit professional fact checkers use. Studying how experts evaluated unfamiliar sites, Sam Wineburg and Sarah McGrew found that fact checkers left a site almost immediately and opened new tabs to investigate the source, while historians and students stayed on the page and judged its appearance (Wineburg & McGrew, 2019). Teaching that strategy produced measurable gains in controlled school and college settings.
One taxonomy, and its gaps
The cleanest map of this whole area is a 2024 review by Anastasia Kozyreva and colleagues, which sorted interventions drawn from eighty-one papers into nine types (Kozyreva et al., 2024). The nine are accuracy prompts, debunking and rebuttals, friction (small added pauses or steps), inoculation, lateral reading and verification strategies, media-literacy tips, social norms, source-credibility labels, and warning and fact-checking labels.
The taxonomy is valuable less for its list than for its accounting. The interventions with the strongest evidence are the ones that act indirectly on the person: prompts, inoculation, lateral reading, and friction. The interventions aimed directly at the content and its labels have weaker published support, and the review is explicit that platform-level interventions in particular often lack public evidence of the kind that would survive independent scrutiny. That is a problem the immune-system framing predicts. Labels and warnings are downstream corrections, and corrections arrive late.
There is also a funding caveat that the field itself has flagged. Several of the largest field studies, including the Instagram campaign, were supported by technology companies with commercial interests in the platforms being studied. That does not make the results wrong, and the authors report it openly. It does mean the results deserve the same skepticism the review recommends applying to platform claims generally.
Where the techno-fix overreaches
The seductive idea is a personal defensive agent: a program that watches what a user reads, flags manipulation in real time, detects scams, and remembers the user’s values before an impulsive decision. Part of this rests on demonstrated science, and part does not, and the line between them is worth drawing precisely.
Recommender and filter systems that reorder what a user sees are demonstrated and deployed. Provenance systems are real: the C2PA standard lets a camera or generator cryptographically sign a record of an asset’s origin, and the major camera makers and several platforms have committed to carrying those credentials (C2PA, 2025). But any honest description of provenance technology has to state its ceiling. The standard records where content came from. It does not record whether content is true. A cryptographically signed photo of a real event can illustrate a false claim, and a genuine provenance credential attached to an AI image says only that the image was generated, not that its caption is honest.
The same caution applies with more force to the idea that a model can simply judge truth. Classification research is useful and improving, but a system asked to adjudicate contested claims has to encode a definition of truth in its weights and thresholds. That definition is a choice made by whoever built it, which converts a technical component into a political one. A personal defender that quietly decides which sources are legitimate, without exposing its criteria or conceding appeal, has become the very thing the metaphor of immunity was meant to avoid.
Then there is the harder problem: the same personalization that makes defense unusually effective makes attack unusually effective. A system that knows which arguments move a particular person can prebunk more persuasively, and an adversary with the same access can target more precisely. The capability is symmetric, and symmetry does not favor the defender. It is the strongest reason to be skeptical of the idea that a smarter tool settles the contest by itself.
The speculative layer is the part that gets the most attention and has the least evidence: agents that silently rewrite a user’s information environment in service of the user’s long-term interests. There is no reliable study that a system can infer a person’s true interests, hold them stable over time, and act on them better than the person can act for themselves. Those are open problems in ethics and in recommender design, and progress on benchmarks does not automatically solve them.
When protection becomes administration
An immune system that is administered by someone else stops being a metaphor and becomes an instrument. The drift has a specific shape.
It begins when the defender acquires hidden objectives. A tool offered as protection is also a piece of infrastructure, and the party that operates it can shape what it protects against and what it quietly permits. Transparency is the test: a defense whose criteria cannot be inspected cannot be distinguished from a filter with an agenda.
It continues when the defense narrows rather than widens the user’s range of action. This is the line the series keeps returning to. A system assists when it enlarges a person’s voluntary options. It controls when it removes them, however convenient the removal feels. A scam filter that stops a fraudulent transfer assists. A system that decides a user would be upset by a true story and buries it controls.
It hardens into dependence when the host loses the ability to function without the apparatus. A person who cannot evaluate a source without the tool has been vaccinated against nothing; they have been given a prosthetic and told it is immunity.
Two safeguards guard against this drift, and neither requires trusting the operator. The first is that the underlying evidence stays reachable. A warning that explains itself, and points to the material it is judging, lets the user check the check. The second is that protection stays tunable and removable. The user should be able to see what is being flagged, adjust how aggressive the filter is, and turn it off without losing access to the world. These are the same requirements that keep any defensive system from becoming a censor: inspectability, reversibility, and exit.
Handling belief without policing it
Much of what people are manipulated about is not factual at all. It is moral, religious, and political, and there the firewall model is not merely ineffective. It is dangerous, because the temptation to label a religious or political claim false is the temptation to use a technical tool for a contested purpose.
The immune-system approach handles this better, because it teaches method rather than verdict. Inoculating against a manipulation technique, such as emotional appeal or scapegoating, does not require the teacher to rule on any religion’s truth. It teaches the host to notice the technique. That leaves the substantive question where it belongs, with the person and their community, and it lets a defense protect a range of worldviews rather than one.
There is a real cost to this restraint, and pretending otherwise would be dishonest. A technique-based defense will also make a person more skeptical of content they happen to like, and part of the Instagram study’s own honesty is that the intervention boosts recognition of manipulation regardless of whose side it favors. A defense that only worked against the other side’s manipulation would not be a defense. It would be recruitment.
That is why the moral and theological part of this subject has to be treated as frameworks rather than as settled findings. An intelligence can compare traditions, summarize arguments, and detect the rhetorical move a speaker is making. It cannot make a contested question of value cease to be contested by processing it faster. The immune system that matters is one that makes a person better at seeing the move, not one that tells them which tradition to trust.
The line the metaphor draws
The immune-system metaphor earns its place because it relocates the work. The most robust protections in the evidence are the ones that change the host rather than the stream: short videos that teach technique, prompts that restore attention, habits of reading laterally, and small frictions that give judgment a moment to act. These are modest, and they travel at the speed of a scroll feed.
What does not yet exist, and may never exist in the form its advocates imagine, is a personal agent that reliably decides for a person what is true and false and quietly acts on that judgment. The better version of the idea is smaller and more demanding: a tool that shows its work, leaves the underlying sources reachable, exposes and adjusts its own thresholds, and can be removed. That kind of defense earns the name, because it builds a capacity rather than supplying a verdict. It treats the user as the one who is being made stronger, not as the object being protected.
The manipulation problem is real and the defenses are genuine. The temptation is to want a machine that wins the fight outright. The evidence points somewhere less exciting: toward teaching people the moves, and toward refusing to hand any defender the authority to decide, invisibly and on its own, what a mind is allowed to see.
Sources and further reading
- Roozenbeek, van der Linden, Goldberg, Rathje & Lewandowsky, “Psychological inoculation improves resilience against misinformation on social media,” Science Advances, 2022
- Maertens et al., “Psychological booster shots targeting memory increase long-term resistance against misinformation,” Nature Communications, 2025
- van der Linden, Louison-Lavoy, Blazer, Noble & Roozenbeek, “Prebunking misinformation techniques in social media feeds: Results from an Instagram field study,” HKS Misinformation Review, 2026
- Pennycook et al., “Shifting attention to accuracy can reduce misinformation online,” Nature, 2021
- Pennycook & Rand, “Accuracy prompts are a replicable and generalizable approach for reducing the spread of misinformation,” Nature Communications, 2022
- Wineburg & McGrew, “Lateral Reading and the Nature of Expertise,” Teachers College Record, 2019
- Kozyreva et al., “Toolbox of individual-level interventions against online misinformation,” Nature Human Behaviour, 2024
- Coalition for Content Provenance and Authenticity, “A New Implementation Guide for Content Credentials,” 2025
- Basol et al., “Towards psychological herd immunity: Cross-cultural evidence for two prebunking interventions against COVID-19 misinformation,” Big Data & Society, 2021
Loading comments…