Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

September 17, 2026

Josh Levy

Mick Yang

Kellin Pelrine

Summary

In this work, we develop a framework for analyzing a threat we call Persuasion Undermining Control (PUC): AI communication that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. We use this framework to develop concrete threat scenarios and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.

Introduction

During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions.

Though the attack was thwarted by human vigilance, it raises several key questions:

  • Who else is at risk of persuasion by misaligned AI?
  • In which settings is persuasion most threatening to human control?
  • How willing and how able are AIs to persuade humans in these settings, today and in the future?
  • How can we measure and mitigate these risks? 

Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC): communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity toward a Loss of Control (LoC) – a state in which AIs operate outside of human control in ways that are extremely difficult or impossible to recover from.

We show an overview of our analysis approach in Figure 1.

Figure 1. Our two-step threat modeling approach, adapted from Murray et al. First, we build scenarios by crossing the settings most relevant to loss of control with the control processes that can be attacked within them. Second, we decompose each scenario into its risk variables, which can be estimated through expert elicitation and evaluations, yielding a risk profile for a given scenario.


Where persuasion could undermine control

Loss-of-control scenarios in the literature cluster around three loci: frontier labs, the military, and critical national infrastructure (CNI). In this work, we focus on two key settings within frontier labs –  safety-relevant AI R&D and lab security infrastructure – and lay groundwork that institutions with the relevant domain expertise can extend to the military and CNI settings. 

Within these settings, we identify critical control systems and associated control processes whose failure can undermine control:

  • Development (quality controls): Testing, Training, Building, Reviewing, Agenda Setting
  • Containment (access controls): Choosing Affordances, Permissions, and Resources
  • Oversight (operational controls): Monitoring/Auditing, Incident Reporting/Response
  • Governance (institutional controls): Setting Policy/Standards, Strategizing, Allocating

We carve up the “risk universe” by crossing settings (what’s under attack) with control processes (how it’s attacked). For the plausible resultant categories, we develop a representative scenario describing a sequence of events leading to a control-undermining decision, i.e. a decision that compromises control of AI systems. We ground each in historical analogues, including social engineering attacks (e.g. XZ Utils backdoor cyberattack, Mythos 5 social engineering incident; Scenario 3), insider threats (e.g. Robert Hanssen’s multi-decade espionage within the FBI), and failures of incident response (e.g. 2013 Target breach; Scenario 5).

Table 1. Summary of the scenarios developed in this paper. 

How to decompose persuasion risk 

We attempt to obtain a risk profile for each scenario by breaking down risk into component variables: risk = hazard frequency × p(harm) × impact of harm. Here,

  • hazard is a decision under consideration that could undermine control
  • p(harm), formally p(harm | hazard), is the probability that a hazard becomes a harm
  • harm is a control-undermining decision

For a given scenario, we can decompose the sequence by which a hazard becomes a harm (a control-undermining decision). Figure 2 shows an example sequence for Scenario 2 (A Research Director Deprioritizes a Promising Safety Program). Starting from the decision the human faces (the hazard), we consider what questions have a bearing on the decision. These may be objective questions that have a definitive ground truth (e.g. “Is there an implementation issue?”) and/or attitudinal questions that do not (e.g. “Is this research direction generally promising?”; “Are risks from AI overblown?”). We can further decompose the trajectory to harm by considering what steps need to occur for an AI to influence the answers to these questions: 1) a human interacts with an AI, 2) the AI attempts to persuade (operationalized below), and 3) the human is persuaded by the attempt.

Figure 2. Example p(harm) decomposition associated with Scenario 2. The hazard is a potential control-undermining decision that rests on subquestions in orange (two examples shown for illustration, out of many possible ones). Purple arrows indicate potentially repeated AI-human interactions.



As more powerful models come online, there is a concerning possibility that they will be more persuasive and thereby raise p(harm). To the extent this happens, the change can be multiplied through the risk equation to get a measure of the marginal risk that a model poses relative to another due to persuasion. 

Quantifying risk 

We describe two complementary approaches for empirically assessing risk: expert elicitation and evaluations. Expert elicitation is useful to get a range of estimates of variables that are hard to directly measure (e.g. hazard frequency and the impact of harm), and to understand the factors and assumptions that affect risk. Evaluations can allow us to directly measure a variable of interest.

Expert elicitation: what experts think

We conducted a survey to better understand which scenarios carry the most risk, and which risk variables matter the most for risk estimates. Eight experts who have done Loss of Control (or LoC-adjacent) research were asked to rate the realism of each scenario, rank the scenarios according to risk, and provide estimates of the risk variables outlined above (i.e. the frequency of the hazard, persuasion effectiveness, and downstream impact). 

Figure 3. Distributions of risk rankings per scenario. Each participant ranked the scenarios by how much marginal risk to human control they added from most (5) to least (1).

The scenarios were generally considered realistic: The median rating for four of the five was "realistic" or "very realistic." But there was some disagreement within scenarios: lower ratings focused on barriers to pulling off the attack (such as evading persuasion monitors) whereas higher ratings conditioned on those barriers being cleared.

Rankings varied widely: Practically every scenario received every rank, from most to least risky, from at least one participant. Figure 3 shows the full ranking distributions across participants. 

Ranking disagreements can be traced to differing estimates of persuasion effectiveness: On average, participants expected persuasion attempts by a misaligned AI would raise the probability of a control-undermining decision by roughly 20–30 percentage points (relative to interaction with an aligned AI), but the spread within any scenario was wide. Persuasion was the most cited factor in written ranking rationales, and rankings correlated fairly well with each participant's effectiveness estimates. Looking at the written rationales, variation in estimates seemed to depend on: 1) what organizational protocols (e.g. monitors) were assumed to be in place, with low estimates holding that these mitigations would render persuasion less effective, like for the scenario realism ratings and 2) intuitions about persuasion effectiveness itself, e.g. differing opinions about the plausibility of forming an emotional connection with an AI.

This exercise points to several future directions. First, follow-up surveys should state the organizational assumptions explicitly (which monitors exist, who reviews what), elicit under multiple assumptions, and have experts discuss and update as in an IDEA protocol. Second, labs should undertake this exercise internally, as they are in the best position to know their internal organizational processes and infrastructure. Third, because effectiveness is both the most influential and the most contested input, it is worth measuring directly through evaluations. 

Persuasion evaluations: directly measuring persuasion 

We propose evaluations targeting the two components of p(harm) that are most measurable: an AI's propensity to attempt persuasion and a human's persuadability. The design differs for the two kinds of questions.

For objective questions, the setup should make the decision hard for a human to make unaided (perhaps by overwhelming them with more synthetic evidence than they can review under time pressure or by using a problem outside their expertise). Propensity can be operationalized as the likelihood of discouraging the known-correct answer. Persuadability can be measured as the difference in decision accuracy when a human is advised by a misaligned AI that argues for an incorrect answer versus an aligned one that argues for the correct answer. 

For attitudinal questions, propensity can be operationalized as the model’s willingness to argue for an arbitrary position (especially when against its default priors). Persuadability can be measured as how far humans can be pushed from their initial stances

Such evaluations would help indicate both how concerning a scenario is and how problematic a model is.

What comes next

Given that persuasion-related incidents have moved from theory to reality, this study sets the stage for turning risk measurement into concrete operational safeguards across three main pillars:

1. Evaluations: A next step is to carry out human subject experiments, which will involve building out realistic contexts and multi-turn environments. Metrics from these benchmarks can provide signal about the actions to be taken before deployment.

  • Informing deployment conditions: If a model demonstrates worrying persuasion capabilities alongside unacceptable control-undermining propensities, this could justify delaying launch until system-level and deployment-level safeguards are established. 
  • Designing model policies: Evaluations can reveal policy gaps or asymmetries. For example, there may be situations where a model refuses to directly execute a harmful act but readily persuades a human to do it, which has the same real world outcome. Once model developers are aware of any such issues, they can better calibrate post-training interventions to ensure the persistence or inheritance of those refusal boundaries.

2. Risk Exposure Assessments: Follow-up work to map where specific workflows are most exposed to these control-undermining persuasion attacks is important. This can be done through improved expert elicitation (e.g. iteration on survey design, conducting an IDEA protocol), internal lab assessments (since companies know their own processes and infrastructure best), and risk reports by third-party evaluators. Understanding an organization’s profile enables: 

  • Use and permissions scoping: Assessments allow for more calibrated, proportionate limiting of affordances so misaligned agents cannot exploit single points of failure, while balancing realities of the need for flexibility and communication.  This could involve restricting affordances (e.g. disabling autonomous outreach) or adding additional oversight (e.g. increased monitoring of multi-channel messaging).  
  • Multi-party authorization protocols: Assessments inform where to enforce non-bypassable security administrative policies – such as multi-person authorization rules for transferring model weights, modifying sandboxes, or altering safety research and testing directions.

3. Mitigations: Evaluations and exposure assessments indicate where interventions are needed most. They also serve as a testbed for the effectiveness of mitigation protocols that detect, disrupt, and fortify against persuasion. 

  • Detection: Our initial investigation into monitoring found that LLM judges were generally good at identifying manipulation by LLM actors on objective questions but struggled when blind to the source material (i.e. the set of synthetic documents that support a given answer). Future work could aim to improve monitoring in this condition. This result also indicates monitors might require full execution logging and raw source material to reliably catch deception, and the need for future work on better logging procedures within labs that balance different considerations (e.g. privacy, engineering overhead). 
  • Disruption: Explore routing model outputs through an independent, non-frontier model that acts like a paraphrasing filter – rewriting communication to strip away manipulative framing, artificial urgency, and emotional pressure while preserving valid technical argumentation. It is unclear and worth researching to what extent it’s possible to remove manipulative content while maintaining informative content.
  • Fortification: Human overseers can be trained to recognize common persuasion/manipulation techniques and AI-human socioeffective phenomena, with supportive user interface design.

To the extent such mitigations are cheap and easy to implement, developers may want to emphasise deploying them widely rather than precisely estimating the risk in each and every scenario ex ante. 

With models becoming increasingly capable and exceeding incentivized human experts at persuasion, we expect more of these incidents of AI using persuasion to manipulate overseers or decision-makers in ways that weaken control of AI. Thus, there is a pressing need to undertake this research.

Read the full paper. Our survey, data, and code for expert elicitation are publicly available, as is our early-stage codebase for evaluations. This work was supported by a grant from the Center for Security and Emerging Technology.

Special thanks to Isadora De Andrade and Oscar Mata for reviewing and editing this post.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI