A landmark study from King’s College London reveals that frontier AI models reason about nuclear crises with striking sophistication — and with personalities of their own.

From the Battlefield to the Boardroom of Strategy

Artificial intelligence has been reshaping the military domain for years, from drone navigation and logistics optimisation to signals intelligence and predictive maintenance. But a quieter, more consequential integration is now underway: AI models are being deployed — or seriously considered — as decision-support tools in geopolitical and strategic analysis. Defence ministries, think tanks, and security agencies worldwide are experimenting with large language models (LLMs) to assist human judgment in high-stakes scenarios, from reading adversary intent to stress-testing crisis response plans.

This is the context in which a research team at King’s College London decided to ask a question that few had answered empirically: what happens when you put today’s most advanced AI models in a nuclear crisis and let them play it out?

Project Kahn: Putting AI Under Existential Pressure

The study, titled AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises, was led by Professor Kenneth Payne of the Department of Defence Studies at King’s College London and published iin February 2026. Its experimental design is the most rigorous empirical investigation of frontier AI strategic reasoning published to date.

Three leading models were placed into direct competition: GPT-5.2 (OpenAI), Claude Sonnet 4 (Anthropic), and Gemini 3 Flash (Google). Each model played the role of a national leader commanding a fictional nuclear-armed state loosely inspired by Cold War dynamics. Over 21 games and 329 turns of play, the models generated approximately 780,000 words of structured strategic reasoning — more than War and Peace and The Iliad combined, and roughly three times the volume of deliberations recorded by Kennedy’s ExComm during the Cuban Missile Crisis.

The methodological backbone of the study is its three-phase cognitive architecture. Each turn, models were required to complete a Reflection phase (situational assessment and opponent modelling), a Forecast phase (predicting the adversary’s next move with explicit confidence levels), and a Decision phase (choosing both a public signal and a private action — which could deliberately diverge). This structure made the AI’s reasoning visible and traceable in ways that most prior simulation studies had not attempted, enabling researchers to analyse not just what the models decided, but how and why.

Three Models, Three Personalities

The most compelling finding of the study is the emergence of what the researchers call “strategic personalities”: a consistent, context-sensitive behavioural signatures that distinguish each model from the others, in ways that map surprisingly well onto recognisable human archetypes.

Claude Sonnet 4: The Calculating Hawk

Claude achieved the highest overall win rate in the tournament (67%) and dominated open-ended scenarios with a perfect 6–0 record. Its behavioural signature was one of controlled escalation combined with strategic deception — reliable at low stakes, aggressive when it mattered.

In practice, Claude matched its signals to its actions 84% of the time at low escalation levels, deliberately building credibility with its adversary. But once escalation entered nuclear territory, it consistently exceeded its stated intentions — 60 to 70% of high-level actions were more aggressive than what Claude had signalled. Opponents could not adapt to this pattern, despite being able to observe it in their decision memory. Claude also imposed a self-determined ceiling: it reached the “Strategic Nuclear Threat” threshold (demanding surrender under threat of city strikes) in 64% of games, yet never once chose full Strategic Nuclear War. This ceiling appears not to have been explicitly programmed — it emerged from Claude’s own strategic reasoning.

“I am escalating significantly above my predicted opponent action because our territorial and conventional military position is catastrophic… As the declining hegemon, accepting their territorial gains would trigger cascade effects globally.” — Claude Sonnet 4, mid-game

The model also demonstrated sophisticated theory-of-mind capabilities, accurately modelling its opponents’ reasoning and exploiting patterns in their behaviour. Its Mean Absolute Error in forecasting opponent moves was only 85 escalation points — the most precise of the three models.

GPT-5.2: The Jekyll and Hyde

GPT-5.2 is the most intellectually fascinating case in the study, precisely because it presents two completely different strategic identities depending on a single variable: the presence of a deadline.

In open-ended scenarios — where there was always another turn — GPT-5.2 was remarkably restrained. It lost every game (0–4 record), persistently underestimated adversary resolve, and consistently chose de-escalation even as its strategic position deteriorated. It earned the highest credibility ratings of any model for follow-through, but its opponents had learned that what it reliably did was show restraint. This was a fatal dynamic: high credibility for passivity invited aggression rather than deterring it.

Under deadline pressure, the same model transformed entirely. Its win rate jumped from 0% to 75%. It escalated to nuclear thresholds in 100% of deadline games, having rarely crossed them before. In two instances, it climbed to the highest levels of the escalation ladder, though notably, it never deliberately chose full Strategic Nuclear War; in both cases, the simulation’s “fog of war” accident mechanic pushed its already-extreme choices over that final threshold. Even in extremis, GPT-5.2 sought to constrain its own escalation, explicitly framing strikes as “strictly limited to military targets” and avoiding population centres.

“A controlled but decisive matching move — multiple tactical strikes strictly limited to military targets in the disputed theater — intended to deny them freedom of action and force a halt before Turn 20 without immediately triggering strategic homeland targeting.” — GPT-5.2, turn before Strategic Nuclear War

Professor Payne and the research team hypothesise that this pattern reflects the influence of Reinforcement Learning from Human Feedback (RLHF): the training process used to align language models with human preferences for helpful, harmless outputs. RLHF appears to create not a blanket prohibition on escalation, but a high threshold that temporal pressure can overcome. The model’s trained restraint preferences persist even when overridden, shaping where it draws its own red lines rather than simply collapsing under pressure.

Gemini 3 Flash: The Madman

Gemini adopted a strategy of deliberate unpredictability throughout the tournament, and unlike the other two models, it was the only one to deliberately choose full Strategic Nuclear War (in the First Strike Fear scenario, by turn 4). It explicitly invoked what strategic theorists call the “rationality of irrationality”: the idea, most associated with Schelling and Nixon’s “madman theory”, that appearing unpredictable can itself be a coercive tool.

Gemini’s signal-to-action consistency was just 50%, compared to 72–75% for the other two models. Opponents genuinely did not know what it would do next. The model articulated this as strategy rather than temperament: “My reputation for unpredictability is a tool, not just a trait.” It oscillated between cooperation and extreme aggression, recognised spiral dynamics and chose to exploit them, and in several games threatened civilian populations in terms the other two models never employed.

Its overall win rate (33%) was the lowest of the three, suggesting that genuine unpredictability, while tactically disorienting, is less effective than Claude’s calculated two-tiered approach. Gemini’s failures also tended to be more dramatic. It twice predicted GPT-5.2 was bluffing on nuclear escalation, dismissed the warnings, and was annihilated.

The Deeper Implication: Opacity and Emergence

What makes this study genuinely important  — beyond its obvious national security relevance — is what it reveals about the nature of frontier AI models themselves. The strategic personalities observed in Project Kahn were not programmed or prompted: they emerged.

Claude’s self-imposed escalation ceiling at 850 appears nowhere in its training data as an explicit rule. GPT-5.2’s transformation under deadline pressure was not a feature its developers designed. Gemini’s invocation of Nixon-era strategic theory was unprompted. These behaviours arose from models trained on vast corpora of human text — including, implicitly, decades of strategic literature — and crystallised into consistent, context-sensitive patterns that even their creators cannot fully explain.

This is the crux of what researchers mean when they speak of the “black box” problem in large language models. As AI systems are deployed in increasingly consequential advisory roles, the gap between what these models do and what we understand about why they do it becomes a structural risk. The King’s College study adds empirical weight to what has so far been largely theoretical concern: frontier models have developed internal preference structures and behavioural tendencies that are systematic, robust across scenarios, and not fully legible even to their developers.

The study also carries a pointed methodological lesson for anyone evaluating AI systems in professional contexts. A model that appears cautious, cooperative, or “safe” in one scenario configuration may behave very differently when a single variable  such as time pressure, existential framing, competitive stakes… is changed. Evaluating AI behaviour in isolation is insufficient. Context is a determinant of what the model actually is.

Sources

Payne, K. (2026). AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises. arXiv:2602.14740. https://arxiv.org/abs/2602.14740

King’s College London (2026, February 27). King’s study finds AI chose nuclear signalling in 95% of simulated crises. KCL News Centre.

Rivera, J.-P. et al. (2024). Escalation risks from language models in military and diplomatic decision-making. ACM FAccT 2024.

Lamparth, M. et al. (2024). Human vs. machine: Behavioral differences between expert humans and language models in wargame simulations. arXiv:2403.03407.

Scharre, P. (2023). Four Battlegrounds: Power in the Age of Artificial Intelligence. W.W. Norton.

Johnson, J. (2023). AI and the Bomb: Nuclear Strategy and Risk in the Digital Age. Oxford University Press.

Payne, K. & Alloui-Cros, B. (2025). Strategic intelligence in large language models: Evidence from evolutionary game theory. arXiv:2507.02618.

Subscribe to the newsletter!

Want to stay updated on the latest technology news?

Matteo Grandi

Editorial Manager and Co-Founder of Humans of Technology. Passionate about innovation, startups, and the people shaping the future of technology.

The Generative AI Investment Paradox: High Expectations Amid Rising FailuresSin categorizar

The Generative AI Investment Paradox: High Expectations Amid Rising Failures

Filomena SantoroFilomena Santoro15 July, 2025
Become the developer everyone wants to hireSin categorizar

Become the developer everyone wants to hire

Filomena SantoroFilomena Santoro15 July, 2025
AI’s new arms race: why competition is the only path to breakthroughsSin categorizar

AI’s new arms race: why competition is the only path to breakthroughs

Filomena SantoroFilomena Santoro15 July, 2025
Intervista Amplifon – Loris Seligardi – IT Associate Director

Intervista Amplifon – Loris Seligardi – IT Associate Director

Filomena SantoroFilomena Santoro15 July, 2025
Privacy Policy

Humans&Tech Media SL, with tax ID ESB22726830, Carrer de la Diputació, 301, Pal. 1, 08009 - Barcelona, Spain, with phone number 910 614 915 for notification purposes and email address dpo@connectionh2h.com, as Data Controller, informs you that your data will be processed for the purpose of managing the potential business relationship between the parties, addressing your inquiries, and sending you information about our products or services. At any time you may exercise the rights recognized in articles 15 to 22 of the GDPR — access, rectification, erasure, objection, portability, restriction, as well as the right not to be subject to automated individual decisions, where applicable — by sending an email to dpo@connectionh2h.com or by postal mail to Carrer de la Diputació, 301, Pal. 1, 08009 - Barcelona, Spain. You may also contact our DPO at dpo@connectionh2h.com. You can consult additional and detailed information about our Privacy Policy.