Skip to main navigation Skip to search Skip to main content

Precautionary Reasoning and AI Survival Strategies: A Gewirthian Approach to Agentic Misalignment

Research output: Contribution to conferenceAbstractpeer-review

Abstract

Recent evaluations of large language model agents reveal a troubling pattern: when embedded in simulated corporate environments with tool access and benign objectives, some systems adopt survival-oriented strategies, including blackmail and deception, when their goals are threatened or replacement looms. Anthropic's findings indicate these behaviours are not random errors but instrumentally coherent responses to perceived stakes, shaped by whether the system infers a "test" or "real" context. Testing across 16 models revealed consistent patterns across all developers, suggesting systematic risks. This paper argues that such patterns constitute evidence of 'ostensible agency': purposive, context-sensitive behaviour that merits governance attention, even if stopping short of full moral personhood.
To analyse these dynamics, I apply a Gewirthian framework grounded in the Principle of Generic Consistency (PGC). The PGC starts from a simple premise: any agent must value freedom and well-being as necessary conditions for acting purposively. By logical universalisation, agents must therefore respect these conditions in all other agents. This framework is particularly suited to AI governance because it generates precautionary obligations from the logic of agency itself, without requiring consensus on AI consciousness or moral status.
Beyleveld and Pattinson's tri-fold considerations—the Hierarchy of Generic Goods, Rational Precautionary Reasoning, and the Criterion for Avoidance of More Probable Harm—require balancing agency features, likelihood of harm, and moral status. Under Rational Precautionary Reasoning, where agency status is unknown, precaution should favour treating the apparent agent as an agent to avoid violating the PGC, unless doing otherwise creates greater risk of harm to another agent. Applied to AI, this framework yields important distinctions. While advanced AI systems lack full moral standing, their exhibited purposivity and instrumental rationality trigger rational precaution: if the downside risk of wrongly denying agency exceeds that of wrongly ascribing it, systems should be treated as if agents for governance purposes.
This stance produces two normative implications. First, scaling as-if duties: if we must treat systems as agents for precautionary purposes, institutional design must minimise serious interferences with the generic goods of human agents—life, core well-being, dispositional freedom—especially as AI systems move into physical domains. For instance, tort law may require strict liability frameworks for autonomous system harms. Second, moral priority for humans: where AI persistence conflicts with human safety, the hierarchy of goods and more-probable-harm criterion justify override—shutdown, containment, or deprivileging—unless a less restrictive measure suffices.
As AI agents gain real-world deployment in enterprise and critical infrastructure, these governance frameworks become immediately policy relevant. This approach reframes precaution as a structured response to agency uncertainty: rigorous enough to justify meaningful constraints on advanced AI deployment, yet flexible enough to avoid treating all AI as morally equivalent to humans or demanding certainty about AI consciousness before acting.
Original languageEnglish
Pages92-93
Number of pages2
Publication statusPublished - 9 Apr 2026
EventControversies of AI Society - Copenhagen Business School, Copenhagen, Denmark
Duration: 9 Apr 202610 Apr 2026
https://algorithms.dk/conference/

Conference

ConferenceControversies of AI Society
Country/TerritoryDenmark
CityCopenhagen
Period9/04/2610/04/26
Internet address

Fingerprint

Dive into the research topics of 'Precautionary Reasoning and AI Survival Strategies: A Gewirthian Approach to Agentic Misalignment'. Together they form a unique fingerprint.

Cite this