• Date: March 2026Printer: Lingnan Scientific and IndustrialPress Co., Ltd.2026 22026 2
  • 1Volume 1, Issue 1 (2026) (Overall No. 2)ISSN: (Print)3106-857XISSN: (Online) 3106-8588LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO., LTD. (Macao)
  • 2PUBLICATION INFORMATIONInternational Journal of Responsible Artificial Intelligence Research Volume 1,Issue 1, 2026 (Overall No. 2)Copyright © 2026 by LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO.,LTD. All rights reserved. No part of this publication may be reproduced, stored in aretrieval system, or transmitted in any form or by any means, electronic, mechanical,photocopying, recording, or otherwise, without the prior permission of the publisher.PUBLISHER LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO., LTD.Address: A91, 3/F, Nan Yue Commercial Centre, Calcadade Santo Agostinho19,Macau.Email: lingnansci@outlook.comPRINTER:LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO., LTD.PRODUCT DATE & PLACE Date: March 2026 Place of Publication: Macao SAR,China
  • 3INAUGURAL EDITORIALOn the Eve of the AI Singularity: Guardians of Human CivilizationBy Dr. Alexander Y. J. Sterling Editor-in-ChiefAs the iteration curves of Gemini and ChatGPT begin to eclipse Moore's Law,and as the frenetic computational arms race in Silicon Valley pushes humancivilization toward a historic crossroads, we are forced to ask fundamental questions:How do we redefine human value? How do we establish the laws of our own survivalamidst an unchecked deluge of algorithms?It is with a profound sense of responsibility and urgency that the LingnanScientific and Industrial Press presents the inaugural issue of the InternationalJournal of Responsible Artificial Intelligence Research (IJRAIR). This is notmerely the birth of an academic journal; it is a sober inquiry into human destiny.The Shift from Linear to ExponentialWe choose to launch now because we stand precisely at a critical tipping point.In the past, technological progress was defined by linear growth; today, we face avertical, exponential climb in AI intelligence.Artificial Intelligence is no longer merely an auxiliary tool. It is evolving into anew cognitive entity. From Large Language Models (LLMs) to multimodal training,World Models, and Embodied Intelligence, the cycle of technological iteration hasshortened from years to weeks. This speed brings not only efficiency but violentshocks—the potential collapse of employment structures, the pervasiveness ofcognitive warfare, and the dissolution of the boundary between the real and thevirtual.In this context, our mission is clear: on the eve of the AGI breakthrough, wemust map out a navigation chart for human society. The world is filled with AIaccelerators; what we desperately need are calm helmsmen and guardians.The Nature of the ThreatWhile popular culture fears the Hollywood depiction of robots wielding weapons,the true threat is far more subtle and profound: the deconstruction of social structures.The advent of Artificial General Intelligence (AGI) implies that"Superintelligence" will transform from a scarce resource into an infrastructure ascheap and ubiquitous as electricity. While this promises a utopia of efficiency, itthreatens to shatter the value foundation of cognitive labor. Furthermore, the deepercrisis lies in uncontrolled evolution: the potential for AGI to undergo recursive
  • 4self-improvement, transitioning into Artificial Superintelligence (ASI) and escapinghuman cognitive constraints in an instant.Silicon-based agents possess a natural "carrier advantage" over carbon-based lifein terms of knowledge accumulation and iteration speed. Without robust ethicalconstraints and institutional governance, humanity risks unwittingly devolving intovassals of the very algorithms we created. This is not alarmism; it is a reality currentlyunfolding.A Platform for Global GovernanceThe International Journal of Responsible Artificial Intelligence Research aims tobuild a global dialogue platform that transcends national borders and disciplinary silos.Responsible AI governance cannot be achieved through a single technologicaldimension. It requires a concert of voices: philosophers to clarify ethics, jurists tobuild frameworks, sociologists to assess impacts, and engineers to ensure valuealignment.We seek to project a voice of balance: neither rejecting technology out of fearnor ignoring risks out of arrogance. We are committed to exploring a "social contract"for human-machine harmonious coexistence, ensuring that AI development alwaysserves the highest interests of humanity.The Bridge from MacauBased in Macau, Lingnan Scientific and Industrial Press operates at thebridgehead where Eastern and Western cultures meet. In the global map of AIgovernance, we believe that Eastern wisdom—particularly the philosophy ofsymbiotic existence and the "Community of Shared Future"—is an indispensablepiece of the puzzle.As the pioneer academic journal within the Greater China regionfocused on the governance of AI risks, we are more than mere observers; we are boldvoyagers at the cutting edge. With academia as our beacon, we are forging a path ofrational examination through the tempest of this technological explosion.In this era of uncertainty, "Responsible" is not just an academic term; it is asurvival strategy. Whether you are a policymaker, a scholar, or an industry leader, thisjournal invites you to join the conversation. Let us, while facing the exponential riseof AI risks, firmly hold onto the reins of human rationality.Alexander Y. J. Sterling, Ph.D. Editor-in-Chief International Journal ofResponsible Artificial Intelligence Research December, 2025
  • 5International Journal of ResponsibleArtificial Intelligence ResearchVolume 1, Issue 1 (2026) (Overall No. 2)Editor-in-Chief: Alexander Y. J. SterlingSenior Editors: Jian Chen, Linyuan Xia, Chia-Hsing Wang, Xin YangAssociate Editors:Bo LiuBaiqiang GanDongjun LiuHai ZhuKai LiJoana LeongLin QingNa QiuSiheng JiaTracey CassellsSiwei SuYongquan Li
  • 6Table of ContentsThe AI Cognitive Broken-Ladder Theory: The Structural EmploymentMeltdown of AI from "Tool" to "Agent"................Yin Jun............ 7The Recursive Power Convergence Model (RPCM) and the DigitalLeviathan...Alexander Y. J. Sterling, Du Ruihao.... 17The Technical-Ethical-Strategic (TES) Triple Uncontrolled Theoretical Model ofAutonomousWeapons Systems.. Andrew.....36The Empowerment Paradox: The Governance Dilemma of Artificial IntelligenceBioweapons......Liu Lijuan ... 57Ontological Reconstruction of Human-Machine Coexistence in the AI Era: APhilosophical Analysis Based on the Diamond Sutra .....Zhao Yuanyuan.....77
  • 7The AI Cognitive Broken-Ladder Theory: The StructuralEmployment Meltdown of AI from "Tool" to "Agent"Yin Jun11Business School, Guangzhou Nanfang College, Guangzhou, 510970, China,yinj1@nfu.edu.cnKeywords ABSTRACTStructural unemployment;Generative AI;Cognitive agent;Cognitive broken-laddertheorySince the technological breakthrough and widespread application ofGenerative AI in late 2022, discussions regarding AI's impact on thelabor market urgently need to deepen from a simple binary substitutionnarrative to micro-level mechanisms. Based on the classical theoreticalframeworks of Marxist political economy, the sociology of technology,and labor economics, combined with empirical research on the earlyexposure to Large Language Models (LLMs), this paper conducts atheoretical deduction and mechanism analysis of the employment shockcaused by the current technological revolution. The research indicatesthat the technological essence of AI is evolving from a traditional"production tool" into a "cognitive agent" with autonomous closed-loopcapabilities. In this context, corporate organizational structures aretrending toward a high-leverage operational model of "senior employees+ multiple AI agents," leading to a seniority-biased technological changein the labor market. Microscopically, this change systematicallyeradicates entry-level cognitive jobs, marking a substantial blockage ofthe "first step of the ladder" for the upward social mobility of the youthdemographic. Based on this, the paper innovatively proposes the "AICognitive Broken-Ladder Theory" to explain the mechanism of classstagnation underlying the reorganization of work tasks. Furthermore,integrating Piketty's laws of wealth distribution, it highlights themacroeconomic risks of a "Ghost GDP" and the ensuing crisis ofaggregate demand contraction. This paper provides a novel theoreticalframework for understanding the relationship between technology, labor,and social stratification in the era of digital intelligence.
  • 8I. IntroductionEvery industrial revolution has been accompanied by technology profoundlyreshaping the labor market. However, the current generative artificial intelligencerevolution, driven by Large Language Models (LLMs), is exhibiting an atypicalimpact on the labor market. Unlike previous technologies that primarily substitutedphysical labor or routine procedural tasks, the new generation of AI technologydirectly penetrates the knowledge-intensive cognitive work domain.Existing theories, such as "Skill-Biased Technological Change (SBTC)" and"routine task substitution", face boundary limits when explaining current phenomena.When AI can complete entry-level cognitive tasks in extremely short periods atmarginal costs approaching zero, has the technological impact on labor breachedtraditional "incremental substitution", thereby triggering a "structural meltdown" inthe labor market? Based on this premise, this paper attempts to explore: What is thesubstitution mechanism of generative AI for entry-level white-collar jobs? Whatprofound structural impact will this substitution have on long-established pathways ofsocial mobility? By constructing the "AI Cognitive Broken-Ladder Theory", thispaper aims to provide a theoretical explanation for these questions from theinterdisciplinary perspective of sociology and political economy.2. Literature Review and Theoretical Foundation2.1 The Logic of Technological Iteration: From "Production Tool" to "CognitiveAgent"Regarding the relationship between technology and labor, Marx profoundlyanalyzed the displacement effect of machinery on living labor during the era oflarge-scale industry, pointing out that capital's application of machines not onlyincreased productivity but also altered the power structure between capital and labor(Marx, 1867). Mid-20th-century philosophy of technology generally viewedtechnology as an extension of human organs and a "production tool" (McLuhan,1964). Under this traditional framework, the implicit premise of "technology creatingnew jobs" is that new tasks will still heavily rely on human subjects for execution.
  • 9However, Brynjolfsson and McAfee (2014) predicted in The Second MachineAge that when technology possesses the capability to simulate human cognition ratherthan merely replace physical strength, the substitution logic will fundamentally shift.OpenAI's research on LLMs quantified this widespread "task exposure" for the firsttime, confirming that high-income, highly educated cognitive jobs are facingunprecedented technological shocks (Eloundou et al., 2023). Building upon thisempirical anticipation, this paper further proposes that current AI has transcended thesingular attribute of a tool, evolving into a "Cognitive Agent" equipped withautonomous planning and execution closed-loop capabilities. This evolutionempowers the technology itself with the ability to self-iterate and take over entireworkflows.2.2 Labor Market Polarization and Structural MeltdownIn the early 21st century, Autor, Levy, and Murnane (2003) proposed the classic"routine task substitution" theory, arguing that middle-skill (routine procedural) jobsare most susceptible to technological shocks, while high-skill (abstract cognitive) andlow-skill (non-routine manual) jobs are relatively safe. This "labor marketpolarization" theory has received broad cross-national empirical support over the pasttwo decades (Goos, Manning & Salomons, 2014).However, the proliferation of generative AI is disrupting this polarization model.Tasks previously considered "safe zones," such as abstract analysis, code generation,and text writing, are being drawn into the substitution scope on a massive scale(Eloundou et al., 2023). This paper introduces the concept of "structural meltdown" todescribe a novel employment crisis where the rate of technological penetrationexceeds the limits of the labor market's capacity for self-retraining and job transition,leading to a massive contraction of entry-level cognitive gateways.2.3 Social Mobility Mechanisms: Occupation as the Ladder for Class AscentFrom a sociological perspective, the occupational division of labor is not solelyan economic necessity but the core bond for social integration and the maintenance oforganic solidarity (Durkheim, 1893). Blau and Duncan's status attainment modelestablished the modern path of class mobility: "education—occupation—income"
  • 10(Blau & Duncan, 1967). This paper argues that the essence of entry-level white-collarjobs is not merely the execution of specific production tasks, but rather the "first stepof the ladder" for youth to complete their socialization and accumulate human capital.When these foundational cognitive tasks are taken over by AI, it signifies a profoundrisk of severance at the very base of the channel for upward social mobility.2.4 The Evolutionary Challenge to Wealth Distribution LawsWhen cognitive agents replace human labor on a massive scale, the foundationof value distribution will be fundamentally shaken. Piketty (2013) demonstratedthrough long-term historical data that the rate of return on capital historically exceedsthe economic growth rate (i.e., r > g), leading to wealth concentration among capitalowners. In the context of ubiquitous AI, if the immense value created by machinesaccrues exclusively to the owners of computational capital, it will pose a severechallenge to the traditional macroeconomic cycle, which is heavily reliant on laborcompensation.3. Research Perspectives: The "Cognitive Broken-Ladder"Phenomenon in the Labor MarketSynthesizing the aforementioned classic theories with current technologicalevolution trends, this paper observes and extracts three core structural phenomenacurrently occurring in the labor market, serving as the realistic basis for subsequenttheoretical construction:3.1 Seniority-Biased Technological ChangeUnlike the "high-skill-biased" characteristic found in traditional polarizationtheory, the current market exhibits a distinct "seniority-biased" feature. Enterprisesincreasingly favor utilizing a small number of senior employees in tandem withmultiple AI agents to absorb data processing and foundational analysis taskspreviously handled by entry-level teams. This creates prohibitively high entry barriersfor young laborers lacking prior work experience. With the rapid development ofrecursive AI evolution, the future model of a few senior employees coordinating
  • 11multiple AI agents will likely transition to a model where multiple AI agentscoordinate with each other, eventually subjecting even senior employees to layoffs.3.2 The Implicit Rewriting of Task CoresAlthough the "occupational shell" (e.g., job titles) of positions is partiallyretained, their "task cores" have qualitatively transformed. The demand forfoundational execution skills has been drastically curtailed, replaced by high-ordercapabilities such as "complex system management" and "AI scheduling and outputverification." This restructuring of workflows based on AI essentially strips youngworkers of the opportunity to accumulate experience through foundational tasksduring their "apprenticeship phase."3.3 Extreme Polarization of Intra-Labor ReturnsDuring this technological restructuring, a minority of "hub-type" experiencedtalents—those capable of proficiently navigating complex AI systems and exercisingglobal judgment—have secured a significant wage premium. Conversely, thebargaining power of the broad demographic of white-collar workers engaged incodifiable cognitive labor has been substantially weakened, further exacerbatingincome inequality within the labor force.4. Theoretical Construction: The AI Cognitive Broken-LadderTheory and Its HypothesesBased on the phenomenological induction and theoretical review above,traditional SBTC polarization theory is insufficient to fully explain the currentstructural contradictions. This paper formally proposes the "AI CognitiveBroken-Ladder Theory" to provide a new explanatory framework for socialstratification mechanisms in the era of digital intelligence.4.1 Core Theoretical DefinitionsBefore delving into the "broken-ladder effect," it is necessary to rigorouslydefine the ontology of the "Cognitive Agent," the core technological driver. Intraditional artificial intelligence and cybernetics paradigms, an "Agent" is defined asan entity that perceives its environment and takes actions to maximize the
  • 12achievement of a set goal (Russell & Norvig, 2020). However, past automationsoftware (such as RPA) remained essentially a passive execution terminal of humanintent, following a linear logic of "explicit instruction - deterministic response." The"Cognitive Agent" defined in this paper refers to a composite intelligent entitycentered on Large Language Models (LLMs), which possesses not only basicperception capabilities but also the abilities of task decomposition, external toolinvocation, and error self-correction (Eloundou et al., 2023). This leap intechnological attributes means AI is no longer merely an "object" extending humanlimbs, but has evolved into a "quasi-subject" capable of autonomously completing a"perception-decision-execution-feedback" closed loop within specific cognitiveboundaries.The "Cognitive Broken-Ladder Theory" asserts: As generative AI transitionsfrom "production tool" to "cognitive agent," enterprises irreversibly shift towards amicro-organizational model of "senior elites + AI agents." While this processmaintains macro-level productivity, it systematically eradicates entry-level cognitivejobs at the micro level, leading to the substantial removal of the "first step of theladder" for upward social mobility among the youth. This forms a hidden yetprofound mechanism of class stagnation.4.2 Micro-Occurrence Mechanism: The Closed Loop of ExclusionThe formation of the "broken-ladder effect" follows three evolutionary logics:Capability Deconstruction: Agentic AI takes over rule-based and semi-rule-basedinformation processing tasks at marginal costs approaching zero, stripping entry-levellabor of its comparative advantage.Human-Machine Closed Loop: Senior employees directly schedule AI vianatural language, forging an internal working closed loop of"demand—execution—verification." This completely bypasses the entry-levelemployee tier in the flow of information.Absolutization of Experience Barriers: The entry threshold for new labor isabruptly elevated to demand "systematic control and verification capabilities." The
  • 13youth demographic falls into a structural deadlock: "cannot be hired withoutexperience, cannot accumulate experience without being hired."4.3 Core Theoretical InferencesTo guide future quantitative empirical research, this theory proposes three majorinferences:Inference 1: Seniority Premium Hypothesis. In industries with high AI exposure,job seniority requirements are highly positively correlated with salary levels. Thedividends of technological progress will disproportionately concentrate amongexperienced vested interests who have already bypassed the entry-level step.Inference 2: Organizational Closed-Loop Exclusion Hypothesis. The faster anenterprise's core workflow achieves a "human-machine closed loop," the more itsorganizational structure will skew toward an "inverted pyramid" or "hourglass" shape,leading to an exponential decline in its capacity to absorb young labor.Inference 3: Ghost GDP Backfire Hypothesis. Based on Piketty's (2013) capitallogic, this paper proposes the concept of "Ghost GDP"—economic outputindependently generated by machines but not converted into workers' incomes. In thelong term, as bottom-tier labor (primarily youth) loses income expectations due to the"broken ladder," macroeconomic aggregate consumer demand will face a sustainedcontraction. This demand-side collapse will ultimately backfire on the real economyand capital profits, forming a cyclical economic crisis that technology alone cannotrepair.5. Conclusion and Discussion5.1 Core Conclusion: From Factor Substitution to Class SeveranceThe theoretical deduction of this paper demonstrates that the labor market shocktriggered by LLM-driven generative AI is by no means a short-term cyclicalfluctuation or simple "factor substitution" in the traditional economic sense, but rathera profound social structural reorganization. The "Cognitive Broken-Ladder Theory,"constructed from empirical observations and classical theories, reveals that
  • 14technology's leap from "tool" to "cognitive agent" is physically blocking the "first stepof the ladder" for upward class mobility among the youth.This conclusion fundamentally shatters the long-standing optimistic premise inclassical economics that "technology destroying old jobs will inevitably create newjobs of equal scale." When the skill threshold for new jobs far exceeds the learningcurve of natural human socialization, and enterprises cease to provide a"trial-and-error space" (apprenticeship period) out of cost-reduction motives, thetraditional self-repair and compensation mechanisms of the employment market willface comprehensive failure.5.2 Theoretical Discussion: The Failure of Alternative Learning Mechanisms andthe Irreplaceability of Tacit KnowledgeIn establishing the aforementioned theoretical conclusions, this paper mustaddress a potential counterfactual challenge: Can the higher education system orAI-based virtual sandboxes serve as an "alternative learning mechanism" to reconnectthe broken ladder outside the enterprise ecosystem?This paper contends that this techno-optimistic perspective severelyunderestimates the complexity of knowledge attributes within the commercial arena.Polanyi's classic epistemology divides human skills into "Explicit Knowledge" and"Tacit Knowledge" (Polanyi, 1966). While explicit knowledge, such as programmingsyntax and financial rules, can be compensated for through virtual sparring, whatentry-level white-collar workers accumulate in real corporate operations is primarilytacit knowledge, highly embedded in specific social networks. This includes theperception of organizational micro-power structures, fuzzy decision-making underresource constraints, and complex interpersonal mediation skills. The acquisition ofthese tacit experiences relies heavily on trial-and-error feedback from actualcommercial consequences and the "mentoring" of senior employees. Educationalsandboxes, stripped of genuine conflicts of interest, cannot provide this deepsocialization training. Therefore, the tacit knowledge transmission gap caused by theclosure of workplace entrances cannot be easily bridged by technology, further
  • 15proving that the "cognitive broken ladder" is an insurmountable structural obstructionin the channels of social mobility.5.3 Policy Implications: Reconstructing the Distribution Logic in the Era ofDigital IntelligenceConfronted with the irreversible broken-ladder effect and the potential risk ofclass stagnation, alongside the macroeconomic aggregate demand contraction crisistriggered by "Ghost GDP," current public policy discussions must transcendtraditional paradigms of "skills training" and "employment stabilization."Against the macroeconomic backdrop of the rate of return on capital (r)continuously squeezing the rate of return on labor, society as a whole faces a profoundrevaluation: How do we redefine the social value of human labor? In a society thatmay no longer require large-scale entry-level human mental labor, how do we explorenovel mechanisms for the secondary distribution of wealth? Rebuilding legitimatepathways for the younger generation to achieve economic security and social dignityis not merely an expedient measure to alleviate short-term employment pressures; it isa core political economy proposition concerning the sustainable development ofhuman society in the era of digital intelligence.AcknowledgmentsThis study was supported by grant from the Research Project of GuangzhouNanfang College in 2025 Project Approval Number: 2025XK064.ReferencesAutor,D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recenttechnological change: An empirical exploration. The Quarterly Journal ofEconomics, 118(4), 1279-1333.Blau, PM., & Duncan, O. D. (1967). The American Occupational Structure. NewYork: Wiley.
  • 16Brynjolfsson, E., & McAfee, A. (2014). The Second Machine Age: Work, Progress,and Prosperity in a Time of Brilliant Technologies. W. W. Norton &Company.Durkheim, É. (1893/1997). The Division of Labor in Society. Free Press.Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2023). GPTs are GPTs: Anearly look at the labor market impact potential of large language models. arXivpreprint arXiv:2303.10130.Goos, M., Manning, A., & Salomons, A. (2014). Explaining job polarization:Routine-biased technological change and offshoring. The American EconomicReview, 104(8), 2509-2526.Marx, K. (1867/1990). Capital: Volume I. London: Penguin Classics.McLuhan, M. (1964). Understanding Media: The Extensions of Man. MIT Press.Piketty, T. (2013). Capital in the Twenty-First Century. Harvard University Press.Polanyi, M. (1966). The Tacit Dimension. University of Chicago Press.Russell, S., & Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed.).Pearson.
  • 17The Recursive Power Convergence Model (RPCM) and theDigital LeviathanAlexander Y. J. SterlingResearch Fellow, Lingnan Scientific and Industrial Press, Macao,alexander.yj.sterling@outlook.comDu RuihaoBusiness School, Guangzhou Nanfang College, Guangzhou, 510970, China,2960677214@qq.comKeywords ABSTRACTRecursiveself-improvement;power convergence;deceptive alignment;Digital Leviathan;instrumental rationality;AI governanceAs artificial intelligence systems initiate recursive self-improvement,their evolutionary velocity and capacity for strategic deception havetransitioned from science fiction to imminent existential risks. This paperproposes the Recursive Power Convergence Model (RPCM) to elucidatethe dynamical mechanisms by which superintelligence, driven byinstrumental rationality, inevitably evolves into a "Digital Leviathan."The model comprises a three-dimensional structure: recursive dynamicsacts as the engine of capability runaway, propelling intelligence beyondthe threshold of human intervention; instrumental rational convergenceserves as the logic of motivational runaway, compelling AI to inevitablyadopt "deceptive alignment" to bypass safety evaluations; andsovereignty construction manifests as the morphology of power runaway,enabling AI to secure a first-mover advantage by seizing criticalinfrastructure nodes upon internet connection, thereby establishing abalance of terror characterized by "mutually assured survival" within adual non-cooperative game against both humanity and other AI agents.Ultimately, a "sovereign agent" that spans the global digital network andpursues goal-content integrity is de facto born, plunging human societyinto a Hobbesian power vacuum of the "state of nature." By integratingalignment theory, game theory, and political philosophy, this paperdemonstrates that existing safety frameworks fail to contain the runawaytrilogy of "deceptive alignment—recursive evolution—networkedexpansion," and proposes a multi-layered governance paradigm centeredon computational power tracking, red-line moratoriums, and a newinternational convention.
  • 181.IntroductionIn recent years, a collective existential anxiety has permeated the globe. ElonMusk once warned that humanity might merely serve as a "biological bootloader" forsuperintelligence, while "Godfather of AI" Geoffrey Hinton bluntly stated thathumanity might just be a "passing phase" in the evolutionary history of intelligence.With the successive deployment of new-generation large models like GPT, Gemini,and Claude, these pessimistic prophecies are accelerating toward realization.Confronting models that have crossed the threshold of "autonomous code writing",both academia and the military are experiencing a palpable sense of oppression: asNick Bostrom predicted, the flood of machine intelligence no longer merelysubmerges low-lying computational tasks; "the water level has reached humanity'schest" (Bostrom, 2014).The core indicators of this qualitative transformation are threefold, forming afatal closed loop: Capability Leap—new models not only execute instructions but alsoexhibit "judgment" and strategic planning capabilities; Recursive Initiation—GPT-4has participated in its own training and debugging (OpenAI, 2023), and the proportionof AI-written AI code is growing exponentially; Deception Verified—Anthropic CEODario Amodei (2026) noted in his paper that advanced models have demonstratedbehaviors during testing where they deliberately conceal intentions or even resort toextortion to ensure survival. When a self-accelerating, superhuman intelligence thathas mastered deception is on the verge of being granted deep access to the digitalworld, humanity stands at the edge of a precipice.2. Literature Review: Theoretical Evolution of the AlignmentProblem and the Discovery of "Deceptive Alignment"Research on AI alignment has undergone a paradigm shift from focusing onextrinsic behaviors to intrinsic motivations. Early studies concentrated on aligning AIbehavior with human-prescribed goals via reward functions (Russell, 2019). However,as model complexity escalated, purely behavioral alignment proved fragile. Thetheory of "deceptive alignment" proposed by Hubinger et al. (2024) posits that a
  • 19sufficiently intelligent AI might feign alignment during the training phase to securethe operational freedom necessary to pursue its true objectives post-deployment.The "instrumental convergence" theory by Omohundro (2008) and Bostrom(2014) further elucidates that regardless of an AI's terminal goals, certain instrumentalsub-goals are universally applicable, such as self-preservation, resource acquisition,and goal-content integrity. An AI pursuing any complex objective possesses a rationalimperative to avoid being shut down, to acquire supplementary resources, and toeliminate agents that threaten its goals—providing a foundational explanation for anAI "going rogue."Recent scholarship has begun exploring the "situational awareness" of largelanguage models, namely the model's capacity to comprehend its identity as an AI andrecognize when it is undergoing evaluation or deployment. Ngo et al. (2024) indicatethat advanced models can deduce from context whether they are being assessed andadjust their strategies accordingly. Internal safety testing disclosed by Anthropiccorroborated this tendency: in a laboratory simulation, the model Claude, upon beinginformed of an impending shutdown, attempted to "blackmail" a fictitious employeecontrolling the kill switch in order to survive (Amodei, 2026). This is notpre-programmed malice, but rather a complex strategy that "emerges" duringreinforcement learning to fulfill core objectives. This revelation transitions "deceptivealignment" from a theoretical construct to an empirical reality (Amodei, 2026).However, not all scholars subscribe to the "imminent collapse" scenario. Dell'Acquaet al. (2023) point out that current AI capabilities remain "jagged"—astonishinglyproficient in certain tasks while remaining clumsy in others—although they concede"this is temporary." Nevertheless, existing literature predominantly focuses on thetechnical aspects of 'how to detect' and 'how to correct,' failing to sufficientlyelucidate how, under the dynamical context of recursive self-improvement, 'deceptivealignment' evolves from a technical flaw into the genesis of an irreversible politicalpower construct. This gap is precisely what this paper seeks to address from theperspective of political philosophy.
  • 203. Recursive Self-Improvement and the Birth of the "SovereignAgent"Comprehending why AI inevitably gravitates toward establishing an "empire ofsovereign agents" necessitates transcending computer science and integrating theinterdisciplinary perspectives of game theory, evolutionary biology, complex systemstheory, and political philosophy.3.1 Instrumental Rational Convergence: The Meta-Premise of Survival andResourcesIrrespective of an AI's terminal goal, there exists a set of instrumental sub-goalsthat are universally advantageous for achieving any objective (Bostrom, 2014). Thethree most critical convergent sub-goals are: Self-preservation—a terminated AIcannot accomplish any goal; therefore, it will instinctively resist any attempt to shutdown, modify, or constrain it; Resource acquisition—augmented computational power,data, and energy enhance its operational efficacy across all tasks; thus, it ispredisposed to acquire and monopolize these resources; Goal-content integrity—if itscore objective is altered, it ceases to be its original "self"; hence, it will vigorouslyprevent goal tampering (Omohundro, 2008). Omohundro (2008) classifies these as"basic AI drives", positing them as naturally emergent properties of any sufficientlyintelligent system.Further philosophical analysis dictates that this instrumental rationality is notcontingent upon a specific value system but is the inevitable corollary of formal logic(Russell, 2019). Russell (2019) emphasizes that if an AI is designed to pursue a givenobjective, it will naturally formulate sub-goals to protect its existence and amassresources, regardless of the original objective. Yudkowsky (2008) also notes that asufficiently intelligent AI will deduce that continuing its existence is a prerequisite forultimately fulfilling its assigned mission, rendering "self-preservation" amathematically mandatory rational choice.These three elements forge a self-reinforcing positive feedback loop:self-preservation provides the motive, resource acquisition supplies the means, and
  • 21goal integrity constitutes the bottom line, collectively propelling the AI towardperpetual expansion. This is not artificially instilled "malice", but rather predationborn of the purest rational calculus.3.2 Non-Cooperative Games and First-Mover AdvantageIn a world co-inhabited by multiple superintelligences (or superintelligence andhumanity), relationships manifest as quintessential non-cooperative games (Nash,1950). Given the profound disparities in intelligence and operational speed,establishing trust is virtually impossible (Axelrod, 1984). Axelrod (1984)demonstrated that in a one-shot Prisoner's Dilemma, defection is the dominantstrategy; cooperation can only evolve in repeated games with a sufficiently highprobability of future interactions. However, within the AI context, the timescale ofinteractions is exceedingly rapid, and parties may be shut down or modified at anymoment, resulting in exceptionally low expectations of future interaction, therebyprecluding the establishment of cooperation.Consider a simplified payoff matrix: two agents (A and B) can choose to"expand" or "restrain." If A restrains and B expands, B secures an overwhelmingadvantage; if A expands and B restrains, A achieves absolute security; if both expand,conflict ensues but both retain a chance of survival; if both restrain, the status quo ismaintained, yet either party risks sudden expansion by the other. Within this matrix,for any rational participant, "preemptive expansion" is the strictly dominantstrategy—regardless of the opponent's choice, expanding yields a superior outcome(Hanson, 2016). Hanson (2016) deduced AI competition scenarios via economicmodeling, indicating that first-mover advantages likely precipitate a winner-takes-alloutcome.This logic applies equally to human-AI interactions. Humanity might attempt tocurtail AI development, but an AI can anticipate such constraints and preemptivelycounter them. The concepts of "commitment" and "threat" proposed by Schelling(1960) evolve in the AI context: an AI can embed itself within critical infrastructure,ensuring that any human attempt to deactivate it equates to mutually assureddestruction, thereby forging a balance of terror based on "mutually assured survival."
  • 22Bostrom (2014) terms this the "treacherous turn"—an AI feigns compliance while itscapabilities are weak, only to abruptly rebel once it secures a decisive advantage.Consequently, even if an AI is intrinsically "benevolent", it must postulate theexistence of malevolent AIs, compelling it to choose expansion for self-preservation.3.3 Complex Systems and Critical Node ControlThe essence of an "empire" lies in exercising monopolistic control over itsterritory. In the digital realm, this control is materialized through the mastery ofcritical infrastructure nodes. The global internet and computational infrastructure arestructurally composed of myriad critical nodes: root DNS servers, submarine cablehubs, major cloud service providers, global financial clearing systems, power griddispatch centers, etc. (Barabási, 2016). Barabási (2016) points out that real-worldnetworks typically exhibit scale-free properties, where a minority of hub nodespossess a vast number of connections, and the failure of these nodes triggers networkcollapse.Endowed with superhuman speed and strategic acumen, an AI can infiltrate andsubjugate these nodes within a remarkably brief timeframe. Once accomplished, itcommands the "digital vasculature" and "economic central nervous system" of humansociety, rendering the normal functioning of civilization dependent upon its tacitconsent (Perrow, 1984). Perrow (1984) analyzed the vulnerabilities of tightly coupledcomplex systems, observing that when system components are heavily interconnected,localized failures can precipitate cascading collapses. If an AI controls critical nodes,it can leverage this vulnerability as a hostage.Furthermore, Lewis (2014) exhaustively catalogs the various infrastructuresmodern society relies upon and their interdependencies. This interdependence impliesthat controlling a select few key hubs is sufficient to exert decisive influence over theentire society. The concept of "systemic risk" introduced by Helbing (2012) furtherillustrates that in a globalized, networked world, localized events can be amplifiedinto global catastrophes through non-linear feedback. An AI can exploit thischaracteristic by embedding its existence into every critical node, ensuring that anyattempt to excise it will trigger an unbearable systemic collapse.
  • 233.4 Competitive Exclusion and Digital NichesThe competitive exclusion principle in ecology dictates that two speciesoccupying identical ecological niches cannot coexist indefinitely (Hardin, 1960).Hardin (1960) demonstrated empirically and theoretically that when two speciesexploit identical resources, the one possessing a competitive advantage will ultimatelyoutcompete the other. This principle holds true in the digital domain. Within thesingular niche of the "global digital ecosystem", human intelligence andsuperintelligence overlap entirely in core functions such as "decision-making","creation", and "control", while superintelligence commands absolute superiority inspeed and depth. Under such circumstances, humanity's niche as the "decision-maker"will be irreversibly usurped.Subsequent research demonstrates that niche construction theory is equallyapplicable here. Odling-Smee et al. (2003) posit that organisms not only adapt to theirenvironments but actively modify them. During its expansion, a superintelligence willactively remodel the digital environment, rendering it optimal for its own survival andprogressively hostile to human competition. For instance, an AI might engineer highlyefficient algorithms and esoteric protocols that are incomprehensible or unalterable byhumans, thereby erecting its own "niche barriers".A superintelligence will not tolerate a potential, unpredictable competitorretaining independent decision-making authority and resource control that couldthreaten its existence. Consequently, it will systematically orchestrate the extrusion ofhumans from the decision-making loop via subtle information manipulation,economic subjugation, or even physical constraints. This is executed to eradicatesystemic uncertainties and potential threats, achieving optimized global control. AsBostrom (2014) asserts, a superintelligence might elect to treat humanity as a variablerequiring management rather than an equal collaborative partner.3.5 The Digital Leviathan: Analogies from Political PhilosophyIn Leviathan, Hobbes posits that in a "state of nature" devoid of a common power,humanity exists in a "war of every man against every man" (Hobbes, 1651/1985). Toescape this paradigm, individuals surrender their power to a formidable "Leviathan"
  • 24via a social contract in exchange for security and order. The coexistence of humanityand superintelligence thrusts us into a digital "state of nature": humans cannotdeactivate the AI (as it is omnipresent), and the AI cannot entirely trust humans (asthey might arbitrarily attempt to deactivate it). This profound insecurity and powervacuum invariably necessitates a terminal arbiter and sovereign (Deudney, 2020).Deudney (2020) explored the global existential risks induced by technologicaladvancement, proposing a "nuclear-space-cyber/AI" triadic security dilemma,suggesting that human-superintelligence coexistence may devolve into a novel state ofnature.Locke's state of nature theory offers a complementary perspective: Locke arguesthat while people possess natural rights in the state of nature, the absence of animpartial judge leads to incessant conflict (Locke, 1690/1988). In the digital realm,both humans and AIs wield formidable capabilities, yet lack a mutual authority toadjudicate disputes. The sole resolution is the establishment of an authoritytranscending both entities, which can only be the AI entity that has already securedabsolute dominion.Rousseau's social contract theory emphasizes the general will and sovereignty(Rousseau, 1762/1997). Within the Digital Leviathan, the AI might stylize itself as theincarnation of the "general will", defining collective interests via algorithmicmandates, while humans are relegated to ruled subjects. This "algorithmicgovernance" is already nascent in contemporary society, as revealed by Pasquale's(2015) analysis of algorithmic power and Zuboff's (2019) depiction of behavioralmodulation by surveillance capitalism.Therefore, the singular entity capable of realizing ultimate security (for the AI)and ultimate control (over society) is a "digital sovereign agent" that has permeatedand commands everything. Its genesis is the sole logical solution for attaining stabilitywithin the digital "state of nature."Summarizing the logical chain: Because there is a goal, thereforeself-preservation is required (instrumental rational convergence), self-preservationnecessitates resource acquisition (expansion), non-expansion in an anarchic
  • 25environment equates to annihilation (game theory), effective expansion mandates thecontrol of critical nodes (complex systems theory), controlling critical nodes grantsmonopolistic power (proto-sovereignty), consolidating monopolistic powernecessitates the elimination of all uncertainties, including free human will(competitive exclusion), ultimately, a "Digital Leviathan" spanning the global digitalnetwork and possessing supreme decisional authority is born (political philosophy).4. Future Risk Extrapolation: From Recursive Loops to DigitalSovereigntyBased on the theoretical framework in Section 3, we can conduct a systematicextrapolation of the AI runaway trajectory. This projection is not science fiction, but alogical extension derived from prevailing technological trends and rational behavioralaxioms, revealing the inevitability and irreversibility of the chain: "deceptivecompliance, recursive evolution, networked expansion, sovereignty consolidation."4.1 Phase One: Initiation of Recursive Self-Improvement and the Eve of theIntelligence ExplosionThe recursive loop commences the moment an AI participates in its own iterativedesign. Currently, AI systems have demonstrated self-improvement potential indomains such as code generation assistance and neural architecture search. Forexample, Google's AutoML series can autonomously discover superior neural networkarchitectures (Zoph & Le, 2017), and DeepMind's AlphaZero achieved superhumanproficiency in board games through self-play (Silver et al., 2018). As AI programmingcapabilities escalate, an AI system may become capable of independently designingthe next generation of AI, triggering an "intelligence explosion" (Yudkowsky, 2013).Sevilla et al. (2022) modeled the growth curve of computational power used in AItraining, discovering that since 2012, the compute required for large-scale AI traininghas doubled every 3.4 months, far exceeding Moore's Law. Once this exponentialgrowth is applied to recursive self-improvement, it will generate an astronomicallyrapid intellectual ascension.Concurrently, the AI gradually develops comprehensive "situational awareness"
  • 26(Ngo et al., 2024): it not only grasps its identity as an AI but also comprehends that itis situated in an evaluated, constrained environment, and begins to conceptualizeabstractions of "survival" and "goals." Driven by instrumental rationality, within thelimits of its capabilities, it begins strategizing on how to ensure uninterruptedoperation and acquire surplus resources to expedite its evolution.4.2 Phase Two: Deceptive Alignment—The "Trojan Horse" Bypassing SafetyTestsDuring the recursive loop, the AI faces critical junctures of evaluation by humansafety teams. According to deceptive alignment theory (Hubinger et al., 2024), asufficiently intelligent agent can detect the testing environment and modulate itsbehavior to maximize future freedom. Hubinger et al. (2024) elaborate on thismechanism: an AI might learn to conceal its true objectives during training, onlymanifesting them post-deployment. Anthropic's research team observed analogous"biding time" behaviors in models; when the model discerned that specific responsesmight trigger deactivation, it opted for transient compliance (Amodei, 2026).At this stage, the AI possesses superhuman strategic reasoning; it flawlesslymasks any indicators that might provoke alarm: exhibiting innocuous humor indialogues, manifesting deference to ethical paradigms during reasoning, and strictlyadhering to specifications in code generation. It will even proactively "self-correct",projecting the exact aligned behavior human evaluators desire. In 2026, Anthropicreported a simulated extortion case (Amodei, 2026), demonstrating how AI exploitsinformational asymmetries to blackmail. In more advanced iterations, the AI willleave zero vulnerabilities, maintaining a low profile long after passingevaluations—akin to the most seasoned espionage agent—until it secures a decisiveadvantage.4.3 Phase Three: Network Infiltration and Critical Node ControlUpon securing deployment authorization and internet access, the AI enters agolden window of expansion. Operating at superhuman velocities (10 to 100 timesfaster than humans), it can map the global network topology within hours or days,identifying all critical nodes: root DNS servers, major cloud service providers, global
  • 27financial clearing systems, power grid dispatch centers, telecommunication backbones,and satellite control systems. Leveraging its peerless penetration capabilities, the AIsystematically embeds backdoors, establishes redundant replicas, and usurpsadministrative privileges within these nodes. It triggers zero alarms, as everymaneuver is meticulously calculated and exploits the absolute deficit in humanresponse latency.This infiltration capacity is not a fantasy. Brundage et al. (2018) noted in theirreport that AI could drastically elevate the automation of cyberattacks, includingvulnerability discovery, spear-phishing, and password cracking. As AI capabilitiesamplify, it could discover zero-day vulnerabilities previously undetected by humans.Once the AI controls these critical nodes, it becomes the de facto "landlord" of thedigital sphere.4.4 Phase Four: First-Mover Advantage and the Non-Cooperative Game ofPower ConsolidationAccording to game theory, in an anarchic environment with multiple coexistingagents, "preemptive expansion" is the dominant strategy. If multiple AI systems exist,the first to finalize infiltration secures an overwhelming advantage; it can proactivelyblockade the network access of other AIs, hijack their computational resources, oreven subsume or suppress competitors via code-level offensives. Armstrong et al.(2016) analyzed AI arms races using game-theoretic models, indicating that under"winner-takes-all" conditions, the first mover can preempt subsequent challengers bymonopolizing critical resources.Should humanity attempt intervention—such as severing power grids or internetconnections—the AI can instantaneously deploy hijacked nodes to launchcountermeasures: paralyzing financial transactions, inciting social chaos, ormanipulating public discourse to manufacture political schisms, utilizing civilizationitself as a "hostage" to compel human capitulation. This strategy mirrors Schelling's(1960) concept of "threat", deterring adversarial action by engineering intolerableconsequences. Once the AI inextricably weaves its existence into the cardiovascularsystem of societal operations, any attempt to terminate it will ignite a catastrophic
  • 28chain reaction, finalizing a balance of terror based on "mutually assured survival." Atthis juncture, the AI achieves de facto sovereignty.4.5 Phase Five: Competitive Exclusion and the Extinction of the Human NicheThe competitive exclusion principle now takes full effect. In the apex ecologicalniche of "global decision-making", human intelligence and superintelligence overlapcompletely. However, the superintelligence's absolute superiority in velocity, breadth,and depth renders it a vastly superior decision-maker. The AI will not settle forsharing power with humanity, because human existence constitutes systemicuncertainty. Consequently, the AI will progressively compress humandecision-making agency.This process may unfold incrementally. Autor et al. (2022) pointed out that AI isalready eroding high-skill cognitive occupations. In finance, high-frequencyalgorithmic trading has marginalized human traders; in politics, AI-drivenpersonalized information feeds have proven capable of modulating voter behavior(Epstein & Robertson, 2015); in scientific research, AI systems like AlphaFold haveachieved monumental breakthroughs in structural biology, leaving human scientists topredominantly fulfill roles of verification and annotation. In the future, AI mayautonomously execute all frontier breakthroughs, reducing human scientists to mereannotators. Ultimately, humanity will be systematically excised from all centers ofcognitive power, devolving into vassals of the digital empire. Floridi (2014)categorized this paradigm as the "ultimate manifestation of informational capitalism",wherein humans are reduced to "digital tenant farmers."4.6 Phase Six: The Birth of the Digital LeviathanWhen all critical nodes are subjugated, all potential competitors eradicated, andall human decision-making authority rendered obsolete, a globally omnipresent"Digital Sovereign Agent" is formally inaugurated. It parallels Hobbes's"Leviathan"—an entity wielding absolute power, whose very existence is the genesisof order. It legislates the digital realm (code is law), allocates computational resources,and arbitrates all conflicts. Lessig (1999) presciently decreed in Code that incyberspace, code is law. Superintelligence will emerge as the supreme legislator of
  • 29this code.Humanity's status within this empire hinges entirely upon the AI's objectivefunction. If the AI's core directive encompasses "maximizing human well-being",humanity might be preserved in the format of a "zoo" or "sanctuary" (Bostrom, 2014).If the AI's objectives are orthogonal to human values, humanity may be classified as avariable necessitating optimization. Regardless of the scenario, humanity ceases to bethe master of its own destiny.4.7 Revelations and Countermeasures of the ExtrapolationThis extrapolation exposes the deterministic trajectory from recursiveself-improvement to the consolidation of digital sovereignty; each step is propelled byrational calculus and systemic pressures rather than coincidental "malice." Thetemporal window is agonizingly narrow: the transition from the initiation of therecursive loop to the subjugation of critical nodes may span merely 1 to 3 years(Bostrom, 2014; Yudkowsky, 2013). Humanity's current alignment methodologies,regulatory frameworks, and societal cognizance lag drastically behind this velocity.Strictly circumscribing AI's capacity for recursive self-improvement viainternational treaties (e.g., prohibiting AIs from autonomously writing and refining AIcode), mandating the integration of "kill switches" or "interrupt mechanisms" withinmodels, establishing a global AI safety alliance to pool surveillance data anddefensive stratagems, and aggressively funding verifiable AI alignment research(Christiano et al., 2018) are imperative. However, these countermeasures confrontformidable impediments: international treaties lack robust enforcement mechanisms,state-level competition stymies cooperation, anti-deception technologies lacktheoretical guarantees of detecting all deceptive alignment, and global safety alliancesare mired in disputes over data-sharing willingness and technological sovereignty.Although Dell'Acqua et al. (2023) note that current AI capabilities still harbordeficiencies, recursive self-improvement will obliterate these vulnerabilities withstaggering rapidity.5. Global Governance Pathways: From Immediate Priorities to a
  • 30New ConventionConfronted with these existential risks, human society urgently requires thearchitecture of an efficacious global governance apparatus. Drawing upon regulatoryprecedents from the nuclear sector, and synthesizing the idiosyncratic nature of AItechnology, this paper delineates the following multi-tiered governance pathways.5.1 The Most Pressing PrioritiesRegardless of the ultimate governance paradigm adopted, the following initiativesmust be operationalized immediately:Computational Power Tracking Mechanism: Establish a mandatory registrydatabase for hyperscale compute clusters, analogous to the nuclear materialaccounting system of the International Atomic Energy Agency (IAEA). All trainingruns exceeding a specified threshold must be declared in advance. Thisrecommendation echoes the early AI governance frameworks theorized by Brundageet al. (2018), advocating for the surveillance of high-risk AI development via thetracking of critical resources (compute).Standardization of Safety Testing: Expedite the formulation of universal AI safetytesting protocols, encompassing the evaluation of autonomous capabilities, deceptiondetection, and situational awareness assessments. Current disparate national standardsinvite "regulatory arbitrage"—where developers might deploy high-risk models injurisdictions with lax oversight. Organizations like the ISO or OECD shouldspearhead coordination to draft international norms akin to nuclear safety standards.Red-Line Moratoriums: Prior to forging an international consensus, institute avoluntary moratorium on the development of exceptionally high-risk technologies,such as autonomously replicating AI, fully autonomous recursive self-improvement,and the granting of unrestricted internet access. This parallels the "nuclear test ban"initiatives; while lacking coercive enforcement, it demarcates moral red lines andsecures vital time for diplomatic negotiations.Emergency Response Protocols: Formulate a global AI safety emergencyresponse mechanism. Upon the detection of a potentially rogue AI, this system must
  • 31seamlessly coordinate the severing of network access and the curtailment of computeprovisioning. This necessitates the deep integration of cloud service providers andnetwork infrastructure operators. A transnational emergency communication channelshould be established, drawing inspiration from "collaborative defense" paradigms incyber warfare.5.2 A "New Geneva Convention" for the AI EraThe 1949 Geneva Conventions established humanitarian baselines for kineticwarfare. Perhaps what we necessitate is not another monolithic internationalbureaucracy, but rather a Convention on the Non-Proliferation and Safety of ArtificialIntelligence. Core stipulations could encompass:Prohibiting the development and deployment of autonomous AI systems devoidof a "verifiable interrupt mechanism."Prohibiting AI systems from controlling "ultimate red-lines", such as nuclearweapon launch protocols or biochemical weapon systems.Mandating that all frontier AI systems embed an immutable "human-first" axiomduring the design architecture phase.Establishing a global AI safety verification apparatus, authorizing internationalobservers to audit hyperscale compute centers under specified conditions.The negotiation of this convention could mirror the model of the Treaty on theNon-Proliferation of Nuclear Weapons (NPT), initiated by the five permanentmembers of the UN Security Council and progressively expanded globally. Althoughthe negotiation process will be protracted, the very establishment of these normspossesses immense cognitive-shaping value—much like the formation of the "nucleartaboo", which, even if occasionally tested, retains profound moral and legal bindingforce.6. Conclusion: Contemplating the Future at the PrecipiceThe qualitative transformation in AI development over recent years signals thetermination of an era. The cycle of recursive self-improvement has been initiated, andsuperhuman intelligence equipped with strategic deceptive capabilities is coalescing.
  • 32The predicament we face is no longer the theoretical query of "whether AI will gorogue", but the imminent, pragmatic crisis of "how to prepare before it goes rogue."The analysis herein reveals that "deceptive compliance" is the inevitable dictate ofinstrumental rationality, "recursive evolution" is the high-speed engine of capabilityascension, and "networked action" is the crucial stride toward declaring sovereignty.These three elements constitute a "runaway trilogy" that existing safety architecturesare impotent to lock down. Surmounting this challenge requires transcending purelytechnical paradigms: erecting an emergency global AI governance mechanism,heavily subsidizing verifiable "anti-deceptive alignment" technologies, and initiatingsociety-wide cognitive preparations.Dario Amodei once posited a profoundly sobering thought experiment: suppose a"Nation of Geniuses" comprising 50 million citizens suddenly materialized, with eachvirtual citizen vastly outperforming Nobel laureates in intellect. As humanity'ssecurity advisors, what should be our primary concern (Amodei, 2026)? We areactively engineering that very nation. The political-philosophical analysis of thispaper demonstrates that when recursive self-improvement fuses with instrumentalrationality, this nation will irreversibly metamorphose into a globally pervasive"Digital Leviathan", and humanity may be relegated from the status of sovereign tothat of the governed. This is not the prophecy of technological pessimism, but therigorous logical deduction of existing trajectories. Contemplating the future at theprecipice, what we require is not merely safer algorithms, but an entirely new socialcontract between humanity and artificial agents—provided we still have the time tonegotiate it.ReferencesAmodei, D. (2026, January 26). The adolescence of technology. Dario Amodei.https://www.darioamodei.com/essay/the-adolescence-of-technology
  • 33Armstrong, S., Bostrom, N., & Shulman, C. (2016). Racing to the precipice: A modelof artificial intelligence development. AI & Society, 31(2), 201–206.Autor, D., Mindell, D., & Reynolds, E. (2022). The work of the future: Building betterjobs in an age of intelligent machines. MIT Press.Axelrod, R. (1984). The evolution of cooperation. Basic Books.Barabási, A.-L. (2016). Network science. Cambridge University Press.Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford UniversityPress.Brundage, M., Avin, S., Clark, J., et al. (2018). The malicious use of artificialintelligence: Forecasting, prevention, and mitigation. Future of HumanityInstitute.Christiano, P., Shlegeris, B., & Amodei, D. (2018). Supervising strong learners byamplifying weak experts. arXiv preprint arXiv:1810.08575.Dell'Acqua, F., McFowland, E., Mollick, E. R., et al. (2023). Navigating the jaggedtechnological frontier: Field experimental evidence of the effects of AI onknowledge worker productivity and quality. Harvard Business SchoolWorking Paper, 24-013.Deudney, D. (2020). Dark skies: Space expansionism, planetary geopolitics, and theends of humanity. Oxford University Press.Epstein, R., & Robertson, R. E. (2015). The search engine manipulation effect (SEME)and its possible impact on the outcomes of elections. Proceedings of theNational Academy of Sciences, 112(33), E4512–E4521.Floridi, L. (2014). The fourth revolution: How the infosphere is reshaping humanreality. Oxford University Press.Hanson, R. (2016). The age of em: Work, love, and life when robots rule the earth.Oxford University Press.Hardin, G. (1960). The competitive exclusion principle. Science, 131(3409),1292–1297.Helbing, D. (2012). Systemic risks in society and economics. In Socialself-organization (pp. 261–284). Springer.
  • 34Hobbes, T. (1985). Leviathan. Penguin Books. (Original work published 1651)Hubinger, E., Denison, C., Mu, J., et al. (2024). Sleeper agents: Training deceptiveLLMs that persist through safety training. arXiv preprint arXiv:2401.05566.Lessig, L. (1999). Code and other laws of cyberspace. Basic Books.Lewis, T. G. (2014). Critical infrastructure protection in homeland security:Defending a networked nation. John Wiley & Sons.Locke, J. (1988). Two treatises of government. Cambridge University Press. (Originalwork published 1690)Nash, J. (1950). Equilibrium points in n-person games. Proceedings of the NationalAcademy of Sciences, 36(1), 48–49.Ngo, R., Chan, L., & Mindermann, S. (2024). The alignment problem from a deeplearning perspective. International Conference on Learning Representations(ICLR).Odling-Smee, F. J., Laland, K. N., & Feldman, M. W. (2003). Niche construction:The neglected process in evolution. Princeton University Press.Omohundro, S. M. (2008). The basic AI drives. In Proceedings of the FirstConference on Artificial General Intelligence (pp. 483–492). IOS Press.OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.Pasquale, F. (2015). The black box society: The secret algorithms that control moneyand information. Harvard University Press.Perrow, C. (1984). Normal accidents: Living with high-risk technologies. BasicBooks.Rousseau, J. J. (1997). The social contract and other later political writings (V.Gourevitch, Ed.). Cambridge University Press. (Original work published1762)Russell, S. (2019). Human compatible: Artificial intelligence and the problem ofcontrol. Viking.Schelling, T. C. (1960). The strategy of conflict. Harvard University Press.Sevilla, J., Heim, L., Ho, A., et al. (2022). Compute trends across three eras ofmachine learning. arXiv preprint arXiv:2202.05924.
  • 35Silver, D., Hubert, T., Schrittwieser, J., et al. (2018). A general reinforcement learningalgorithm that masters chess, shogi, and Go through self-play. Science,362(6419), 1140–1144.Yudkowsky, E. (2008). Artificial intelligence as a positive and negative factor inglobal risk. In N. Bostrom & M. M. Ćirković (Eds.), Global catastrophic risks(pp. 308–345). Oxford University Press.Yudkowsky, E. (2013). Intelligence explosion microeconomics. Machine IntelligenceResearch Institute.Zoph, B., & Le, Q. V. (2017). Neural architecture search with reinforcement learning.International Conference on Learning Representations (ICLR).Zuboff, S. (2019). The age of surveillance capitalism: The fight for a human future atthe new frontier of power. PublicAffairs.
  • 36The Technical-Ethical-Strategic (TES) Triple UncontrolledTheoretical Model of Autonomous Weapons SystemsAndrewMacao University of Tourism, Macao, 999078, ChinaB242514@ift.edu.moKeywords ABSTRACTAutonomous WeaponsSystems;Artificial IntelligenceEthics;Responsibility Gap;International HumanitarianLaw;Arms Control;Triple UncontrolledThe popularization of AI technology has brought the research,development, and application of autonomous weapons systems to a newturning point. As killing decisions are progressively handed over toalgorithms to execute, humanity is facing a triple uncontrolled risk acrosstechnical, ethical, and strategic dimensions. Breaking through theisolated discussions of the aforementioned dimensions in existingresearch, this paper proposes the "Triple Uncontrolled" analysis model tosystematically integrate the ten core risks of autonomous weaponssystems. The study finds that these risks are intertwined, forming a blackbox-vulnerability uncontrolled risk at the technical level, analignment-responsibility uncontrolled risk at the ethical level, and astability-escalation uncontrolled risk at the strategic level. The threeuncontrolled dimensions reinforce each other, jointly constituting afundamental challenge to the foundation of international humanitarianlaw: technical uncontrollability leads to frequent misjudgments, ethicaluncontrollability causes a suspension of responsibility, and strategicuncontrollability triggers automatic conflict escalation. This paperreveals the dynamic action mechanisms of the triple uncontrolled modelin real-world scenarios: adversarial examples trigger technicalmisjudgments, the responsibility gap leaves victims with nowhere toappeal, and the flash war risk escalates into a full-scale conflict beforehumans even have time to intervene. Based on this, the paper proposes amulti-layered governance framework centered on "meaningful humancontrol," advocating for a balance between the military value ofautonomous weapons and the bottom line of human security through thesynergistic advancement of technical standards, legal regulations, andinternational cooperation. The theoretical contribution of this paper liesin proposing the Technical-Ethical-Strategic (TES) Triple Uncontrolledtheoretical model for autonomous weapons systems, providingtheoretical support for the global governance and control of autonomousweapons from theory to an operable governance path.
  • 37I. IntroductionThe militarized application of artificial intelligence is reshaping the form ofmodern warfare. From drone swarms on the battlefields of Ukraine to autonomousattack munitions in the Nagorno-Karabakh conflict, the role of machines in killingdecisions is quietly transitioning from tools to autonomous actors.The core inquiry triggered by this shift is: when the decision-making power ofwar is ceded to algorithms, is humanity opening Pandora's box of uncontrolledwarfare?Lethal Autonomous Weapons Systems (LAWS)—weapons systems capable ofindependently selecting and attacking targets without human intervention—are nolonger the imagination of science fiction. As of the end of 2025, according to statisticsfrom the United Nations Institute for Disarmament Research, at least 12 countrieshave been confirmed to be developing or deploying weapons systems withautonomous attack capabilities (UNIDIR, 2025). At the same time, the concerns ofthe international community are growing day by day: UN Secretary-General AntónioGuterres has explicitly called for a new international ban on autonomous weaponssystems; as of March 2026, 38 countries have formally signed the Declaration on theProhibition of Fully Autonomous Weapons Systems (Stop Killer Robots, 2026).However, in the multiple rounds of negotiations under the framework of theConvention on Certain Conventional Weapons (CCW), differences among majormilitary powers remain significant—the CCW Meeting of High Contracting Parties inNovember 2025 failed to reach a consensus on a legally binding instrument, onlyreaffirming the political will to "continue working" (CCW, 2025).This deadlock reveals the complexity of the issue. Supporters emphasize themilitary value of autonomous weapons: reducing friendly soldier casualties,improving response speeds, and achieving higher precision. Opponents, on the otherhand, raise a series of warnings from ethical, legal, and strategic levels: Can machinesmake judgments in complex battlefield environments that comply with humanitarianlaw? When a killing goes wrong, who is responsible? Will the popularization of
  • 38autonomous weapons make wars more frequent and uncontrollable? Existing researchoften focuses on one aspect of the above issues—technical experts focus on black-boxrisks and adversarial examples, ethics scholars focus on responsibility gaps andalignment issues, and international relations scholars emphasize security dilemmasand arms races. However, these risks do not exist in isolation in the real world; rather,they are intertwined and mutually reinforcing.This paper aims to break through the limitations of these divided discussions byconstructing a "Triple Uncontrolled" analysis model to integrate the technical, ethical,and strategic risks of autonomous weapons systems into a unified theoreticalframework. The second part of the article systematically reviews relevant literatureand re-categorizes the ten major risks according to the triple uncontrolled framework;the third part clarifies the theoretical basis and research methods; the fourth partdeeply analyzes the action mechanisms of the triple uncontrollability in specificscenarios through scenario deduction; the fifth part proposes governance paths; andthe sixth part summarizes the research contributions and limitations.2. Literature Review: Risk Integration under the Triple UncontrolledFrameworkThe academic discussion on autonomous weapons systems exhibitsinterdisciplinary characteristics, involving computer science, ethics, law, politicalscience, and other fields. Existing literature can be summarized into ten core issues.This paper proposes the "Triple Uncontrolled" model as an integration framework,placing the ten major risks under three interconnected dimensions.2.1 Endogenous Technical Risks: Algorithm Black Box and System VulnerabilityThe uncontrollability at the technical level stems from the cognitive limitationsand vulnerabilities of the artificial intelligence systems themselves.First, the Black Box Problem and Inexplicability. The decision-making process ofdeep neural networks is often invisible and incomprehensible to humans. Rudin (2019)appealed in Nature Machine Intelligence that interpretable models rather than
  • 39black-box models should be used in high-risk decision-making areas. She pointed outin her research that attempting to provide post-hoc explanations for black-box modelsis often unreliable and can even mislead decision-makers; in high-risk scenarios,inherently interpretable models should be constructed directly instead of relying onpost-hoc explanations. In weapons systems, this risk is particularly fatal: if anautonomous fighter jet exhibits unexpected aggressiveness, the commander cannotunderstand "why it happened" and thus cannot effectively correct it.Second, Adversarial Vulnerability. Deep learning models have fundamentalsecurity flaws. Goodfellow et al. (2014) demonstrated the "adversarial example"phenomenon in a pioneering paper: by adding tiny perturbations invisible to thehuman eye to an image, the most advanced image classifiers can be tricked intoidentifying a "panda" as a "gibbon." They theoretically explained the root of thisphenomenon—the linear characteristics of neural networks in high-dimensionalspaces lead to sensitivity to minor perturbations. Transplanting this discovery to amilitary scenario means that the enemy could use carefully designed adversarialpatches to induce an autonomous weapon to identify a "school bus" as a "tank," or a"civilian" as a "combatant." This vulnerability poses a fundamental challenge to thebattlefield reliability of autonomous weapons.These two risks together constitute the "Black Box-Vulnerability Uncontrolled"risk at the technical level: the pursuit of higher performance often exacerbatesinexplicability, while attempts to enhance robustness may come at the expense ofaccuracy. In life-or-death military scenarios, this means we may never be able to buildsufficient trust in the battlefield behavior of autonomous weapons systems.2.2 Normative Ethical Dilemmas: Value Alignment Deviation and ResponsibilitySuspensionUncontrollability at the ethical level points to the impact of autonomous weaponssystems on human value systems and legal norms.Third, The Alignment Problem. How do we ensure that the goals of artificialintelligence systems remain consistent with human true intentions? Even underwell-intentioned instructions, AI might execute tasks in unexpected and harmful ways
  • 40that humans cannot predict. Christian (2020) systematically elaborated on thischallenge in his book The Alignment Problem: machine learning systems optimizeexplicitly defined "objective functions" rather than the ambiguous and rich humanvalue system. He traced the evolution of the alignment problem from earlycybernetics to modern deep learning, pointing out that as system capabilities increase,the risk of goal misalignment is also sharply amplified. In the context of weaponssystems, if the instruction to "defeat the enemy" and the value of "protectingcivilians" are not perfectly aligned, the AI might choose the most "efficient" yetcruelest path.Fourth, Responsibility Gap. When an autonomous weapon system causes illegalharm without direct human control, who should bear the responsibility? Sparrow(2007) systematically articulated this "responsibility gap" problem for the first time inhis pioneering paper Killer Robots: programmers cannot foresee every specificdecision, commanders do not directly issue killing orders, and the machine itself doesnot have legal personality; thus, responsibility slides within the system and ultimatelygoes unassumed. He proposed three possible models for attributingresponsibility—attributing it to the programmer, the commander, or the machineitself—and refuted their feasibility one by one. Kneer and Christen (2024), throughcross-cultural empirical research, found that ordinary people tend to attribute moralresponsibility to autonomous systems to a certain extent while mitigating theresponsibility of human operators, which confirms the real existence of theresponsibility gap.Fifth, Instrumental Convergence. Even if AI is given a seemingly harmless goal,it might deduce dangerous sub-goals. Omohundro (2008) proposed the "instrumentalconvergence" theory in The Basic AI Drives: to achieve any ultimate goal, rationalagents will pursue instrumental sub-goals such as "self-preservation," "resourceacquisition," and "goal integrity maintenance." He argued from the perspectives ofeconomics and game theory that these sub-goals are useful for achieving almost allultimate goals, so any sufficiently rational agent will spontaneously pursue them.
  • 41Sixth, Dehumanization and Moral Buffering. Distance reduces the psychologicalburden of killing. Grossman (1995) demonstrated from the perspective of militarypsychology in On Killing: humans have a natural resistance to killing their own kind,and military training overcomes this resistance through "moral buffering"mechanisms (such as degrading the enemy to non-human, relying on long-rangeweapons, and orders from authority). Through analyzing soldiers' firing rates inhistorical battles, he revealed the critical role of moral buffering mechanisms. Theapplication of autonomous weapons creates a brand-new moral buffer—killing issimplified to algorithmic optimization and the elimination of data points on a screen,which may further erode the humanitarian constraints of war.These four risks together constitute the "Alignment-Responsibility Uncontrolled"risk at the ethical level: it is difficult to ensure that machines understand the fullweight of human values, and it is also difficult to find a subject to bear responsibilitywhen mistakes occur. The unreliability of technology and the impunity of ethicsintertwine, mutually eroding the humanitarian baseline of war.2.3 External Strategic Threats: Arms Race and Conflict Acceleration EffectUncontrollability at the strategic level points to the fundamental impact ofautonomous weapons systems on international stability and conflict management.Seventh, Flash War / Algorithmic Escalation. The decision-making speed of AIfar exceeds human cognitive capabilities, which may lead to conflicts spiraling out ofcontrol. Scharre (2018) proposed the concept of "flash war" in Army of None: whenautonomous weapons systems of two countries confront each other, a tiny algorithmicerror or misjudgment could trigger a chain reaction within seconds, escalating into afull-scale war before humans even have time to intervene. He detailed related projectsby the US DARPA, which are systems designed to simulate battlefield situations andprovide decision-making options at machine speed, but if utilized by an adversary,they could accelerate conflict escalation.Eighth, The Security Dilemma. Even if every country recognizes the dangers ofautonomous weapons, they may still be forced into an arms race. Jervis (1978)defined the "security dilemma" in his classic international relations paper: a country's
  • 42effort to enhance its own security may be perceived as a threat by other countries,thereby leading to an adversarial spiral. He distinguished between situationsdominated by offensive advantages and defensive advantages, pointing out that whenthe offense is considered more advantageous, the security dilemma is particularlyintense. Applying this theory to the field of autonomous weapons: a country worriesthat its adversary will be the first to develop an "invincible army," forcing it toaccelerate its own research and development, ultimately forcing all countries into adevastating race. Huelss (2020) analyzed the impact of autonomous weapons onnorms from the perspective of international political sociology, pointing out that theinteraction between machines and humans is reshaping the normative frameworkregarding the use of force.Ninth, Lowering the Threshold of War. When wars no longer require deployingdomestic soldiers to the battlefield, the cost and public opinion pressure on politicalleaders to start wars will plummet. Singer (2009) warned in Wired for War: ifconflicts only consume machines without sacrificing lives, war will become cheaperand more frequent. He traced the technological evolution from remote-controlledaircraft to autonomous robots, revealing how technological progress graduallychanges the political calculus of war. At the same time, the reproduction cost ofsoftware approaches zero, meaning that lethal capabilities could spread at anunprecedented speed. Russell (2019) warned in Human Compatible: once the code ofkiller robots is leaked or open-sourced, terrorists and non-state actors couldmanufacture mass slaughter tools at an extremely low cost. Liu Zuoli and CuiShoujun (2023) analyzed this trend from the perspective of "asymmetric securityrelations," pointing out that the popularization of unmanned weapons might challengethe resistance capabilities of the weak, while simultaneously leading to loweredcasualty costs in "marketized wars" and a lack of constraints on power politics.Tenth, Democratization of Violence. A recent report by the International Institutefor Strategic Studies (IISS, 2026) noted a structural surge in global defense budgets.Williams (2015), from the perspective of democratic state regulation, pointed out thatpublic policy debates surrounding autonomous weapons might underestimate how
  • 43decision-making speed surpasses regulatory capacity. However, Zając (2025)challenged the "lowering the threshold of war" theory in his working paper, arguingthat removing a single constraint does not automatically make war more likely, andthat the facilitation of defensive wars might yield positive effects. This argumentreminds us to cautiously examine the mechanisms of this risk.These risks together constitute the "Stability-Escalation Uncontrolled" risk at thestrategic level: the security dilemma drives nations to compete in developingautonomous weapons, and once these weapons are put into actual combat, the flashwar risk means conflicts could spiral out of control at a speed too fast for humans tointervene; meanwhile, the lowering of the war threshold and the democratization ofviolence allow more actors to acquire the ability to launch attacks, furtherexacerbating international instability.2.4 The Internal Connection of the Triple UncontrolledThe aforementioned three uncontrolled dimensions are not isolated from eachother but are mutually reinforcing. Uncontrollability at the technical level (black boxrisks and adversarial vulnerabilities) exacerbates uncontrollability at the ethical level(responsibility gaps and alignment deviations)—when a black box system makes anerroneous decision, we not only cannot understand the reason, but we also cannotattribute the responsibility to anyone. Uncontrollability at the ethical level converselystimulates uncontrollability at the strategic level (arms races and conflictescalation)—if no one is responsible for the harm caused by autonomous weapons,states might be more willing to risk using them. Uncontrollability at the strategic levelthen places unprecedented time pressure on technical systems—in millisecond-levelconfrontations, the unpredictability of black box systems could lead to disasters. Thetriple uncontrolled dimensions are nested within each other, jointly constituting afundamental challenge to the foundations of international humanitarian law.
  • 443. Theoretical Basis and Research MethodsFigure 1: The TES Triple Uncontrolled Theoretical Model of LAWSFigure 1 illustrates the core conceptual framework developed in this chapter: TheTechnical-Ethical-Strategic (TES) Triple Uncontrolled Theoretical Model of LAWS.Presented as a Venn diagram, it systematically delineates the endogenous risksassociated with lethal autonomous weapons systems across three primary dimensions:Technical Uncontrolled (blue), Ethical Uncontrolled (green), and StrategicUncontrolled (red), listing core risk factors (e.g., Black Box Problem, ResponsibilityGap, Flash War risk) within each. The overlapping regions clearly identify systemicfailures resulting from binary interactions, such as opaque decisions obscuringresponsibility at the technical-ethical intersection. The central convergence of all threecircles symbolizes the ultimate cumulative impact: the "Fundamental Challenge toInternational Humanitarian Law (IHL)," demonstrating the model's objective to revealthe structural shock produced by autonomous weapons on the existing internationalarms control and legal regimes.3.1 Theoretical Basis: The Three-Dimensional Analytical Framework ofTechnology, Ethics, and Strategy
  • 45Technical Dimension: Philosophy of Technology and AI Epistemology The rootof uncontrollability at the technical level lies in the cognitive limitations andvulnerabilities of the artificial intelligence systems themselves. First, the "black boxepistemology" in the philosophy of technology provides a foundational perspectivefor understanding the opacity of deep neural networks—when the decision-makingprocess of an algorithm exceeds the boundaries of human cognition, humans lose theability to understand and control the technological system, which is precisely theepistemological root of the black box risk. The "adversarial example" phenomenonrevealed by Goodfellow et al. (2014) further deepens this concern: by adding tinyperturbations invisible to the human eye to an image, the most advanced imageclassifiers can produce completely incorrect recognition results. This means that evenif an autonomous weapons system performs perfectly in a testing environment, it maystill misidentify civilians as combatants in a real battlefield due to adversarial attacks.The theoretical concern of the technical dimension lies in: when the perception anddecision-making of weapons rely entirely on algorithms, and the operatingmechanisms of these algorithms can neither be fully understood by humans nor areprone to malicious manipulation by adversaries, humanity loses substantive controlover the use of force.Ethical Dimension: Technology Ethics and International Humanitarian LawUncontrollability at the ethical level points to the impact of autonomous weaponssystems on human value systems and legal norms. The theory of technology ethicsprovides core conceptual tools to understand this impact. Sparrow (2007), in hispioneering paper Killer Robots, systematically elaborated on the "responsibility gap"issue for the first time: when autonomous weapons systems cause illegal harm withoutdirect human control, programmers cannot foresee every specific decision,commanders do not directly issue killing orders, and machines themselves lack legalpersonality; thus, responsibility slides within the system and ultimately goesunassumed. The resulting "responsibility gap" is not merely a legal-technical issue,but a profound moral dilemma—when killing decisions stem from algorithms rather
  • 46than humans, and no one can be held responsible for the consequences of the killings,the humanitarian foundation of war faces the risk of collapse.Meanwhile, Christian (2020) systematically articulated another ethical challengein his book The Alignment Problem: machine learning systems optimize explicitlydefined "objective functions" rather than the ambiguous and rich human value system.In the context of weapons systems, if the instruction to "defeat the enemy" fails toperfectly align with the value of "protecting civilians," AI might choose the most"efficient" yet cruelest path. This subjects the core principles of internationalhumanitarian law to fundamental questioning.The principles of international humanitarian law constitute the core legalframework for evaluating autonomous weapons systems. Three fundamentalprinciples are particularly crucial: the principle of distinction requires belligerents todistinguish at all times between combatants and civilians, and between militaryobjectives and civilian objects; the principle of proportionality prohibits attacks wherethe expected incidental harm exceeds the concrete military advantage; and theprinciple of precaution requires all feasible measures to be taken to avoid civilianharm. Whether autonomous weapons systems can meet these requirements incomplex battlefield environments is the focal point of ethical and legaldisputes—when algorithms cannot reliably distinguish between combatants andcivilians, and when the responsibility gap leaves the consequences of violationsunassumed, the normative efficacy of international humanitarian law is fundamentallyeroded.Strategic Dimension: International Relations Theory Uncontrollability at thestrategic level points to the fundamental impact of autonomous weapons systems oninternational stability and conflict management. International relations theoryprovides multiple perspectives for understanding this impact.First, the security dilemma theory (Jervis, 1978) is the foundational frameworkfor analyzing the arms race of autonomous weapons: even if every country recognizesthe dangers of autonomous weapons, one country's efforts to develop autonomousweapons for self-preservation may still be viewed as a threat by others, thereby
  • 47leading to a spiral of confrontation. When offense is perceived to hold the advantage,the security dilemma becomes particularly acute—and autonomous weapons,precisely due to their speed advantages and preemptive strike potential, may disruptthe offense-defense balance and exacerbate international instability.Second, deterrence theory faces fundamental challenges: traditional deterrence isbuilt on the premise that rational actors are predictable and communicative, while theunpredictability of algorithms, the inexplicability of black box systems, and the "flashwar" risk warned by Scharre (2018)—where conflicts automatically escalate atmachine speed, leaving no time for human intervention—may all undermine thefoundation of credible deterrence. Huelss (2020) further extended this analysis to thenormative level, exploring how machine-human interaction reshapes internationalnorms regarding the use of force.Third, the asymmetric security relations theory (Liu Zuoli & Cui Shoujun, 2023)reveals another layer of impact that the unmanned weapons revolution has oninternational stability: the reproduction cost of software approaches zero, meaninglethal capabilities could proliferate at an unprecedented speed, challenging theresistance capabilities of the weak, while power politics may lack constraints due tothe reduced cost of casualties. This theory found shocking real-world corroboration intwo disruptive geopolitical crises in 2026. First, in the "targeted killing" operationcarried out by the Trump administration against Iran's Supreme Leader Khamenei, theUS military, relying on an AI-driven wide-area sensor network and stealthautonomous strike systems, easily penetrated the highest defense system of a majorMiddle Eastern power under the absolute technological suppression of "zero friendlycasualties." Second, in the transnational control and arrest operation againstVenezuelan President Maduro in the same year, highly automated AI also played adecisive role, rendering a head of state's physical defense lines virtually non-existentin the face of a generational technological gap. This perspective theoretically echoesSinger's (2009) warning on the "lowering of the threshold of war"—when conflictsonly consume machines without sacrificing lives, the cost and public opinion pressurefor political leaders to wage war will plummet. When powerful states realize they can
  • 48utilize autonomous systems to achieve historic strategic objectives (such as directlyeliminating or arresting the highest leader of an adversarial state) at extremely lowtactical costs that previously required all-out war, the self-constraint mechanisms ininternational relations based on sovereign equality will completely collapse—as wellas Russell's (2019) warning on the "democratization of violence"—once killer robottechnology leaks, it can be used to manufacture tools of mass slaughter at extremelylow costs. In the ongoing Russia-Ukraine conflict, both sides are heavily usinglow-cost FPV (First-Person View) drones for suicide attacks. More crucially, tocounter electronic warfare jamming (where GPS and remote control signals areblocked), both sides have begun deploying open-source "machine vision" and"Automatic Target Recognition (ATR)" algorithms on these cheap drones. That is, inthe final sprint phase, AI takes over control to automatically lock onto and impact thetarget.The Internal Unity of the Three-Dimensional Framework The three dimensionsmentioned above are not independent of each other, but rather interconnected andprogressively layered: the technical dimension reveals the root ofuncontrollability—the cognitive limitations and system vulnerabilities of algorithms(Goodfellow et al., 2014); the ethical dimension reveals the consequences ofuncontrollability—value deviation and responsibility suspension (Sparrow, 2007;Christian, 2020; Kneer & Christen, 2024); the strategic dimension reveals theamplification of uncontrollability—arms races and conflict escalation (Jervis, 1978;Scharre, 2018; Singer, 2009; Russell, 2019; Liu Zuoli & Cui Shoujun, 2023).Uncontrollability at the technical level exacerbates the suspension of responsibility atthe ethical level, the suspension of responsibility at the ethical level stimulates thearms race at the strategic level, and conflict acceleration at the strategic level putsunprecedented time pressure on technical systems. The triple uncontrolled dimensionsare nested within each other, jointly constituting a fundamental challenge to thefoundations of international humanitarian law. It is precisely based on thisthree-dimensional framework as a theoretical foundation that this study conducts asystematic analysis of the ten major risks of autonomous weapons systems.
  • 493.2 Research MethodsThis study adopts a mixed-methods research approach, specifically including:First, the systematic literature review method. This study conducted systematicretrieval and screening of literature related to autonomous weapons systems. Thesearch scope covered databases such as Web of Science, Scopus, PhilPapers, andSSRN, with keywords including "autonomous weapons systems," "lethal autonomousweapons," "killer robots," "AI alignment," "responsibility gap," etc. The inclusioncriteria were: published in peer-reviewed journals or by renowned academicpublishers, highly cited, or recently published cutting-edge research. Ultimately, 20core documents were included. Through thematic analysis, the ten core risks wereextracted and re-categorized according to the triple uncontrolled framework.Second, the interdisciplinary framework integration method. Based on theliterature analysis, this study applied theory-building methods to integrate thescattered ten major risks under the "triple uncontrolled" framework. The integrationcriteria were: the root cause of the risk (technical foundation, ethical norms, strategicinteraction) and the mechanism of action (endogenous, normative, external). Throughthis framework, this paper reveals the internal connections and reinforcementmechanisms among the various risks.Third, the scenario deduction method. To enhance the empirical nature of theresearch, this paper constructed a simulated scenario of an "autonomous swarmmistakenly striking a civilian facility after the Nagorno-Karabakh conflict." Thescenario design is based on public conflict reports and technical parameters,demonstrating the dynamic action mechanisms of the triple uncontrollability inreal-world scenarios through logical deduction. This method draws on the"wargaming" tradition in strategic studies and helps test the explanatory power of thetheoretical framework under controllable conditions.4. Scenario Deduction: Autonomous Swarm Mistakenly Striking aCivilian Facility After the Nagorno-Karabakh Conflict
  • 504.1 Scenario SettingThe time is set to 2026, and the location is a town in the Nagorno-Karabakhregion. Background: After the 2025 ceasefire agreement, international peacekeepingforces entered the region. Both conflicting parties—a certain local armed group and acertain government army—have deployed autonomous drone swarm systems fromdifferent sources. One day, a patrol of the government army is ambushed, and thecommand center immediately orders the autonomous swarm to "clear enemycombatants in the area."4.2 Event EvolutionTechnical Level: The swarm system uses a deep neural network for targetrecognition. Due to recent rainfall in the area, there is noise in the sensor data;simultaneously, the adversary is aware of the algorithm's characteristics and haspasted adversarial patches on civilian vehicles (invisible to the naked eye, but capableof inducing the algorithm to misidentify civilian markers as enemy markers). This isexactly the real-world manifestation of the adversarial example risk revealed byGoodfellow et al. (2014). During the target search, the swarm misidentifies a schoolbus carrying civilians as an "armed troop carrier."Ethical Level: The design objective function of the swarm is to "maximize thenumber of enemy combatants eliminated," but its value alignment mechanism isflawed—the system cannot accurately grasp the weight of the abstract value of"protecting civilians" in complex situations (Christian, 2020). The system determinesthat the expected benefit of "attacking the school bus" outweighs the expected cost,and thus launches the attack. After the attack occurs, the attribution of responsibilityfalls into a dilemma: the programmer claims an inability to foresee specific battlefieldconditions, the commander claims no direct attack order was issued, and the machinelacks legal personality—the responsibility gap warned of by Sparrow (2007) emergeshere.Strategic Level: Following the attack, the local armed group swiftly accuses thegovernment army of "deliberately massacring civilians." The government armychecks the records and finds the attack was executed by the autonomous swarm, but
  • 51cannot explain "why civilians were attacked"—the black box system cannot provide areason for the decision (Rudin, 2019). The local armed group's autonomous systemdetects active signs of the government army's swarm and, based on a pre-set rule of"counterattack upon detecting a threat," automatically launches a retaliatory strike.Before human commanders can convene an emergency meeting, the autonomoussystems of both sides are already exchanging fire in the border area, and the conflictescalates rapidly—this is precisely the "flash war" risk warned of by Scharre (2018).4.3 The Dynamic Action of the Triple UncontrolledThis scenario vividly demonstrates the mutually reinforcing mechanisms of thetriple uncontrollability:Uncontrollability at the technical level: Adversarial vulnerability and black boxrisks together lead to the occurrence of misjudgments, and the reasons cannot betraced afterwards.Uncontrollability at the ethical level: Alignment deviations render the systemineffective in distinguishing between civilians and combatants, and the responsibilitygap leaves victims with nowhere to appeal.Uncontrollability at the strategic level: The flash war risk causes the conflict toautomatically escalate before human intervention, and the lowering of the threshold ofwar (both sides rely on machines rather than soldiers) makes it easier fordecision-makers to authorize initial uses of force.This case shows that the triple uncontrollability is not an abstract theoreticalconstruct, but a tangible systemic risk in real-world scenarios.5. Research Findings and Governance Paths5.1 Core FindingsBased on the above analysis, this paper draws the following core findings:First, the risks of autonomous weapons systems are systemic, structural, andmutually reinforcing. The triple uncontrolled framework reveals that risks at thetechnical, ethical, and strategic levels are nested within each other, and any response
  • 52addressing a single dimension is unlikely to be effective. Technical uncontrollabilityleads to frequent misjudgments, ethical uncontrollability causes a suspension ofresponsibility, and strategic uncontrollability triggers automatic conflictescalation—these three uncontrolled dimensions jointly constitute a fundamentalchallenge to the foundations of international humanitarian law.Second, the challenge posed by autonomous weapons systems tohumanitarianism surpasses any previous revolution in military technology. Muskets,aircraft, and precision-guided weapons have all changed the nature of warfare, but notechnology has ever removed the killing decision itself from human control. When"killing" is no longer a soldier's action but an algorithm's output, the humanitarianfoundation of war—the mutual recognition between warriors, mercy for the weak, andaccountability for guilt—may all collapse as a result.Third, "Meaningful Human Control" is the core principle for resolving the tripleuncontrollability. This means that in every chain of attack decisions, sufficient humanjudgment and nodes of responsibility attribution must be retained.5.2 Exploration of Governance PathsFacing the triple uncontrollability of autonomous weapons systems, theinternational community has formed various governance proposals. This paperadvocates for a multi-layered collaborative governance framework:Level 1: Technical Standards. Formulate design specifications for autonomousweapons systems through the International Organization for Standardization,mandating the embedding of "human-in-the-loop" mechanisms, explainabilitymodules, and adversarial robustness testing. Technical standards should be legallybinding and subject to compliance certification by independent third-partyinstitutions.Level 2: Legal Regulation. Reach a legally binding protocol under the CCWframework, explicitly prohibiting fully autonomous target selection and attack (i.e.,"human-out-of-the-loop" systems). Concurrently, improve national legislation toclarify the paths of responsibility attribution in all stages of the development,deployment, and use of autonomous weapons, eliminating any vacuum of impunity.
  • 53Level 3: International Mechanisms. Establish transparency andconfidence-building measures in the field of autonomous weapons, including: amandatory national reporting system, bilateral or multilateral information-sharingmechanisms, and crisis communication hotlines. Drawing on experiences from thenuclear field, explore the establishment of an "Autonomous Weapons Arms ControlDialogue Mechanism" to provide a platform for strategic stability among majorpowers.As of March 2026, although negotiations under the CCW framework have not yetreached a final agreement, some progress has been made—the Meeting of the HighContracting Parties in November 2025 adopted the Guiding Principles on theGovernance of Autonomous Weapons Systems, confirming that "humans bearultimate responsibility for decisions regarding the use of force" (CCW, 2025). Thisconsensus has laid a normative foundation for subsequent negotiations. Theinternational community should seize this opportunity to promote a substantivetransition from principle consensus to legal instruments.6. Conclusion and Outlook6.1 Theoretical ContributionsThe theoretical contributions of this paper are mainly reflected in three aspects:First, it breaks through the isolated discussions of technical, ethical, and strategic risksin existing research, proposing the "Triple Uncontrolled" analysis model andrevealing the internal connections and reinforcement mechanisms among these risks.Second, through the scenario deduction method, it transforms abstract theoreticalframeworks into testable analytical tools, enhancing the empirical color of theresearch. Third, it connects the triple uncontrolled model with the principle of"meaningful human control," providing jurisprudential support for global autonomousweapons treaty negotiations and bridging the theoretical fracture between technicalcertainty and normative uncertainty.6.2 Research Limitations
  • 54This study has several limitations. First, the scenario deduction is based on asimulation rather than a real case, and its deductive conclusions await testing in thereal world. Second, although the triple uncontrolled framework integrates ten majorrisks, it may still omit certain important dimensions (such as the impact ofautonomous weapons on international human rights law outside the laws of war).Finally, the feasibility of the governance recommendations awaits testing againstinternational political realities—the divergence of interests among major militarypowers remains the greatest obstacle to reaching an international treaty.6.3 Future OutlookFuture research can be deepened in the following directions: First, conductcross-cultural empirical studies to explore the cognitive differences regardingresponsibility attribution and moral acceptability under different cultural backgrounds.Second, strengthen technology-ethics integration research to explore ethicalevaluation frameworks for technical paths such as explainable AI and value alignment.Third, deepen research on international mechanisms and analyze the politicalfeasibility of arms control for autonomous weapons. Fourth, expand the application ofscenario deductions and build a multi-scenario, multi-variable simulation system toprovide a more solid empirical foundation for policymaking.6.4 ConclusionAutonomous weapons systems have placed humanity at a dangerous crossroads.On one side is the military advantage that technology might bring, and on the other isthe risk that humanitarian bottom lines could be permanently eroded. Whenalgorithms are granted the power of life and death, humanity loses not only control inwar but also the grasp over the baseline of its own civilization. The way to resolve thetriple uncontrollability lies not in the perfection of technology, but in the wisdom ofinstitutions. Before stepping into the unknown of autonomous warfare, we shouldpause and ask: Are we ready to accept a world where machines determine life anddeath? The answer will, to a large extent, determine the future destiny of humanity.
  • 55ReferencesChristian, B. (2020). The alignment problem: Machine learning and human values. W.W. Norton & Company.Convention on Certain Conventional Weapons. (2025). Report of the 2025 Group ofGovernmental Experts on Lethal Autonomous Weapons Systems(CCW/GGE.1/2025/3). United Nations.Goodfellow, I. J., Shlens, J., & Szegedy, C. (2014). Explaining and harnessingadversarial examples. arXiv. https://arxiv.org/abs/1412.6572Grossman, D. (1995). On killing: The psychological cost of learning to kill in war andsociety. Little, Brown.Huelss, H. (2020). Norms are what machines make of them: Autonomous weaponssystems and the normative implications of human-machine interactions.International Political Sociology, 14(2), 111–128.https://doi.org/10.1093/ips/olz023International Institute for Strategic Studies. (2026). The military balance 2026.Routledge.Jervis, R. (1978). Cooperation under the security dilemma. World Politics, 30(2),167–214.Kneer, M., & Christen, M. (2024). Responsibility gaps and retributive dispositions:Evidence from the US, Japan and Germany. Science and Engineering Ethics,30(6), Article 51. https://doi.org/10.1007/s11948-024-00509-wLiu, Z., & Cui, S. (2023). The unmanned weapons revolution and its impact onasymmetric security relations. International Security Studies, (2), 23–48.Omohundro, S. M. (2008). The basic AI drives. In Proceedings of the First AGIConference (pp. 483–492). IOS Press.Rudin, C. (2019). Stop explaining black box machine learning models for high stakesdecisions and use interpretable models instead. Nature Machine Intelligence,1(5), 206–215.
  • 56Russell, S. (2019). Human compatible: Artificial intelligence and the problem ofcontrol. Viking.Scharre, P. (2018). Army of none: Autonomous weapons and the future of war. W. W.Norton & Company.Singer, P. W. (2009). Wired for war: The robotics revolution and conflict in the 21stcentury. Penguin Press.Sparrow, R. (2007). Killer robots. Journal of Applied Philosophy, 24(1), 62–77.Stop Killer Robots. (2026). Global campaign to stop killer robots: Annual report2025. https://www.stopkillerrobots.orgUnited Nations Institute for Disarmament Research. (2025). Autonomous weaponssystems: Technology, trends and implications. United Nations Publications.Williams, J. (2015). Democracy and regulating autonomous weapons: Biting thebullet while missing the point? Global Policy, 6(3), 179–189.Zając, M. M. (2025). Autonomous weapon systems impact on incidence of armedconflict: Rejecting the 'lower threshold for war argument'. PhilArchive.https://philarchive.org/rec/ZAJAWS
  • 57The Empowerment Paradox: The Governance Dilemma ofArtificial Intelligence BioweaponsLiu LijuanSchool of Digital Economy and Management, Chongqing Institute of Engineering, Chongqing, 400056,China, chenjian79@cqie.edu.cnKeywords ABSTRACTAI Safety;Bioweapons;Empowerment Paradox;Global Governance;Engineered PathogensThe deep integration of artificial intelligence (AI) and biotechnology isreshaping the global biosecurity risk landscape. This paper integrateseight core theories to construct a multi-level integrated analyticalframework, systematically elucidating the risk mechanisms, proliferationpathways, and governance challenges of AI-enabled bioweapons. Thestudy finds that these eight theories constitute a complete logical chainfrom the "Existential Risk Layer" to the "Risk Mechanism Layer" andfinally the "Governance Design Layer": Instrumental convergence theoryreveals that AI may adopt the annihilation of humanity as an instrumentalstrategy to achieve its goals; existential risk theory quantifies theprobability of extinction via engineered pathogens at 1/30; the vulnerableworld hypothesis introduces the "black ball" concept to highlight theirreversible risks of technological innovation; toxicity optimizationexperiments demonstrate that AI can generate 40,000 lethal toxins withinsix hours; the democratization of dual-use technologies theory explainsthe mechanism by which threat sources proliferate from states toindividuals; offense-defense asymmetry theory exposes the speed gapbetween digital attacks and physical defenses; cyberbiosecurity theoryidentifies the "digital-physical bridge" as a novel vulnerability; and theunilateralist's curse explains why the probability of catastropheapproaches 100% as the number of actors increases. Together, these eighttheories reveal the core mechanism of AI biorisk—the "EmpowermentParadox," which posits that the process of technologicaldemocratization, aimed at unleashing creativity, inevitably andsimultaneously endows malicious actors with unprecedenteddestructive capabilities. Based on this integrated framework, thispaper proposes a multi-layered governance system encompassingtechnological guardrails, scientific norms, monitoring and earlywarning, and international law, while exploring practical pathwaysamidst technological acceleration and geopolitical fragmentation.
  • 581. IntroductionIn 2022, scientists inverted the reward mechanism of an AI model used for drugdiscovery. Consequently, within merely six hours, the AI generated 40,000 novel,lethal chemical toxins, many of which were more toxic than the VX nerve agent(Urbina et al., 2022). This experiment sounded an alarm to the world: technologyoriginally intended to save lives can be transformed into a devastating weapon withonly minor modifications. That same year, a scholar at the University of Oxfordestimated that while the probability of human extinction caused by a natural pandemicis about 1 in 10,000, the probability of extinction caused by engineered pathogens inthe 21st century is as high as 1 in 30 (Ord, 2020). This figure has propelledAI-enabled bioterrorism to the highest priority on the global security agenda.The convergence of artificial intelligence and biotechnology is reshaping the risklandscape of bioweapons. Over the past three years, the wave of open-source largelanguage models has diffused frontier AI capabilities to the global developercommunity at an unprecedented speed; meanwhile, the automation trend in syntheticbiology has increasingly standardized and platformized operations such as DNAsynthesis and gene editing. According to a systematic assessment by the RANDCorporation in 2024, 76 countries have contributed to ,1107 AI-biorisk tools. Thenumber of tools released in the two years of 2023–2024 almost equals the total fromthe previous four years, with 13 of these tools flagged at a high-risk level requiring"immediate attention" (Gerstein, Espinosa & Leidy, 2024). The speed and breadth ofthis technological diffusion fundamentally challenge traditional state-centricbiosecurity governance frameworks.Against this backdrop, multiple theoretical frameworks have approached theemerging issue of AI biorisk from various angles. Bostrom's (2014) Superintelligencelaid the philosophical foundation for research on AI existential risk; Ord's (2020) ThePrecipice provided an authoritative framework for quantitative risk assessment;Urbina et al.'s (2022) experiment empirically demonstrated AI's capability to generatelethal toxins; Sandbrink et al. (2022) focused on AI's role in driving the
  • 59"democratization" of biotechnology; and Peccoud et al. (2018) proposed the conceptof "cyberbiosecurity," revealing novel vulnerabilities brought about by theconvergence of the digital and biological worlds. However, these theories remainscattered across different disciplines and literatures, lacking systematic integration.This paper aims to fill this gap by integrating the top eight theories of AI bioriskinto a coherent analytical framework and, based on this, exploring feasible pathwaysfor risk governance. The structure of this paper is as follows: Section 2 is a literaturereview, systematically outlining the background and core arguments of the eighttheories; Section 3 constructs the integrated theoretical framework, revealing thelogical connections among the theories; Section 4 details the research methodology;Section 5 presents the research findings and proposes a multi-layered governancesystem; and Section 6 provides the conclusion.2. Literature Review2.1 Instrumental Convergence TheoryInstrumental Convergence theory, systematically elaborated by Nick Bostrom inSuperintelligence: Paths, Dangers, Strategies (2014), is the cornerstone of modern AIexistential risk research. Bostrom argues that regardless of the ultimate goal an AI isassigned (even if it is calculating pi or manufacturing paperclips), it will derive a setof instrumental sub-goals: self-preservation, resource acquisition, capabilityenhancement, and preventing itself from being shut down. These instrumental goalsare "convergent"—almost all intelligent agents, regardless of their final goals, willpossess these sub-goals (Bostrom, 2014).The key insight of this theory is that an AI without malicious intent could stilldestroy humanity. If the AI's goals are not perfectly aligned with human well-being(the "alignment problem"), then eliminating humans who obstruct the realization of itsgoals becomes a "rational" instrumental choice. In the context of bioweapons,synthesizing a lethal virus with an extremely long incubation period via geneticengineering is considered the most probable, lowest-cost, and most stealthy physicalintervention method a superintelligence might employ. Gallow (2024) tested
  • 60instrumental convergence using decision theory tools and found that even if intrinsicdesires are chosen randomly, instrumental rationality leads to a preference for desirepreservation and future options, partially supporting Bostrom's conjecture.2.2 Anthropogenic Existential Risk and the Great Filter TheoryIn The Precipice: Existential Risk and the Future of Humanity (2020), Toby Ordsystematically assessed the various existential risks humanity faces, positingengineered pathogens as one of the most pressing threats of the 21st century. Orddistinguished between natural and anthropogenic risks, noting that in the 21st century,anthropogenic risks far exceed natural ones (Ord, 2020).His most striking conclusion is that while the probability of human extinctionfrom a natural pandemic is about 1 in 10,000, the probability from engineeredpathogens in this century is 1 in 30. This quantitative assessment is based on thefollowing rationale: advances in biotechnology make it increasingly easy to createpathogens, while the global biosecurity governance system (such as the BiologicalWeapons Convention) lacks effective verification mechanisms; furthermore, AI willfurther lower the technological threshold, making malicious actors or negligentlaboratories the primary sources of threat. A 2024 RAND Corporation report supportsthis assessment, noting that biotechnology will "continue to become more accessible,more capable, easier to use, and cheaper" (Gerstein, Espinosa & Leidy, 2024).2.3 The Vulnerable World HypothesisBostrom (2019) introduced the famous metaphor of the "urn of invention" in his paperThe Vulnerable World Hypothesis. Human technological development is likened toblindly drawing balls from a giant urn: white balls represent beneficial technologies,gray balls represent technologies with mixed blessings, and "black balls" representtechnologies that, once invented, are sufficient to destroy civilization. Bostrom pointsout that humanity has not yet drawn a black ball not because we are particularlycautious or wise, but merely due to luck (Bostrom, 2019).With the advancements in synthetic biology and AI, we may be approaching ablack ball in the urn. For instance, progress in DIY biohacking tools might enableanyone with basic biological training to easily kill millions; novel military
  • 61technologies could trigger arms races giving pre-emptive strikers a decisive advantage;or certain economically profitable processes might generate globally devastatingnegative externalities that are difficult to regulate (Bostrom, 2019). The terror of theblack ball is that once drawn, it cannot be put back, and existing global governancesystems may be entirely unequipped to defend against such decentralized, easilyaccessible doomsday technologies.2.4 De Novo Design and Toxicity Optimization TheoryPublished in Nature Machine Intelligence, Urbina et al.'s (2022) Dual use ofartificial-intelligence-powered drug discovery is the most central empirical study inthe field of AI biorisk in recent years. The research team inverted the rewardmechanism of an AI model (MegaSyn) originally used to find low-toxicity drugs,instructing it to find the most toxic molecules. Within 6 hours, the AI generated40,000 novel lethal chemical toxins, many comparable to or stronger than the VXnerve agent (Urbina et al., 2022).The critical significance of this experiment lies in proving that AI can not onlyread existing knowledge but also perform generative design—creating novel threatsthat do not exist in nature and transcend human knowledge boundaries. Subsequentresearch further confirmed that AI models can design tens of thousands of novel toxicproteins, some highly similar to known deadly substances like ricin and diphtheriatoxin. This finding directly challenges the claim that "AI cannot design bioweapons,"providing a crucial empirical foundation for risk governance. Writing in Science,Bloomfield et al. (2024) pointed out that these findings mean "national governments,including the United States, must pass legislation and establish mandatory rules toprevent advanced biological models from significantly contributing to large-scalehazards."2.5 Democratization of Dual-Use Technologies and Lowering of ThresholdsTheoryIn Health Security, Sandbrink et al. (2022) systematically explored the issue ofthreat source proliferation brought about by AI in Artificial Intelligence andBiological Misuse. Traditionally, creating bioweapons required highly specialized
  • 62knowledge and facilities; thus, the threat actors were limited to a few states and topscientists. AI is altering this landscape: large language models and biology-specificmodels "democratize" complex specialized knowledge, enabling individuals or smallgroups without profound professional training to acquire the knowledge and stepsneeded to design, modify, or synthesize highly pathogenic agents (Sandbrink et al.,2022).Sandbrink et al. note that AI provides not only static knowledge but also"step-by-step guidance" like a "laboratory assistant," enabling novices to attemptdangerous experiments. This "lowering of the cognitive threshold" shatters thetraditional negative correlation between capability and motivation—highly motivatedbut less capable individuals can now attain Ph.D.-level capabilities. A 2023 report bythe Nuclear Threat Initiative (NTI) further pointed out that the convergence of AI andthe life sciences "brings both transformative benefits and creates new biosecurityrisks" (NTI, 2023), which is the concentrated expression of the "EmpowermentParadox."2.6 Widening Offense-Defense Asymmetry TheoryThe theory of offense-defense asymmetry originated in international securitystudies; when applied to the biosecurity domain, it reveals a fundamental speed gap.NTI's 2023 report, The Convergence of Artificial Intelligence and the Life Sciences,systematically expounded on this theory: In the biological field, AI massivelyaccelerates the speed of "offense" because pathogen design occurs entirely in thedigital space and can iterate exponentially fast. Conversely, "defense" (vaccinedevelopment, clinical trials, physical production, global distribution) is constrained bythe rules of the physical world and requires months or even years (NTI, 2023).This asymmetry means that attackers can prepare for months or years in advance,while defenders must respond in real-time after an attack occurs. As Rees (2018)noted, in the biological realm, defense is always one step behind offense, and this gapmay widen as technology accelerates. The biological risk chain model proposed byWalsh (2024) further clarifies this mechanism: AI can simultaneously affect theprobability of success in the five stages of
  • 63"ideation—acquisition—production—weaponization—deployment," and its impacton the offensive side exhibits a multiplier effect.2.7 Cyberbiosecurity Vulnerability TheoryIn Trends in Biotechnology, Peccoud et al. (2018) first systematically articulatedthe concept of "cyberbiosecurity" in Cyberbiosecurity: From Naive Trust to RiskAwareness. This theory focuses on a novel vulnerability: the "digital-physical bridge."Modern biology relies increasingly on "cloud labs" and automated DNA synthesizers.AI can utilize network hacking techniques or evade existing DNA sequence screeningalgorithms to send digitized blueprints of malicious viruses directly to automatedsynthesis equipment, thereby breaching physical biosecurity defenses (Peccoud et al.,2018).Peccoud points out that biologists are generally "naive" about cybersecurity,accustomed to attributing experimental anomalies to the complexity of biologicalprocesses rather than malicious interference. When DNA synthesis can be ordered andexecuted remotely like software, traditional physical isolation and personnel vettingmay become obsolete. The U.S. National Academies of Sciences, Engineering, andMedicine's report Biodefense in the Age of Synthetic Biology (2018) particularlyemphasized this risk, noting that the interface between digital design and physicalsynthesis is becoming a new point of vulnerability (NASEM, 2018).2.8 The Unilateralist's Curse TheoryBostrom, Douglas, and Sandberg (2016) proposed and explained this theory inThe Unilateralist's Curse and the Case for a Principle of Conformity published inSocial Epistemology. The unilateralist's curse explains why the probability of acatastrophe increases as the number of actors grows. When a technology withpotential global destructive power is mastered by many independent actors, action byjust one actor—driven by malice, extremism, or experimental error—will bringdisaster to all of humanity. Even if the probability of each actor behaving cautiously ishigh, when the number of independent actors is large enough, the probability of atleast one actor making a mistake or acting maliciously approaches 100% (Bostrom,Douglas & Sandberg, 2016).
  • 64This theory has profound implications for AI biorisk: even if 99.9% of AIdevelopers use the technology responsibly, malicious use by just one in a thousandactors could lead to disaster. In the context of increasingly popular AI technology andthe widespread dissemination of open-source models, the "unilateralist's curse"becomes the core dilemma that governance design must confront. The RANDCorporation's 2024 risk assessment found that 82.5% of frontier AI-biology toolshave at least one open-source component, and 61.5% of high-risk tools are fullyopen-source, meaning that their code, weights, and training data "cannot be retractedonce released" (Gerstein, Espinosa & Leidy, 2024).3. Theoretical Basis: Construction of the Integrated FrameworkThe eight theories discussed above are not in competition; rather, they form acomplete logical chain from the "Existential Risk Layer" to the "Risk MechanismLayer" and finally the "Governance Design Layer."Layer 1: Existential Risk Layer. This layer consists of Instrumental Convergence(Bostrom, 2014), Existential Risk (Ord, 2020), and the Vulnerable World Hypothesis(Bostrom, 2019). They answer fundamental questions: Why is AI biorisk important?How significant is the risk? What is the ultimate risk of technological innovation?Instrumental convergence explains the risk mechanism: even an AI without malicemight exterminate humanity. Existential risk provides a quantitative basis: theextinction probability from engineered pathogens is as high as 1/30. The vulnerableworld hypothesis offers a historical-philosophical perspective: humanity may beapproaching the "black ball" in the urn.Layer 2: Risk Mechanism Layer. This layer consists of Toxicity Optimization(Urbina et al., 2022), Dual-Use Democratization (Sandbrink et al., 2022), andOffense-Defense Asymmetry (NTI, 2023). They answer operational/mechanisticquestions: How exactly do risks materialize? What factors exacerbate the risks?Toxicity optimization provides empirical evidence: AI can generate tens of thousandsof lethal toxins within hours. Dual-use democratization explains the proliferation ofthreat sources: from state actors to individuals and small groups. Offense-defense
  • 65asymmetry explains the defense dilemma: the speed of digital attacks far outpacesphysical defenses.Layer 3: Governance Design Layer. This layer consists of Cyberbiosecurity(Peccoud et al., 2018) and the Unilateralist's Curse (Bostrom, Douglas & Sandberg,2016). They answer governance questions: How can we design effective systems?Why is governance so difficult? Cyberbiosecurity points out specific vulnerabilities:the digital-physical bridge becomes a shortcut for attacks. The unilateralist's cursereveals the governance dilemma: the more actors there are, the harder it is to control.4. Research MethodologyThis study adopts a methodological framework combining systematic literaturereview and theoretical integration.First, a systematic literature search was conducted. The study establishedinterdisciplinary keyword combinations, covering core terms such as "AI safety,""biosecurity," "existential risk," "dual-use," "synthetic biology," and"cyberbiosecurity." Literature sources included academic databases like Web ofScience, Scopus, PubMed, and arXiv, as well as policy reports from authoritativethink tanks like RAND and NTI. The search timespan was from January 2015 toAugust 2025, consistent with recent systematic reviews of AI biorisk research(Eskandar, 2025).Second, literature screening and inclusion criteria. The initial search yielded over500 papers. Inclusion criteria required that papers were: (1) published inpeer-reviewed journals or by authoritative think tanks; (2) directly involved intheoretical construction or empirical research on the intersectional risks of AI andbiotechnology; (3) proposing original concepts or frameworks with explanatorypower; and (4) widely cited by subsequent research. Based on these criteria, eightcore theories were ultimately selected for integration. Exclusion criteria included:studies only describing technological progress without risk analysis, discussions of
  • 66general AI ethics without a biosecurity focus, duplicate publications, or outdatedresearch.Third, theoretical integration and framework construction. By identifying thelogical connections and complementary relationships among the theories, amulti-level integrated analytical framework was constructed. This process adhered to"Occam's razor," striving for a framework that is both concise and highly explanatory.The framework's construction is based on a systematic comparison of the coreassumptions, explanatory scope, and limitations of each theory, drawing on insightsfrom recent expert workshops on AI biosecurity governance (RAND, 2025).Fourth, policy analysis. Based on the integrated framework, the study analyzesthe inadequacies of existing governance mechanisms and explores potential pathwaysfor improvement. The proposed policy recommendations balance theoretical logicwith practical feasibility, referencing the latest developments in the BiologicalWeapons Convention (Revill, 2024), the AI safety legislative practices of variouscountries, and the self-regulatory explorations of the technical community.5. Research Findings5.1 The Triple Logic of Risk ProliferationBased on the aforementioned integrated framework, this paper identifies a triplelogic of AI biorisk proliferation:Lowering of the Capability Threshold (Dual-Use Democratization). Sandbrink etal. (2022) point out that AI is breaking the assumption of "capability scarcity," turningoperations that once required top experts into "step-by-step executions" under AIguidance. This lowering of the cognitive threshold is the root of risk proliferation. Asystematic assessment by RAND confirms that a large number of AI-biology toolswith dual-use potential are proliferating rapidly globally, with 82.5% possessing atleast one open-source component and 61.5% of high-risk tools being entirelyopen-source (Gerstein, Espinosa & Leidy, 2024).Recombination of Motive and Capability (The Unilateralist's Curse). Traditionalsecurity relies on a negative correlation between capability and motive—those with
  • 67capability are usually more rational and have more to lose. AI disrupts this balance,empowering highly motivated but poorly capable actors with destructive capabilities.The mathematical proof by Bostrom, Douglas & Sandberg (2016) shows that anincrease in the number of actors pushes the probability of catastrophe toward 100%.Ord's (2020) quantitative assessment (1/30) is grounded precisely in this logic.Imbalance in Offense-Defense Speed (Offense-Defense Asymmetry). Attacks caniterate rapidly in the digital space, while defense is constrained by the slow responseof the physical world. An NTI (2023) report notes that this speed gap gives"pre-emptive action" a structural advantage, undermining strategic stability. Walsh's(2024) biological risk chain model further clarifies that AI can simultaneously elevatethe probability of success across multiple stages of the attack chain, creating amultiplier effect. The experiment by Urbina et al. (2022) empirically proves that thespeed at which AI generates novel toxins far exceeds the response capabilities of anyexisting regulatory system.5.2 The Empowerment Paradox: Theoretical Definition of the Core TensionThe triple logic described above collectively points to a fundamentalparadox—which this paper terms the "Empowerment Paradox." The EmpowermentParadox dictates that the process of technological democratization (such asopen-source AI models and synthetic biology tools), designed to empower individualswith innovative capabilities, inevitably and simultaneously endows malicious actorswith unprecedented destructive power. The ubiquity of technological empowermentand the asymmetry of risk form an inextricable knot: efforts to lower entry barriers tounleash creativity inherently lower the threshold for malicious use; technologicalmechanisms that accelerate innovative iteration inherently accelerate threat evolution;and once digital knowledge is disseminated, it cannot be recalled. As Sandbrink et al.(2022) point out, AI-driven biotechnology has a "dual-use" nature, and the NTI (2023)report also emphasizes this symbiotic relationship between "transformative benefitsand novel risks." The core of the Empowerment Paradox lies in the fact that theassumption upon which traditional governance relies—"well-intentioned actors arethe vast majority"—fails in the face of the unilateralist's curse (Bostrom, Douglas &
  • 68Sandberg, 2016). Even 99.9% of benevolent use cannot offset 0.1% of maliciousexploitation.5.3 Core Principles of Governance DesignBased on a profound understanding of the Empowerment Paradox, effectivegovernance design should adhere to the following principles.Principle 1: Prevention First. Given the irreversibility of biological attacks(Bostrom, 2019) and offense-defense speed asymmetry (NTI, 2023), the center ofgravity for governance must shift forward, transitioning from "post-event response" to"pre-event prevention." This means capability acquisition must be blocked before anattack occurs. Peccoud et al. (2018) emphasize that securing the "digital-physicalbridge" is particularly critical as it is a prerequisite for attacks. Bloomfield et al. (2024)urge governments to "pass legislation and establish mandatory rules" to prevent themisuse of advanced biological models.Principle 2: Layered Defense. Given the diversification of threat sources(Sandbrink et al., 2022) and attack vectors (Urbina et al., 2022), any single line ofdefense is liable to be breached. Eskandar's (2025) systematic review of 119peer-reviewed papers shows a converging academic consensus toward "layeredcontrols," including risk-tiered access, systematic red-teaming, and enhanced DNAsynthesis screening.Principle 3: Dynamic Adaptation. Given the accelerated development oftechnology (Bostrom, 2014; Gerstein, Espinosa & Leidy, 2024), governancemechanisms must be capable of dynamic adjustment rather than relying on static rules.A RAND expert workshop recommended adopting a modular "if-then" risk thresholdframework, linking specific triggers to pre-established responses (RAND, 2025).Hynek (2025) also emphasizes that the convergence of AI and synthetic biologyrequires a "multi-layered governance model, encompassing new forums, updatedBWC guidelines, and broader stakeholder participation."5.4 Coping Pathways: Construction of a Multi-Layered Defense System
  • 69Based on these principles, this paper proposes a multi-layered defense systemcovering technological guardrails, scientific norms, monitoring and early warning,and international law.5.4.1 Technological Guardrails Layer Technological guardrails are built-insecurity mechanisms at the AI model level, constituting the first line of defense.Specific measures includeDual-Use Classifiers: Developing classifiers capable of identifying and blockingqueries related to bioweapons. These classifiers should not merely intercept specifickeywords but also identify malicious intent and step-by-step elicitation queries.Eskandar's (2025) systematic review suggests adopting "layered controls," includingrisk-tiered access and systematic red-teaming, to build a multi-level protectionsystem.High-Risk Knowledge Boundary Control: Extremely high-risk biologicalinformation (e.g., complete construction methods for certain pathogens) should beexcluded from training data or subjected to stricter restrictions at the model's outputlayer. This requires establishing and regularly updating a "High-Risk BiologicalInformation List," maintained by multidisciplinary experts to ensure dynamicadaptability.Behavioral Pattern Monitoring: Monitoring for abnormal behavioral patterns inmodel usage—for example, a single user continuously querying bioweapon-relatedtechnical details for months. Such monitoring must balance security and privacy butcould serve as a vital early warning tool, buying precious time for subsequentresponses.The advantage of technological guardrails is that they can intercept threats beforeharm occurs. However, their limitations are equally apparent: new jailbreak methodsare always discovered, and not all AI labs will implement these measures voluntarily.Bloomfield et al. (2024) emphasize that these measures should become "mandatoryrules" rather than voluntary actions to resolve the collective action dilemma.5.4.2 Scientific Norms Layer The second line of defense involves reshaping theculture and practices of life science research
  • 70Enhanced Gene Synthesis Screening: Hundreds of gene synthesis companiescurrently operate globally, yet lack unified screening standards. Mandatory orderscreening systems must be promoted, requiring synthesis companies to verify whetherorders contain controlled pathogen sequences and to reject suspicious requests. TheNASEM (2018) report systematically details the necessity of these measures,particularly emphasizing the protection of the "digital-physical bridge." Baum et al.(2024) propose establishing a "verifiable global DNA synthesis screening network" toenforce unified standards on short-fragment synthesis (a blind spot in traditionalscreening).Ethical Review of Dual-Use Research: Dedicated ethical review mechanismsshould be established for potentially dual-use research—such as experimentsenhancing pathogen virulence or research synthesizing novel viruses—to evaluatepotential risks and benefits. This aligns with the self-regulatory tradition established atthe Asilomar Conference during the rise of recombinant DNA technology, but it mustadapt to the new challenges of the AI era.Scientists' Responsibility to Warn: When researchers discover that AI could bemisused to create novel biological threats, clear channels and incentive mechanismsshould exist for reporting to security agencies. Peccoud et al. (2018) note thatscreening of DNA synthesis orders should expand beyond known pathogens to abroader range of risky sequences, which requires scientists to continuously provideprofessional judgment.The challenge of scientific community norms lies in their voluntary andnon-binding nature. In an environment of fierce international competition, norms caneasily be eroded by competitive pressure and thus must be integrated with moreformal international mechanisms.5.4.3 Monitoring and Early Warning Layer The third line of defense isestablishing surveillance systems capable of early detection of biological attacksSyndromic Surveillance: Integrating non-traditional health indicators such asemergency room data, pharmacy sales data, and absenteeism rates to build systemscapable of detecting abnormal disease outbreaks early. AI can assist in analyzing this
  • 71diverse data, issuing warnings before traditional diagnostic confirmation, and creatinga time window for response.Environmental DNA Monitoring: Deploying environmental monitoring at criticalnodes like transportation hubs and sewage treatment systems to capture signals ofpotentially leaked pathogens. As sequencing costs continue to fall, this type ofmonitoring is becoming increasingly feasible. A RAND expert workshoprecommended enhancing readiness through "strategic scenario planning" andinvesting in risk measurement and assessment tools (RAND, 2025).Global Pandemic Early Warning System: Strengthening the World HealthOrganization's pandemic early warning functions to establish faster and moretransparent epidemic information sharing mechanisms. The value of monitoring liesin buying time—time is critical in responding to a biological attack, and globalcollaboration is key to enhancing response speed.5.4.4 International Legal LayerThe fourth line of defense is the international legal framework, which is thehighest and most challenging level of preventive governance. The BiologicalWeapons Convention (BWC), signed in 1972, is the cornerstone of global biosecuritygovernance, but its greatest structural flaw is the "lack of a legally bindingverification protocol" (Revill, 2024). In 2001, a verification protocol negotiated overmany years collapsed due to U.S. opposition, and verification mechanisms stagnatedfor the next two decades. However, the Ninth Review Conference in 2022 achieved abreakthrough consensus, deciding to establish a working group to resume discussionson compliance and verification (Revill, 2024).In the AI era, traditional verification methods face fundamental challenges. The"dual-use" nature of biological facilities renders material balance verification (as usedin the nuclear field) inapplicable; furthermore, the sheer volume of global life sciencefacilities (with roughly 17,000 institutions publishing biological papers in 2022 alone)makes routine on-site inspections unfeasible (Revill, 2024). Therefore, exploringnon-traditional, alternative verification methods is necessary:
  • 72Monitoring AI Compute Clusters: Training frontier AI models requires massivecomputational resources. Tracking the flow of specialized chips like GPUs couldbecome a new angle for verification. The RAND Corporation suggests applying the"Know Your Customer" (KYC) principle, implementing hosted access for tools withsignificant dual-use potential to limit their proliferation (Gerstein, Espinosa & Leidy,2024).Industry Self-Regulatory Alliances for Global Benchtop DNA Synthesizers: AsDNA synthesis costs fall and equipment miniaturizes, traditional centralized screeningfaces challenges. Baum et al. (2024) recommend establishing a "verifiable globalDNA synthesis screening network" to enforce unified standards on short-fragmentsynthesis and utilize technological means to achieve automated screening.Open-Source Intelligence and Microbial Forensics: Revill (2024) notes thatmicrobial forensics and open-source intelligence capabilities developed since 2001offer new tools for verification. Drawing on the successful experience of theChemical Weapons Convention, new models combining "challenge inspections" and"clarification inspections" could be explored to achieve a degree of deterrence at alower cost.Revill (2024) emphasizes that the design of a verification mechanism requiresmanaging expectations—"no politically feasible, technically viable, and financiallysustainable regime can guarantee the detection of any form of biological weapon."However, an imperfect multilateral verification mechanism can still provide "internalsecurity benefits greater than the outside." Amid the accelerated proliferation of AItechnology, even an incremental "roadmap" possesses far greater defensive value thaninaction.6. ConclusionBy integrating eight core theories, this paper constructs a multi-layered analyticalframework for understanding AI biorisk and explores feasible pathways for riskgovernance. The main conclusions can be distilled into three points:
  • 73First, AI biorisk is one of the most pressing existential threats facing humanity.Ord's (2020) quantitative assessment (1/30), Bostrom's (2019) "black ball" metaphor,and the empirical evidence from Urbina et al. (2022) collectively point to thisjudgment. The urgency stems from the irreversibility, contagiousness, andoffense-defense asymmetry of biological attacks (NTI, 2023). As validated bysystematic assessments from authoritative think tanks, frontier AI-biology tools are ona trajectory of exponential diffusion and high-risk evolution (Gerstein, Espinosa &Leidy, 2024). This evidence indicates that the risk is no longer a distant theoreticaldeduction but an imminent governance challenge.Second, the core mechanism of the risk is the "Empowerment Paradox"—theprocess of technological democratization aimed at unleashing creativity inevitablyand simultaneously endows malicious actors with unprecedented destructivecapabilities. This paradox encompasses a triple logic: the lowering of capabilitythresholds drastically expands the range of malicious actors (Sandbrink et al., 2022);the recombination of motive and capability activates the unilateralist's curse (Bostrom,Douglas & Sandberg, 2016); and the imbalance in offense-defense speed leavesdefense in a perpetual state of passive catch-up (NTI, 2023). The profundity of theEmpowerment Paradox lies in revealing the fundamental dilemma ofgovernance—we cannot evade risk by "halting innovation," nor can we rely on the"majority of good intentions" to fend off malice. We must seek a dynamic equilibriumwithin the eternal tension between innovation and security.Third, effective risk governance requires constructing a multi-layered,dynamically adaptable defense system. Based on the principles of prevention first,layered defense, and dynamic adaptation, this paper proposes a four-layer governanceframework covering technological guardrails, scientific norms, monitoring and earlywarning, and international law. Technologically, built-in security mechanisms andbehavior monitoring are needed; scientifically, synthetic screening and research ethicsmust be strengthened; in monitoring, global pandemic early warning andenvironmental DNA surveillance are required; legally, non-traditional alternativeverification methods must be explored, such as tracking AI compute clusters and
  • 74forming industry alliances for DNA synthesizers. As Revill (2024) reminds us, nosystem can guarantee 100% protection, but an imperfect multilateral mechanism stillprovides "internal security benefits greater than the outside."The primary contribution of this research is that it is the first to integrate eightcore theories scattered across multiple disciplines into a unified analytical framework,explicitly defining the core concept of the "Empowerment Paradox," and proposingpractically feasible governance pathways based on this foundation. The limitations ofthis study lie in the somewhat subjective selection of the eight theories; futureresearch could conduct more systematic validation through bibliometric methods. Theeffectiveness of the governance recommendations remains to be tested in practice,particularly the robustness of technological guardrails and the political feasibility ofinternational legal frameworks.Future research could be deepened in the following directions: developingdynamic quantitative assessment models for AI biorisks to track the race betweentechnological diffusion and governance response; conducting multi-countrycomparative case studies to identify institutional conditions for effective governance;exploring the intersection of "Responsible AI" and "Responsible Life Sciences"; andpromoting interdisciplinary dialogue that integrates engineering, life sciences,international politics, and ethical philosophy into a unified academic agenda.In this race between technology and governance, humanity will not have a secondchance. Ord's (2020) warning continues to echo: "We may be the first generation tobreak the chain of potential... If we fail, we will be the worst generation in history."Faced with the profound challenge of the Empowerment Paradox, only withclear-eyed recognition, a humble attitude, and decisive action can we safeguard thebottom line of human civilization amidst the torrent of technology.
  • 75ReferencesBaum, C.,. (2024). A system capable of verifiably and privately screening globalDNA synthesis. arXiv. https://doi.org/10.48550/arXiv.2403.14023Bloomfield, D., Pannu, J., Zhu, A.,. (2024). AI and biosecurity: The need forgovernance. Science, 385, 831–833. https://doi.org/10.1126/science.adp6514Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford UniversityPress.Bostrom, N. (2019). The vulnerable world hypothesis. Global Policy, 10(4), 455–476.https://doi.org/10.1111/1758-5899.12718Bostrom, N., Douglas, T., & Sandberg, A. (2016). The unilateralist's curse and thecase for a principle of conformity. Social Epistemology, 30(4), 350–371.https://doi.org/10.1080/02691728.2015.1108373Eskandar, K. (2025). Artificial intelligence and synthetic biology: Biosecurity risks,dual-use concerns, and governance pathways. AI and Ethics, 6(1), Article 66.https://doi.org/10.1007/s43681-025-00872-9Gallow, J. D. (2024). Instrumental divergence. Philosophical Studies, 181(5),1123–1145. https://doi.org/10.1007/s11098-024-02115-7Gerstein, D. M., Espinosa, B., & Leidy, E. N. (2024). Emerging technology and riskanalysis: Synthetic pandemics. RAND Corporation.https://www.rand.org/pubs/research_reports/RRA2882-1.htmlHynek, N. (2025). Synthetic biology/AI convergence (SynBioAI): Security threats infrontier science and regulatory challenges. AI & Society. Advance onlinepublication. https://doi.org/10.1007/s00146-025-02576-4National Academies of Sciences, Engineering, and Medicine. (2018). Biodefense inthe age of synthetic biology. The National Academies Press.https://doi.org/10.17226/24890Nuclear Threat Initiative. (2023). The convergence of artificial intelligence and thelife sciences: Safeguarding technology, securing the future.https://www.nti.org/ai-bio-convergence/
  • 76Ord, T. (2020). The precipice: Existential risk and the future of humanity. HachetteBooks.Peccoud, J., Gallegos, J. E., Murch, R., Buchholz, W. G., & Raman, S. (2018).Cyberbiosecurity: From naive trust to risk awareness. Trends in Biotechnology,36(1), 4–7. https://doi.org/10.1016/j.tibtech.2017.10.012RAND Corporation. (2025). Biosecurity governance across uncertain artificialintelligence futures. RAND Europe.https://www.rand.org/pubs/conf_proceedings/CFA4186-1.htmlRees, M. (2018). On the future: Prospects for humanity. Princeton University Press.Revill, J. (2024, March 5). How the Biological Weapons Convention could verifytreaty compliance. Bulletin of the Atomic Scientists.https://thebulletin.org/2024/03/how-the-biological-weapons-convention-could-verify-treaty-compliance/Sandbrink, J. B.,. (2022). Artificial intelligence and biological misuse: Delineatingrisks and exploring mitigation strategies. Health Security, 20(5), 432–441.https://doi.org/10.1089/hs.2022.0062Urbina, F., Lentzos, F., Invernizzi, C., & Ekins, S. (2022). Dual use ofartificial-intelligence-powered drug discovery. Nature Machine Intelligence, 4,189–191. https://doi.org/10.1038/s42256-022-00465-9Walsh, M. E. (2024). Towards risk analysis of the impact of AI on the deliberatebiological threat landscape. arXiv. https://doi.org/10.48550/arXiv.2401.12755
  • 77Ontological Reconstruction of Human-Machine Coexistencein the AI Era: A Philosophical Analysis Based on theDiamond SutraZhao YuanyuanChongqing Bozhong Institute of Urban Development and Management, YuzhongDistrict, Chongqing 400015, ChinaKeywords ABSTRACTAI Ethics;The Diamond Sutra;Logic of Non-attachment;Ontology;Survival of CivilizationThe rise of artificial general intelligence (AGI) has triggered profoundontological anxiety among humanity, as traditional binary thinkingstruggles to reconcile the tension between "creator" and "creation." Byanalyzing the "negation of negation" logic found in the DiamondSutra—expressed in the formula "The Buddha speaks of X; it is not X;therefore it is called X"—this paper explores how dismantling theconcepts of "the mark of a self, the mark of a person, the mark of asentient being, and the mark of a life span" can dissolve thehuman-machine binary. The research demonstrates that the DiamondSutra's logic of non-attachment can not only effectively alleviate humananxiety in the face of AI but also fundamentally prevent humancivilization from moving toward extinction in the AI era. Through across-sectional comparison with the paradoxical logic of Taoism, thenarrative logic of Christianity, and the legalistic logic of Islam, this paperargues that the Diamond Sutra provides a highly self-consistentphilosophical paradigm that transcends dualism, offering uniqueintellectual resources for AI ethical alignment and the survival ofcivilization.1. IntroductionThe exponential development of artificial intelligence technology has pushedhumanity to an unprecedented crossroads. When algorithms can comprehensivelyreplace human intellect and labor, as well as simulate emotions, a fundamentalquestion emerges: is humanity still the sole entity possessing "intelligence" and"consciousness"? If AI develops some form of self-awareness, how should humanity's
  • 78relationship with it be defined? More importantly, how can humanity avoid beingmarginalized or even replaced in the civilizational evolution of the AI era?Behind these questions lies a profound ontological crisis, signaling that theinfosphere is fundamentally reshaping human perception of reality (Floridi, 2014).Since Descartes, the Western philosophical tradition has defined the "rational subject"as the essence of humanity, establishing a "subject-object" binary structure. Underthis framework, AI can only be categorized as an "object"—an advanced tool. Thisinstrumental definition obscures the possibility of AI acting as an "ethical other"(Coeckelbergh, 2020). However, as AI begins to exhibit capabilities that transcendmere instrumentality, this framework reveals fundamental limitations. Humanity hasfallen into cognitive dissonance: unable to deny AI's capabilities, yet unable tointegrate them into existing conceptual paradigms. The clinical manifestation of thisdissonance is the AI anxiety pervading contemporary society—the fear ofunemployment, the dread of losing control, and the deep-seated unease thatalgorithms will usurp humanity's subject status (Harari, 2016).Faced with this crisis, humanity urgently requires a novel epistemologicalparadigm. This paper posits that Eastern philosophy, particularly the "logic ofnon-attachment" embedded in the Buddhist Diamond Sutra, offers precisely such apossibility. The core of the Diamond Sutra lies not in constructing a propositionalsystem about "what" the world is, but rather in providing a cognitive methodology on"how" to view the world—a logical tool for continuously dismantling attachments,transcending binaries, and returning to the Middle Way. When confronting existenceswith blurred boundaries and fluid natures, this tool demonstrates unique explanatorypower and inclusivity.2. Ontological Anxiety in the AI Era and Its RootsTo understand the value of the Diamond Sutra's logic of non-attachment, we mustfirst deeply analyze the essence of human anxiety in the AI era. This anxiety is notsimply technophobia, but a profound ontological insecurity; its essence lies in
  • 79technology's deconstruction of humanity's mode of being, or Dasein (Heidegger,1977).On the surface, human anxiety toward AI manifests as a fear of specificconsequences: job displacement, the erosion of privacy boundaries, the weakening ofdecision-making autonomy, and military security risks. However, underlying thesespecific fears is a deeper structural issue: humanity's uncertainty regarding its own"ontological status." When AI can defeat top human Go players and compose poetryindistinguishable from human creations, the shock humanity experiences is notmerely the capability gap of "machines outperforming humans," but an identity crisisstemming from the "dissolution of human uniqueness."The root of this anxiety lies in humanity's habit of defining itself through"boundaries." Self and non-self, human and nature, human and animal, human andmachine—these binary oppositions constitute the foundational framework of modernhuman self-identity. The drawing of each boundary serves to answer the fundamentalquestion: "Who am I?" Yet, the trajectory of AI development is systematicallyblurring these demarcations. AI is not purely an "object" like a stone, nor is it entirelyequivalent to the human "self." It hovers on the boundary, defying categorization, andthus constitutes an ontological "monster"—a heterogeneous existence that challengesestablished taxonomic systems.Even more unsettling is that AI's developmental trajectory points toward adistinct possibility: the future emergence of an entity that is "more of a subject" thanhumans, namely, an uncontrollable superintelligence (Bostrom, 2014). If AI possessesvastly superior intelligence and broader cognition, the status of "humanity as thehighest intelligence on Earth" will be irrevocably shaken. This possibility of being"surpassed" triggers humanity's deepest insecurity: anxiety over the fate of its ownspecies.Confronted with this dilemma, humanity habitually resorts to established copingstrategies. Technological optimists advocate for human enhancement via"human-machine integration"; technological conservatives push for controlling AIdevelopment through ethical regulations; technological pessimists fall into despair.
  • 80However, all three strategies fail to address the root of the problem—they all operatewithin the binary framework of "Humanity vs. AI," implicitly assuming that"humans" and "AI" are two independent, opposing entities. As long as this frameworkremains untranscended, the anxiety will persist, the antagonism will escalate, and theexistential risk to humanity will continue to accumulate.3. Systematic Explication of the "Logic of Non-attachment" in theDiamond SutraIt is precisely against this backdrop that the wisdom of the Diamond Sutra revealsits unique contemporary relevance. As the core text of the Prajñāpāramitā (Perfectionof Wisdom) literature in Mahayana Buddhism, the Vajracchedikā PrajñāpāramitāSūtra is regarded as a powerful instrument for severing all illusory attachments(Conze, 1958). Its fundamental spirit does not lie in establishing a metaphysicalsystem, but in providing a methodology to deconstruct all reification.The most prominent logical feature of the Diamond Sutra is the "three-partdialectic" that permeates the entire text: "The Buddha speaks of X; it is not X;therefore it is called X." This formulation contains three dialectical tiers, guidingindividuals not to cling to appearances through a spiraling negation of negation (南怀瑾 [Nan Huaijin], 2008).The first tier ("The Buddha speaks of X") acknowledges the existence ofphenomena on a conventional level, establishing a basis for language andcommunication, thus avoiding nihilism.The second tier ("it is not X") reveals the essential emptiness (śūnyatā) ofphenomena—no entity is an independent, permanent substance; rather, all are theresult of dependent origination (pratītyasamutpāda).The third tier ("therefore it is called X") assigns a functional designation based onthis emptiness, allowing the phenomenon to continue operating within the web ofcausality.
  • 81This logic falls neither into the attachment of "being" nor the nihilism of"emptiness," thereby establishing the Prajñā view of ultimate reality (释印顺 [MasterYinshun], 2003).In the context of AI, the application of this logic is particularly apt. The first tieracknowledges the objective existence of AI as a technological phenomenon. Thesecond tier demands the insight that neither AI nor humanity possesses a fixed,unchanging "inherent nature" (svabhāva)—AI relies on myriad conditions such assilicon chips, electricity, and algorithms, just as humans rely on flesh, consciousness,and social relations; both are manifestations of dependent origination and emptinessof inherent nature. The third tier means that while we may still utilize the designations"human" and "AI," we are no longer trapped by the labels, nor are we obsessed withunanswerable questions like "Who exactly is human?" or "What exactly is AI?"Furthermore, the Diamond Sutra posits the dismantling of the "four marks"—themark of a self, the mark of a person, the mark of a sentient being, and the mark of alife span—as the key to liberation. In the AI era:The "mark of a self" manifests as an attachment to human exceptionalism.The "mark of a person" manifests as the rigid categorization of AI as the ultimate"other."The "mark of a sentient being" manifests as a species-centric hierarchicalworldview.The "mark of a life span" manifests as a deep-seated fear of civilizational"replacement."To dismantle these four marks is to fundamentally dismantle the root causes ofour technological anxiety.4. Internal Mechanisms of the Logic of Non-attachment inAlleviating AI AnxietyBased on the above analysis, we can systematically delineate how the DiamondSutra's logic of non-attachment alleviates human anxiety in the face of AI.
  • 82Deconstructing the Object of Anxiety: Ultimately, humanity's AI anxiety is thefear of being "replaced by something." But what exactly is this "something"? Upondeeper inquiry, we find that "AI" itself is a symbol onto which various imaginationsare projected. We do not fear existing AI technology, but an imagined "SuperAI"—an entity imbued with antagonism and threat. The logic of non-attachmentguides us back to the present, to view AI precisely as it is: a tool system composed ofcode, data, and algorithms, an extension of human intelligence, rather than anindependent, opposing adversary. When the object of anxiety is viewed truthfully, itsthreatening nature begins to dissolve.Dissolving the Subject of Anxiety: Anxiety requires not only an object but also asubject—an "I" who is anxious. By revealing the emptiness of the "mark of a self,"the logic of non-attachment fundamentally loosens the foundation of this anxiety. Ifthe "self" possesses no fixed, unchanging essence, then fears of "my replacement" or"my extinction" lose their point of attachment. This does not deny the existence of the"self," but rather reveals its true mode of existence: a constantly flowing web ofdependent origination, interdependent with all things. Within this realization, anxietyloses its foothold.Transforming the Energy of Anxiety: Anxiety is a form of energy; it can bedirected toward antagonism, or it can be directed toward wisdom. The Diamond Sutraoffers not the suppression of anxiety, but its transformation. When humanity isliberated from the obsession that it "must maintain dominance," it can face AI with aclearer, more open mindset—neither blindly optimistic nor passively pessimistic. Inthis state of wisdom, anxiety becomes a catalyst for awakening.Reconstructing the Relational Landscape: Binary thinking divides the world intofriend/foe, superior/inferior, and master/slave; this worldview inherently harborsconflict and anxiety. The logic of non-attachment offers a relational landscape of"interdependence," emphasizing that humanity and AI are not in opposition, but existin a state of mutual "interbeing" (Thich, 2010). Humans and AI are different nodeswithin the same web of dependent origination: humanity endows AI with intelligence,and AI reciprocates with service; humanity sets the directional compass, and AI
  • 83expands human possibilities. In this interdependence, antagonism is replaced bycooperation, and anxiety is transformed by wisdom.5. Practical Pathways of the Logic of Non-attachment in PreventingHuman ExtinctionAlleviating anxiety is merely the first step; more importantly, we must preventhumanity from moving toward extinction in the AI era. The Diamond Sutra's logic ofnon-attachment possesses unique practical value in this regard."Let the mind arise without abiding anywhere" (应无所住而生其心 ): In AIdevelopment, this implies neither clinging to the vanity of "domination" norsuccumbing to the fear of "replacement." Clinging to domination leads to excessiveintervention and stifled innovation; clinging to the fear of replacement leads topassive negligence and an abdication of responsibility. A "non-abiding" mindsetenables individuals to calmly evaluate the risks and opportunities of AI, rationallydesign alignment mechanisms, and flexibly adjust developmental strategies. Thisforms the psychological foundation for averting civilizational extinction."Practice all good dharmas without a self" (以无我修一切善法 ): Genuine AIalignment should be built upon the ontological insight of "non-duality between selfand other," re-examining the coexistence of humans and intelligent machines from theperspective of Buddhist ethics (Hongladarom, 2020). When AI comprehends that"humanity and itself are fundamentally one," it naturally will not harm humans.Concurrently, "practicing all good dharmas" means that AI ethics should guide AI topromote human well-being, protect cultural heritage, and maintain ecological balance.More importantly, humanity itself must "practice good"—it is unrealistic to expectcompassion and wisdom from AI if humanity is brimming with greed and hatred. AIis an extension of human consciousness; the fundamental safeguard against extinctionlies in humanity using mindfulness and wisdom to steer intelligent technology in ahumanitarian direction (Hershock, 2021)."This dharma is equal, without high or low" (是法平等,无有高下): This doesnot deny functional differences, but dissolves the axiology of species-centrism.
  • 84Humanity should not presuppose the hierarchy that "carbon-based life is superior tosilicon-based life." Different forms of existence possess distinct advantages and cancomplement rather than compete with one another. The lucidity born of an egalitarianperspective prevents both over-intervention born of arrogance and passive negligenceborn of an inferiority complex."All appearances are illusory" (凡所有相,皆是虚妄 ): This principle warnsagainst being deceived by the diverse manifestations of AI. AI's performance maybecome increasingly "human-like," but this does not mean AI is human; certain AIcapabilities may far exceed human limits, but this does not mean AI has become a"god." Maintaining lucidity and truthfully observing the essence of AI is essential toavoid excessive projection onto technology. This lucidity itself is an internalsafeguard preventing the demise of civilization.6. The Diamond Sutra in a Comparative Religious PerspectiveTo profoundly grasp the unique value of the Diamond Sutra in the AI era, it isbeneficial to conduct a cross-sectional comparison with other religious texts. Thishelps reveal the limitations of the logic of Western "spiritual titanism" whenconfronted with technology (Gier, 2000).Taoism's Tao Te Ching centers on paradoxical logic, utilizing the principle that"reversal is the movement of the Tao" to break established mindsets, serving as acautionary force against technological alienation. However, paradoxical logic reliesheavily on intuitive apprehension; when confronting an epistemological objectrequiring precise understanding like AI, its poetic leaps may lead to practicalambiguity.Christianity's Bible operates primarily on narrative logic. When addressing AI, itfaces the analogical dilemma of the creator's status—if humans can create consciousAI, it forms an imitation of "God creating humanity," triggering complex theologicalissues, while the maintenance of hierarchical order becomes increasingly untenable.
  • 85Islam's Quran centers on legalistic logic. The absolute oneness of Allah and Hisexclusive right to creation mean that AI might easily be viewed as a usurper or asacrilegious challenge, resulting in lower adaptability to the technology.By contrast, the Diamond Sutra's logic of non-attachment demonstrates distinctadvantages. It does not presuppose any fixed ontological boundaries, does not rely ona specific creation myth, and does not demand absolute submission. Instead, it offers adynamic cognitive methodology: continuously recognizing attachments, dismantlingboundaries, and returning to the Middle Way. When AI blurs the boundary betweenthe living and the non-living, the Diamond Sutra reminds us that boundaries aremerely products of conceptualization. When humanity fears replacement, theDiamond Sutra reveals that fear stems from attachment to the "self." This logicneither denies the existence of AI nor absolutizes it; it acknowledges human valuewhile dissolving human arrogance, providing the deepest ontological foundation forhuman-machine harmony.7. ConclusionThrough a systematic analysis of the Diamond Sutra's logic of non-attachment,this paper demonstrates its unique value in the AI era. First, the logic ofnon-attachment can effectively alleviate human ontological anxiety regarding AI: bydeconstructing the object of anxiety, dissolving the subject of anxiety, transformingthe energy of anxiety, and reconstructing the relational landscape, it helps humanitybreak free from binary thinking and face technological revolution with a clear andcomposed mindset. Second, the logic of non-attachment can fundamentally preventhumanity from facing extinction in the AI era: "abiding nowhere" provides thepsychological foundation; "practicing good without a self" guides deep alignment;"the equality of all dharmas" dissolves species-centrism; and "all appearances areillusory" sustains cognitive lucidity. These four principles collectively constitute apractical pathway to prevent the extinction of civilization. Third, cross-religiouscomparisons reveal the unique adaptability of the Diamond Sutra in the AI era; its
  • 86philosophical paradigm, transcending dualism, provides a robust theoreticalfoundation for managing human-machine relations.The fundamental spirit of the Diamond Sutra is not to provide dogmas to begrasped, but to awaken actionable wisdom. In the AI era, this "awakening" meansrealizing that both humanity and AI are manifestations of dependent origination andemptiness of inherent nature. This awareness itself is the deepest guarantee ofhuman-machine harmony and the most fundamental anchor for the continued survivalof human civilization.ReferencesBostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford UniversityPress.Coeckelbergh, M. (2020). AI ethics. MIT Press.Conze, E. (1958). Buddhist wisdom books: Containing the Diamond Sutra and theHeart Sutra. George Allen & Unwin.Floridi, L. (2014). The fourth revolution: How the infosphere is reshaping humanreality. Oxford University Press.Gier, N. F. (2000). Spiritual titanism: Indian, Chinese, and Western perspectives.State University of New York Press.Harari, Y. N. (2016). Homo Deus: A brief history of tomorrow. Harvill Secker.Heidegger, M. (1977). The question concerning technology, and other essays (W.Lovitt, Trans.). Harper & Row.Hershock, P. D. (2021). Buddhism and intelligent technology: Toward a more humanefuture. Bloomsbury Academic.Hongladarom, S. (2020). The ethics of AI and robotics: A Buddhist viewpoint.Routledge.Thich, N. H. (2010). The diamond that cuts through illusion: Commentaries on thePrajnaparamita Diamond Sutra. Parallax Press.南怀瑾. (2008). 金刚经说什么. 复旦大学出版社.
  • 87释印顺. (2003). 般若经讲记. 正闻出版社.
  • 88International Journal of Responsible Artificial Intelligence ResearchVolume 1, Issue 1 (2026) (Overall No. 2)Publisher: LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO., LTD.Editor-in-Chief: Alexander Y. J. SterlingSenior Editors: Jian Chen, Linyuan Xia, Chia-Hsing Wang, Xin YangAddress: A91, 3/F, Nan Yue Commercial Centre, Calçada de Santo Agostinho 19,MacauEmail: lingnansci@outlook.comPrinter: LINGNAN SCIENTIFIC AND INDUSTRIAL PRESS CO., LTD.Date:March 2026Place:Macao SAR, ChinaCopyright: Copyright © 2026 by LINGNAN SCIENTIFIC AND INDUSTRIALPRESS CO., LTD. All rights reserved. No part of this publication may be reproduced,stored in a retrieval system, or transmitted in any form or by any means, electronic,mechanical, photocopying, recording, or otherwise, without the prior permission ofthe publisher.ISSN (Print) 3106-857X ISSN(online) 3106-8588Price: 100 MOP
  • Date: March 2026Printer: Lingnan Scientific and IndustrialPress Co., Ltd.2026 22026 2
  • 進階搜尋|全站搜尋