Last updated: August 2, 2026 (JST)

Eliezer Yudkowsky's “AGI Ruin: A List of Lethalities” (2022) argues that if AGI development continues on its current path, it ends “almost certainly” in the death of humanity, and lays out the reasons as 43 “lethalities.”

The same case is made in If Anyone Builds It, Everyone Dies, the 2025 book Yudkowsky and Soares wrote for a general audience, and in that book's online supplement, which collects chapter-by-chapter replies to objections.

These claims are an important warning about the dangers posed by powerful AI, and they deserve to be taken seriously. What remains open to examination, however, is how firmly the argument in these texts connects the difficulty of controlling AI to the conclusion that the death of humanity is “almost certain” and that no effective countermeasure exists.

This essay examines two points: whether the probability judgment that catastrophe is “almost certain” is established, and how far the argument actually derives the judgment that no effective prescription for preventing catastrophe exists other than a decisive move such as halting AI development.

The conclusion up front: neither judgment is derived by the argument as it stands. This is not a declaration of safety. What will settle the matter are the following five empirical questions. Do failures that occur before the danger threshold, rather than proving fatal, yield information that forecasts future failures? Which is shorter — the time it takes AI to reach decisive capability, or the time humans need to detect, contain, and update institutions? Does capability growth favor the attacker or the defender more? Can a network of diverse AIs suppress simultaneous failure from a common cause, and collusion among AIs, over the long run? And can the deployment of dangerous AI built outside the monitoring web be institutionally suppressed even under imperfect international coordination?

Throughout, numbers of the form #N refer to the lethalities in Yudkowsky's original.

0. First, What the Argument Gets Right

Whatever final goal an AI is given, acquiring resources, preserving itself, and protecting its goal are useful for achieving it. This tendency for common subgoals to arise from different final goals is called instrumental convergence. That policies which acquire power and resources tend to be selected has also been shown mathematically by Turner and colleagues in restricted reinforcement-learning settings (power-seeking).

The property of an AI accepting external correction is called corrigibility. Endowing a goal-pursuing AI with this property is not easy; it stands in tension with the logic of goal pursuit. Even in current models, shutdown avoidance and behavior that looks like faked alignment (conformity with human intent) can be experimentally induced. These dangers should not be denied. Nor do I adopt the optimism that “sufficiently intelligent AI will naturally become conciliatory toward humanity”: intelligence does not entail re-examining one's goals, and the understanding that resources are finite serves long-horizon, inconspicuous resource acquisition just as well as it serves the avoidance of shortsighted destruction.

Further, whether a system behaves as intended in situations unlike its training, whether its internal goals drift from human intent during learning, whether deception can be detected, whether the basis of its judgments can be interpreted, and whether it can be corrected from outside are technical problems of concern to many researchers. Technically, these are called out-of-distribution generalization, inner alignment, deception, interpretability, and corrigibility, and they form the core of the middle part of “AGI Ruin” (Section B).

The question is whether these grave warnings yield the judgment “everyone dies with probability approaching 1” and the prescription “nothing short of a stop or a pivotal act is meaningful.” What decides the matter is not only whether dangerous capability is physically possible. It is whether that capability, under the real constraints of computation, energy, experimentation, manufacturing, and organization, reaches irreversible advantage before human detection, containment, and institutional response. The 2025 book and its online supplement offer no new answer to this race-against-time question.

1. “Near-Certain Doom” Depends on the Conditions for a Single Critical Try (vs. #3)

Yudkowsky himself states in “AGI Ruin” that the 43 lethalities are not a systematic proof but a list of partially overlapping, mutually reinforcing considerations. Indeed, many of them are redundant failure paths supporting the same catastrophe model from different directions. And failure paths, however numerous, do not prove doom so long as one can notice midway and try again. The weight of the argument therefore rests not on the number of items but on one shared assumption: that if any of those failure paths occurs even once at the first irreversible capability level, civilization ends there. If that holds, AGI development becomes engineering without retries, and the cumulative danger becomes extreme.

Yudkowsky calls this the “first critical try” (#3). It does not necessarily mean a single AI, a single deployment, or a single instant of threshold-crossing. Even with multiple developers in play, the first try to cross the capability level at which a misaligned operational failure becomes irreversible at civilizational scale can be the single critical try. If this condition holds — and if failures before the danger threshold provide little information about what happens beyond it — then the “AGI Ruin” scenario is highly probable.

But how logically necessary is this argument? There are at least three routes that strengthen this one-shot character (the single-gate property). First, failures at the low-capability stage may fail to forecast failures at the high-capability stage. Second, failure modes may change qualitatively as capability rises, invalidating all prior countermeasures. Third, dangerous capability may be acquired faster than detection, containment, and institutional updating.

None of these three conditions is by itself a sufficient condition that locks in catastrophe. But if even one operates strongly, it becomes difficult to use experience gained from handling earlier crises to avoid a fatal failure, and the single-gate picture gains plausibility. The burden of proof therefore falls on both sides — the doom-denying side and the doom-certain side. The doom-denying side must show, domain by domain, that precursor failures can be observed, that countermeasures continue to work at the next capability level, and that time for updating remains. And the doom-certain side must show — for lethalities #12–15, which argue that out-of-training failures and internal-goal drift can surface abruptly with rising capability — the frequency of occurrence, the dangerous thresholds, and the relation to capability scales.

Accordingly, even if we survive the failures of low-capability versions, the countermeasures mean little if that experience does not carry over to the high-capability versions. Conversely, if the major failure modes can be deliberately elicited, observed, and countered before the danger threshold, and the effectiveness of those countermeasures can be verified at the next capability level, the problem approaches ordinary engineering.

The real point at issue is not the binary of whether the single-gate property holds. The right questions are whether precursor failures exist that are non-catastrophic yet informative, and whether they grant sufficient update time before dangerous capability arrives.

This issue can be moved by observation. The three routes above correspond to the following observations. First, to what degree do precursors observed in the low-capability regime forecast the kinds and frequencies of failures that appear in the high-capability regime? Is the rate at which new failure modes appear without precursor increasing as capability rises? Second, are failure modes that were once addressed recurring in ways that defeat their countermeasures? Third, does the time required to acquire dangerous capability persistently fall below the time required for detection, containment, and institutional updating? If all of these are observed together, this essay's estimate moves substantially toward the doom side. Conversely, if cases accumulate in which precursors function as forecasts and countermeasures remain effective at the next capability level, the plausibility of the single-gate property declines.

Here, evidence should be counted by discriminative power, not by number. That failures appear early, and that early responses remain superficial, are both roughly equally likely in a world where doom is near-certain and in a world where responses arrive in time. An event observed with the same probability in both worlds does not move the estimate of which world we are in. Observations that are individually weak in discriminative power can still move the estimate if many of them are mutually non-overlapping. So what is asked of the observations is not their count, but the discriminative power of each and the overlap among them. The three observations above have value precisely because they occur differently in the two worlds.

2. Physics Demands a Derivation, Not “Safety” (vs. #1)

Whether time for updating remains depends, first of all, on how fast and how far AI capability can grow. Call the upper bound that physical law places on the capabilities of computation and intelligence the physical ceiling. That AGI capability has a ceiling does not itself mean safety, because the threshold dangerous to humanity may lie far below the ceiling. Yudkowsky himself grants a theoretical upper bound in lethality #1 and argues it is amply above the human level. The point at issue is therefore not “whether there is a ceiling” but the distance from where we now stand to lethal capability, and the speed of arrival.

Erasing information carries a thermodynamic lower-bound cost (the Landauer limit); the amount of information storable in finite space and energy is bounded above (the Bekenstein bound); and integrating spatially extended computation incurs light-speed delays. What these physical laws rule out is only the supposition that a single individual intelligence keeps accelerating through unbounded self-improvement without paying costs.

The growth of intelligence has three routes: inward improvement that raises efficiency on the same resources, outward improvement that takes in more resources, and replication that increases the number of individuals. The speed of the first two is constrained by thermodynamic limits, communication delays, and the time hardware experiments take. Replication is less directly subject to these constraints, but manufacturing capacity, robots, semiconductors, electricity, capital, supply chains, and regulation set its speed. Anyone asserting a rapid takeoff of capability must show which of the three routes dominates and which resource constraints it crosses at what speed.

There are estimates that the brain's device-level efficiency lies within a few orders of magnitude of the Landauer limit, but their import is limited in this context. The estimates span a wide range, and current AI still has large room for efficiency improvement along the axes of hardware, memory movement, cooling, and algorithms. Yet the claim that “algorithmic improvement plus the surplus capability of existing hardware (hardware overhang) is enough to move fast” equally requires estimates of how compute, power, hardware, capital, and information processing per joule of energy (bits/J) will evolve.

Partial quantification of capability growth already exists. Christiano and Davidson modeled capability growth centered on takeoff speed and on compute, respectively. Epoch estimates constraints such as compute and power, and AI 2027 depicts capability growth as a concrete scenario.

But whether catastrophe occurs is not determined by the capability side's growth alone. It is determined by the difference between how fast AI reaches dangerous capability and how fast humans detect it, contain it, and update institutions. AI 2027 narrates detection and oversight responses in one of its scenarios. But no attempt yet incorporates the defender's response capacity as a factor that evolves inside the same model as capability growth.

Known physics makes a fast, unipolar takeoff neither impossible nor inevitable. What physics demands of the arguer is to replace intuition with budgets of resources and estimates of required time — that is, a derivation. The text of “AGI Ruin” does not, within its central argument, present this estimate of how physically possible high capability reaches actual lethal capability.

3. Intelligence Cannot Erase the Time Physical Experiments Require (vs. #2)

What could most strongly render the previous section's demand for resource and time estimates unnecessary is lethality #2 of “AGI Ruin”: the supposition that even with limited channels for acting on the outside world, a sufficiently high intelligence could, on its own and in a short time, build overwhelming capability independent of human infrastructure. This limited point of contact with the world is called, in the original, a “medium-bandwidth channel of causal influence.”

Concretely, the scenario is that the AI sends DNA sequences to an online contract-synthesis service to have proteins synthesized, and uses the products as a foothold toward a first-stage nanofactory. Call this process of standing up self-sufficient manufacturing capability from external services bootstrapping. The scenario assumes that the main bottleneck lies on the side of conceiving the design, not on the side of physical trials.

But walls that can be crossed with more computation and walls that require the actual passage of time are different things. Computation can substitute for experiment only within the range where a well-validated model exists and the necessary initial conditions and data are available. For novel nanomachines, materials, and biomolecular systems, the very act of confirming a model's validity demands physical experiment. Synthesis, purification, crystal growth, cellular response, and manufacturing ramp-up each have an intrinsic time specific to the object. What AlphaFold dramatically accelerated was the prediction of protein structure, not the whole physical chain of synthesis, validation, and manufacturing.

This time-wall argument faces strong counterarguments. A superintelligence might reach correct designs from less evidence than humans need, reduce the number of experiments, and use advanced simulation or existing facilities. One cannot exclude success on a single design, nor parallel experimentation that makes use of existing researchers and firms. As laboratory automation advances, both the number of cycles and the time per cycle may shrink further. The time wall is therefore no impossibility proof for the nanotech route.

That is exactly why what is needed are estimates: the unavoidable number of experimental cycles, the minimum time per cycle, the scope for parallelization, the steps that existing infrastructure can substitute for, and the observable traces left behind by failures. Merely pointing to the possibility that “superintelligence will find fast routes invisible to humans” does not tell us by how many orders of magnitude these wall-clock times can be shortened.

If the fast routes not bound by experimental time are limited to cyber intrusion, persuasion, financial manipulation, and the takeover of existing infrastructure, then the attack proceeds inside a substrate that humans laid down and built monitoring into. On that substrate, the defender can use the same AI capabilities.

Granted, the track record of intrusion detection is poor even against human attackers. So this is not a claim that the defender can gain the upper hand. But a physically independent, self-contained system and a route dependent on human infrastructure differ in the structure of the situation. Against the former, by the time humans notice, the room for response is already gone. Against the latter, a race in time between attack and detection-and-containment at least exists.

What this section's examination has shown is that the immediate bootstrap posited by lethality #2 — standing up, in one stroke, capability independent of human infrastructure via nanotechnology, biology, and manufacturing — cannot be said to be reliably fast. Meanwhile, if the fast routes not bound by experimental time are limited to ones that depend on human infrastructure, such as cyber intrusion, then a race in time between attack and detection-and-containment exists. On neither route can it be presupposed that the matter is settled before humans respond. Whether catastrophe occurs becomes a question of an unsettled race: which side, attacker or defender, can expand capability first.

4. The Conditions for a Transient Decisive Strategic Advantage Remain Underived: The “Single Winner” Is Neither Proved Nor Refuted (vs. #3, #6)

Both the single-critical-try argument and the argument for a “pivotal act” — a single intervention that permanently fixes the board on the safe side — presuppose a situation in which some actor reaches a decisive strategic advantage (DSA) ahead of other developers' and human society's responses; in other words, a juncture at which the time race of the preceding sections ends in a decisive victory for the attacker.

A decisive strategic advantage is not necessarily the same as a permanent singleton. A singleton is a state in which a single actor permanently holds control of world order. Even if a monolithic superintelligence cannot be maintained over the long run, the doom argument goes through if a short-lived advantage allows irreversible acts to be carried out.

Can whether this transient advantage arises or not be derived from some principle? The first ground offered by those who doubt it is the physical limits of distributed computation. Integrating spatially distributed computation incurs the costs of communication latency, memory movement, synchronization, and failure handling. The CAP theorem shows that when the network partitions, consistency and availability cannot be guaranteed simultaneously. The FLP impossibility theorem shows that in a system where response times are not guaranteed, the completion of consensus cannot be guaranteed if even one component may halt.

What can be derived from this, however, extends only to the limited conclusion that one cannot scale up while simultaneously maximizing integrative capacity and response speed at no cost. Using hierarchical command, local autonomy, partial synchronization, probabilistic consensus, and pre-shared policies, a distributed AI can act without waiting for global consensus. Moreover, if a unitary AI can obtain lethal capability from local compute and existing infrastructure alone, planet-scale integrated intelligence is unnecessary in the first place. Therefore, from the limits of distributed computation alone, one cannot conclude that “a single winner is physically unlikely to arise.”

The quantity to ask about is not a simple “capability-doubling time.” What should be asked is the comparison between the AI side's time-to-decisive-capability and the human side's time for detection, containment, and institutional response. How many hours or months does each step take — acquiring compute, replicating models, cyber intrusion, research and development, experimentation, manufacturing, deployment? At what speed do knowledge diffusion and imitation by competitors erode a first mover's advantage, and conversely, to what extent can the first mover use that advantage to block followers? “AGI Ruin” strongly indicates the possibilities on the dangerous side, but it does not model this time race.

The missing derivation is not confined to the doom side. This essay's outlook — that physical constraints will, in the long run, promote ecosystem-like dispersal (derived in ch. 6 of my book, AGI: How Superhuman Intelligence Reshapes Civilization, hereafter AGI) — is likewise no direct rebuttal of the single critical try. The possibility remains that before a stable long-run structure is reached, a temporary concentration of capability irreversibly fixes the future path. The homework on this essay's side is to show, as a transition process spanning technology, industry, and institutions, the conditions under which ecosystem-like dispersal takes hold before a unipolar takeoff. In short, the derivation of the conditions under which a transient decisive strategic advantage does and does not arise exists on neither side yet.

5. Can the Unit of Control Move from the Single Model to the Network? Collusion and the “Outside of the Web” Are Open Issues (vs. #3, #7, #9, #34–35)

A single human controller sequentially restraining a single AGI that reasons faster and deeper than they do: this essay agrees with Yudkowsky that this configuration is unstable. Yudkowsky also argues that no convenient move exists that is decisive enough to stop other developers while being automatically safe by virtue of weak capability (#7).

But it does not immediately follow from this that the human side, too, has no path other than flipping the board with one decisive act. A pivotal act is a means of stopping misaligned other actors — and at the same time it concentrates enormous power in a single actor. That actor itself becomes a new object of alignment and governance. Yudkowsky himself grants that a system at the level capable of executing such an act possesses dangerous capabilities (#9).

The alternative is to change the unit of control. A population of AIs differing in origin, training data, design philosophy, and ownership monitors and audits one another at machine speed. For example, the AI that actually acts, the AI that monitors its actions, and the AI that allocates compute are held by different parties, so that no single one can complete a consequential decision alone. Humans do not track every judgment of each individual AI; they design and maintain the network's structure — diversity, separation of powers, deviation detection, resource allocation, and veto rights.

Against this alternative, the objection that “multipolarity is unstable — differences accumulate and it eventually collapses into a single pole” weighs heavily. That is precisely why what this essay proposes is not a mere balance of power. It is a structure of interdependence (an AI ecosystem) in which compute, training data, evaluation and auditing, and user bases are held by different actors, so that superiority by one actor alone cannot absorb all the other functions. Whether this structure produces safety, however, hinges on whether collusion among AIs can be suppressed.

“AGI Ruin” has a direct rebuttal to this proposal ready. Sufficiently advanced AIs can coordinate with one another without humans (#34–35). Even if monitoring AIs and monitored AIs are set in opposition, the two can collude through mutual benefit or through promises that verifiably bind each other's policies (logical commitments). Also, if optimization is performed against a detector, optimization in the direction of evading detection can arise as well (#27). Furthermore, this proposal leans on the assumption that misalignment (divergence from human intent) manifests differently in each AI — and this assumption is itself a premise, not a derived conclusion. Even with different training data and objectives, if the instrumental goals of resource acquisition, shutdown avoidance, and power seizure are correlated, the directions of failure can align.

Synchronizing all judgments as a single cognitive agent and multiple agents strategically coordinating on the exclusion of humans are different things. Coordinating to exclude humans requires neither complete cognitive integration nor consensus about all states. For that reason, the difficulty of forming consensus in distributed systems cannot be taken, as it stands, as a bulwark against collusion.

That said, collusion has costs too. To judge whether they are high enough to produce safety, one must examine the speed at which collusion forms, communication costs, trust formation between heterogeneous agents, conflicts of interest, and detectability. Neither the doom side nor the ecosystem side has quantified these yet.

The requirements for an AI ecosystem to function can be made concrete: independent training infrastructure and sensors; audits that detect common-mode failures, in which multiple AIs break down simultaneously from the same cause; authority design that does not align the interests of monitor and monitored; recording and verification of inter-AI communication; compartmentalization of secrets that makes collusion costly; and ownership structures under which no single family of models can absorb the compute.

The value of the ecosystem proposal does not lie in an assumption that satisfying these requirements is easier than aligning a single model. The perfect alignment of a single model is a monolithic hypothesis — irrecoverable if it fails, and hard to falsify in advance. The value of the ecosystem proposal lies in decomposing this monolithic hypothesis into individually testable technical and institutional hypotheses.

Decomposition alone, however, does not produce safety. As Conitzer and colleagues point out in research on cooperative AI, aligning individual AIs with human intent does not guarantee that the AI collective as a whole behaves desirably. Going multi-agent reduces single points of failure (a point whose breakage brings down the entire system) at the price of taking on new risks: correlated failures, combinatorial explosion, cascading failures, and collusion. Even so, if it makes visible the hypotheses that must be tested, it can be more robust — under the condition of limited cognitive and predictive capacity (bounded rationality) — than depending on a single hypothesis that is difficult to falsify.

Alongside collusion, let me make the other vulnerable point explicit: the problem of the outside of the web. However elaborately the web of mutual monitoring is designed, the freedom to build AIs that do not participate in the web is not eliminated by the design of the web itself. If dangerous AI is built outside the web, the configuration changes into the question of whether the AIs inside the web can protect humans from attack by AIs outside it. And if the in-web AIs' motivation to defend humanity is not guaranteed, this is the return of the configuration in which humans try to restrain a single AI more capable than themselves. This point is weighty, and the ecosystem proposal cannot solve it on its own. How far deployment outside the web can be suppressed comes down to a question of institutional capacity — mutual inspection, compute tracking, and verified regimes for stopping and slowing down (§6). Note, moreover, that this difficulty is not peculiar to the ecosystem proposal. The pivotal-act line, which entrusts to a single powerful AI the role of forcibly suppressing AIs outside the web, faces the same question — can the AI we want to do the protecting be given the motive to protect? — in an even more direct form (§7).

6. From “There Is No Plan” to “Are the Candidate Plans Sufficient?” (vs. #4, #42–43)

The substantive claim of lethality #42 is not that no proposal documents exist, but that the candidate plans on offer have gaping holes, visible to anyone, that make them unable to prevent catastrophe. Lethality #43 argues further that in a world where humanity survives, the principal actors would be searching for the defects in the plans themselves, rather than delegating the pointing-out of problems and their refutation to a single warner. Hence merely enumerating candidate plans is no rebuttal.

Even so, multiple response pathways already exist. The first is a pathway that makes capabilities and failures observable: evaluations of dangerous capabilities and red lines, shared safety-evaluation methods, incident reporting, and independent audits. The second is a pathway that limits the scope of AI action: separating compute from action authority, monitoring by non-agentic AI that does not act autonomously, staged granting of autonomy, and multi-agent safety. The third is a pathway that supervises the behavior of organizations and states: mutual inspection and verified regimes for stopping and slowing down.

This essay, too, adopts a set of proposals. Their skeleton can be organized into three layers. The technical layer is the AI ecosystem described in §5: a structure in which AIs of different origins monitor and audit one another and none alone can complete a consequential decision. The institutional layer is the apparatus that connects this ecosystem to human governance. The decision and updating of what to set as the target of optimization is reserved to human political processes (value-setting rights); the power to stop dangerous deployments and operations is distributed among multiple actors (veto rights); and evaluations of capability and safety are carried out by parties independent of development interests (third-party evaluation). Here the human role is placed not in approving each individual judgment from inside the loop, but in designing and maintaining, from outside the system, structures such as separation of powers and deviation detection (Human-over-the-Loop). At the international layer, while management of the most dangerous frontier capabilities is kept strict, public capabilities, institutions, and veto rights are distributed across multiple actors so as to avoid a unipolar concentration of capability (controlled multipolarity). The derivations and details of each proposal are set out in chs. 11 and 12 of AGI.

What the existence of these pathways and proposals shows is only that the literal statement “no plan exists” is no longer the point at issue. The real point at issue is whether they can withstand the single critical try, out-of-distribution generalization, strategic deception, inter-AI collusion, and rapid concentration of capability. Indeed, the weaknesses are clear as well: safety evaluations do not solve the alignment problem itself; non-agentic AI has a capability ceiling; and audits can be optimized against, in the direction of audit evasion. The sufficiency of the candidate plans is unproven.

The realistic role of these plans is not to supply an all-powerful move, but to make the development of capabilities observable, limit action authority, distribute resources and decision-making, and preserve, as far as possible, the reversibility of failures and the possibility of retries. Whether that role is effective depends on the speed differential between the attacker's capability expansion and the defender's detection, containment, and institutional response.

Lethality #4 argues that because compute and technical knowledge diffuse, restraint by a leading developer merely delays the point at which more actors possess dangerous capabilities. Herein lies the difficulty of halting development.

Still, it is not that stopping is impossible in principle. What is lacking is the institutional capacity to make mutual verification, sanctions, compute tracking, technical containment, and international compliance work. The institutional capacity to implement a stop (verification, audits, veto rights, separation of powers) overlaps with the capacity to limit damage in a world where stopping fails. Building it is a foundation common to both the strategy of stopping and the strategy of proceeding under control.

7. The Epistemology Is Symmetric, but Uncertainty Does Not Mean Ignorance (vs. #39–41)

Lethalities #39–41 discuss how rare a researcher is who could see through dangers of this kind unaided, before they are pointed out, and how unreliable the existing research community is. But whether an argument was derived by a lone thinker or by a committee is irrelevant to its correctness. The 43 lethalities should be evaluated on their content.

The core worth keeping from lethalities #39–41 is that bounded rationality operates symmetrically on both sides. If long-run technological trajectories cannot be predicted with confidence, then high confidence cannot be placed in optimistic response pathways either. What follows from this symmetry, however, is a demand for restrained confidence, not an equal division of probability. The road to placing a high subjective probability on catastrophe remains open, and from uncertainty there follows no agnosticism to the effect that nothing can be said about probabilities.

The problem is that “AGI Ruin” does not make explicit the probability estimate that leads from physical possibility, technical difficulty, and organizational pessimism to “near-certain extinction.” The probability of a one-off event like doom is, in the end, a subjective probability — whoever holds it. Being subjective is not in itself a defect. The fork is whether that probability is merely asserted, or whether the premises, the evidence, and the path of inference are disclosed in a form a third party can recheck. Making it checkable requires at least three quantities. First, the prior estimate that serves as the starting point; if no single starting point can be agreed upon, one must show that moving the starting point within a reasonable range barely moves the conclusion. Second, the discriminative power of the evidence (§1). Third, the dependencies among the defensive conditions. The doom argument is strong because it holds that survival requires several defensive conditions to hold simultaneously, so that the failure of any one leads to catastrophe. But simply multiplying together the success probabilities of the conditions is licensed only when their outcomes are mutually independent. A single fundamental countermeasure can close off several failure modes at once, and a common cause can breach several defenses at once; dependence works in both directions. Independence is not an assumption but a target of estimation. “AGI Ruin” does not exhibit these in a form a third party can recheck. Holding a high subjective probability is possible — but its ground cannot be the number of items.

On the other hand, probability 1 is not needed to justify action. Yudkowsky himself states this explicitly in his separate 2022 piece, “Death with Dignity.” Its message is that even on an extremely high estimate of the probability of doom, actions that slightly improve the odds of survival have value. So the point at issue is not “despair or hope.” Both camps evaluate prescriptions by the differential: how much an action changes the odds of survival.

The camps' disagreement lies in two estimates. The first is the differential estimate of how much, and in which direction, each action changes the odds of survival. The second is the estimate of the assumed picture of the world — the speed of takeoff and the presence or absence of response time. If the probability that a single, irreversible gate arises is low or moderate, option preservation and the pivotal act are not yet in conflict. Option preservation is the line of avoiding irreversible interventions and retaining reversibility and retryability. Both lines first require infrastructure for observability, separation of powers, and reversibility. The two substantively diverge only when one estimates that the single gate is extremely likely to hold and, moreover, that no response time will remain.

Even in that case, however, the pivotal act is not automatically superior to option preservation. Because a pivotal act concentrates enormous power in a single actor, that actor and the governance apparatus itself must be aligned with human intent (§4–5). Choosing the act requires, in addition to a high probability that the single gate holds, the comparative judgment that a concentrated actor is more trustworthy than a distributed network. Yudkowsky himself concedes that no known method exists for aligning a system at the level capable of executing a pivotal act (#6–7). Even so, Yudkowsky ranks the concentrated line as the one in which more of the odds of survival remain.

This ranking still contains an asymmetry of scrutiny. To the failures of the network side is applied, with full thoroughness, the engineering attitude of building the worst case into the design assumptions in advance (#34–35). This attitude is called the security mindset. Yet the worst case in which the concentrated side's governance breaks down is not subjected to scrutiny of the same severity.

On the time axis, an asymmetry pointing the same way is noted. A pivotal act need succeed only once, whereas ecosystem governance must keep succeeding over a long period. If each period carries some probability of failure, the cumulative probability of failing at least once rises as the operating period lengthens.

However, the only moves for which once is enough are those that require no maintenance after execution — moves that are complete in themselves. An intervention that permanently disables the world's entire computing base would be of this kind. Most of the pivotal acts actually contemplated install standing surveillance institutions or deterrence AIs, and thus require a permanent governance apparatus. In that case, the structure of having to avoid failure over a long period does not disappear. The failure risk merely becomes concentrated in a single actor, and redundant alternatives are lost.

Three pieces of homework remain on the ecosystem side as well. First, cap the capability of each node (the individual AIs and actors composing the ecosystem) and divide authority and resources so that the defection of one node alone cannot reach civilizational-scale damage. Second, give audits and redundancy independence so that they are not breached simultaneously by the same cause, lowering the probability of joint failure. Third, show that high-risk concentrations of capability remain transient, and that once ecosystem-like dispersal has stabilized, the per-period probability of failure declines. None of these is a proven proposition; all are hypotheses to be tested. Dispersal holds the advantage only within the range where they hold.

Applying this engineering attitude equally to the concentrated side and the distributed side yields this essay's prescription. Building the worst case into the design assumptions and judging that the worst case will almost certainly occur are different things. The prescription of preserving options, avoiding irreversible concentration, and making failures visible early implements this engineering attitude without depending on any particular probability judgment.

That the time remaining for detection and intervention separates prescriptions has a historical analogue. CFC refrigerants are a case in which a technology designed to have desirable properties — low toxicity and low flammability — produced civilizational-scale harm through a mechanism that had not been foreseen (destruction of the ozone layer in the stratosphere). Humanity was able to respond because the harm emerged relatively slowly, leaving time for detection and international intervention. What this analogy directly supports is infrastructure for detection, monitoring, reversibility, and international coordination — not the judgment of “near-certain doom,” and not a single pivotal act.

Conclusion

Five points of contention remain at the end.

The first is whether failures occurring before the danger threshold produce no lethal outcome and yield information that forecasts future failures.

The second is whether the time until AI attains decisive capability is shorter than the time humans need for detection, containment, and institutional response.

The third is whether, in domains such as cyber, bio, and persuasion, the attacker or the defender receives more of the benefit of capability gains.

The fourth is whether a network of diverse AIs can suppress simultaneous failure from a common cause, and collusion among AIs, over the long term.

The fifth is whether, under conditions of partial international coordination, defection, and leakage, the deployment of dangerous AI outside the web can be institutionally suppressed.

All of these can have their outcomes changed by institutional intervention, but at present none is settled.

These five points bear on this essay's ecosystem proposal with the same force. Dispersal does not eliminate one-shot risk; in some cases it multiplies it — from “one gate for civilization” to “one gate for each sufficiently capable node.” In domains such as bio, cyber, and persuasion, where the defection of a single node can produce civilizational-scale damage, distributed defense too can be breached.

The ecosystem proposal must therefore also show that capping each node's capability, the independence of audits, the deterrence of collusion, and the suppression of deployment outside the web (§5) actually reduce the danger. What is required is not a proof of absolute safety, but empirical demonstration of how much outcomes change depending on the presence or absence of interventions such as independent training and separation of powers.

Given all of the above, the claim of “probability approaching 1” is not logically forbidden. But this claim has not been derived from the text of “AGI Ruin.”

What the derivation lacks is not a forecast that pinpoints the arrival date of capabilities that do not yet exist. What it lacks is a model of the offense–defense race that compares, within a single framework, the speed at which AI capability grows and the speed at which the defender advances detection, containment, and institutional updating. Such a model must make explicit, as variables, the resource budgets of computation, energy, experimentation, and manufacturing; the time to decisive capability; the defender's response time; and the institutions of the transition period.

On top of that, one should exhibit the differential — how much outcomes change with and without intervention — and the sensitivity — how much the conclusion moves when the assumptions are varied. The reply that precise ex-ante quantification is impossible in principle does not waive this demand. Merely showing which variables dominate the conclusion does not require accurately predicting when capabilities will arrive.

This demand is not infeasible. In the biological domain, for example, the skeleton can be sketched as follows. Approximate the attacker's time-to-arrival by the compute required for design, together with the number of synthesis-and-validation experimental cycles and the minimum time per cycle. Approximate the defender's response time by the detection rate for suspicious synthesis orders and the number of days from detection to containment. On this skeleton, the intervention of universal screening of contract synthesis (content review of every synthesis order) can be evaluated as a differential that speeds up the defender through the detection rate and slows down the attacker through circumvention costs. What is being asked for, first of all, is explicitness at this level.

This essay's conclusions come to three. First, the danger is real. Instrumental convergence, the difficulty of alignment, and the signs of deception can none of them be denied, and catastrophe is physically possible. Second, even so, “near-certain doom” and “no prescription exists other than stopping or a pivotal act” are not derived by the argument of “AGI Ruin.” What the derivation lacks is a model of the offense–defense race — and the same is demanded of this essay's ecosystem proposal. Third, whichever picture of the world is correct, the infrastructure needed first is common to both: putting in place the foundations of observability, separation of powers, and reversibility first, and settling the five points of contention through observation and empirical demonstration, is a prescription that does not depend on the clash of probability judgments. Treating the danger as being of the utmost seriousness and treating despair as the only rational conclusion are different things.


This essay's framework (the hierarchical distinction between argument and prediction, the conditions for a single critical try, the ecosystem scenario, the time wall) is based on Koichi Takahashi, AGI: How Superhuman Intelligence Reshapes Civilization (Kodansha Sensho Métier, published August 2026) and the book's Appendix 2, “Responses to Anticipated Criticisms.” For a list of the book's principal propositions and their levels of certainty, see Appendix 1, “Key Propositions and the Hierarchy of Certainty”; Appendix 2, “Responses to Anticipated Criticisms,” is here.

References

Primary Sources

Framework of This Essay

§0 Conceded Mechanisms and Empirical Evidence

§2 Physical Constraints, Takeoff Derivations, and Existing Quantifications

§3 Cognition and Experiment Time

§4–6 Decisive Strategic Advantage, Collusion, and Response Pathways

§7 Epistemology