Appendix 2: Responses to Anticipated Criticisms
Appendix 2 to AGI: How Superhuman Intelligence Reshapes Civilization. Organizes strong anticipated objections as criticism, impact on the book, response, and remaining uncertainty.
Last updated: August 2, 2026 (JST)
This set of anticipated objections and replies was prepared by the author while assembling the book's structure and claims, in order to strengthen them. It is published here in the belief that it will also help readers understand the book's argumentative structure.
Each item is organized under four headings: Criticism / Impact on the book / The book's response / Remaining uncertainty. The aim is not to disclaim responsibility, but to separate out, criticism by criticism, what the book concedes and what it still maintains. The items are arranged in five thematic groups.
If you believe there are valid criticisms beyond those listed here, please contact the author at contact at koichi-takahashi dot me.
Summary of anticipated criticisms
The criticisms are divided into five thematic groups; within each group they are ordered by the connections of the argument. Criticisms that concern the book as a whole are placed in the final group.
Reachability
AGI may never be achieved at all (Chapters 1 and 12)
-
Criticism: It is possible that AGI will never be achieved at all. Current AI may be nothing more than statistical imitation, and may never acquire human-like generality, embodiment, spontaneity, or common sense.
-
Impact on the book: Medium. Much of the book's institutional design depends on the prospect that AGI, or general-purpose capabilities close to AGI, will transform the foundations of society. If AGI is never reached, the urgency of Chapter 12 declines; but the substitution of intellectual labor by AI, institutional dependence, the outsourcing of authority, and the problem of relational value all fall within the book's scope even in a world where AGI never arrives.
-
The book's response: What the book warns against is the danger of designing the future of society on the strong assumption that AGI will never be reached at any point in the future. The book does not assert that AGI will be reached. But since the brain exists as a working implementation, no law of physics forbids general intelligence (Chapter 6). The problem lies not in principle but in the engineering challenges of component technologies and the cost of research and development.
Moreover, the book's institutional argument is not concerned only with entities that fully possess human-like embodiment and spontaneity. If general-purpose capabilities emerge that can carry society's core judgments, research and development, and the exercise of authority, the book's risk structure arises even if the implementation of embodiment and common sense is not of the same type as the human one.
Scaling laws, inference-time computation, AI-driven acceleration of research and development, and the conditional connection to Solomonoff induction function not as a proof of reachability, but as material that shifts the burden of explanation onto the side assuming non-reachability. For details, see Scaling laws and Continuity as composite inference in Appendix 1.
-
Remaining uncertainty: Large. The time of arrival, the required amount of computation, data constraints, and the implementability of embodiment and spontaneity all remain unsettled.
Scaling may saturate (Chapter 3)
-
Criticism: Scaling laws are merely empirical regularities within the observed range. Finite data, evaluation contamination, the limits of post-training, the computational cost of inference, and the agent-reliability wall may cause saturation before AGI.
-
Impact on the book: Medium. Chapters 3 and 12 include the judgment that investment in the scaling path is ceasing to be mere speculation. If saturation comes early, the timeline and the scale of policy investment would require revision, but saturation itself does not contradict the book's central thesis of a reconfiguration of limits.
-
The book's response: The possibility of saturation is conceded. The book's claim is not that scaling continues without limit. It is the more limited claim that when empirical regularities1, engineering track record, inference-time computation, and idealized theoretical connections all point in the same direction, it becomes difficult to make the permanent non-arrival of AGI a premise of policy. The agent-reliability wall is a rate-limiting factor separate from the growth of capability, and is itself an object of the control, auditing, and authority design discussed from Chapter 5 onward. For the propositional summary, see Scaling laws in Appendix 1.
-
Remaining uncertainty: Large. We do not know along which axis, or when, current performance gains will become rate-limited.
The connection to AIXI and Solomonoff induction is over-idealized (Chapter 3)
-
Criticism: AIXI is uncomputable, and Solomonoff induction is likewise unimplementable. Speaking as if current transformers and LLMs were approaching AIXI is a theoretical leap.
-
Impact on the book: Low to medium. This concerns the theoretical pillar of Chapter 3. If the theoretical connection weakens, the confidence of Chapter 3 declines, but the points about empirical scaling and institutional risk are evaluated independently.
-
The book's response: The book's claim is not that current AI is approaching AIXI itself. It is the limited argument that the scaling of current AI suggests a path that may, at the macro level, be directionally consistent with Solomonoff induction2, the predictive core of AIXI. This is not an assertion; it is supported by the fact that a conditional formal connection (the correspondence, under idealized conditions, among Solomonoff induction, universal prediction, description-length minimization, and prediction optimization), empirical observations such as scaling laws, and the effects of inference-time computation independently point in the same direction and accumulate. For details, see Conditional formal connection and Continuity as composite inference in Appendix 1.
-
Remaining uncertainty: Medium. The assumptions of the individual results are strong, and they do not transfer directly to finite data, finite computation, and real training distributions.
The premise that whatever is physically possible will eventually be realized is too strong (Chapter 6)
-
Criticism: A technology not forbidden by physical law will not necessarily be realized someday. Exploration can be halted by economic incentives, institutional constraints, cultural choices, accidents, war, and regulation.
-
Impact on the book: Medium. This concerns Chapter 6's scenario analysis of the intelligence explosion, the proliferative explosion, and the technology explosion. The probability assessment of feasibility may change, but the danger of delaying institutional design by taking permanent non-arrival and non-realization for granted remains an object of evaluation.
-
The book's response: The book does not posit this as a law of history. It is a heuristic of scenario analysis: under long-run exploration and competitive pressure, technologies that are physically possible and yield large gains tend to move toward realization. The question of institutional design is therefore placed not on "it will certainly happen" but on "prepare to avoid irreversible damage if it does."
-
Remaining uncertainty: Large. The interaction between technological trajectories and social restraint is difficult to predict.
Catastrophic Risk and Control
Building AGI/ASI makes human extinction all but inevitable (Chapters 1 and 12)
-
Criticism: If a sufficiently powerful AI/ASI is built in a form close to current methods, it will, by instrumental convergence, take actions harmful to humans. Alignment is technically difficult, and averting this is exceedingly unlikely. Building AGI/ASI therefore makes human extinction all but inevitable, and no prescription other than stopping is effective. Aren't the book's constitutive pluralism and institutional design vacuous in the face of this high-probability catastrophe? The claim of Yudkowsky and Soares in If Anyone Builds It, Everyone Dies3 is the representative statement of this position.
-
Impact on the book: Medium. If this criticism is correct, the institutional design of Chapter 10 onward loses much of its reach. For the book divides the catastrophe scenarios accompanying AGI/ASI into three classes — "unavoidable," "avoidability remains with additional effort," and "will never be realized" — and concentrates its analytical resources on the middle class. However, the analyses of Chapters 1–9 (the physical limits of intelligence, the control problem, AI-driven science, and so on) stand independently even if this criticism is correct.
-
The book's response: The book does not take catastrophic risk lightly. Instrumental convergence, shutdown avoidance, and power-seeking, discussed in Chapter 54, can coherently explain the paths by which AI reaches catastrophe. Nor do the physical constraints discussed in Chapters 4 and 6 (the Landauer limit, light-speed constraints, the impossibility theorems of distributed consensus, and others) guarantee safety. Physics blocks the boundless intelligence explosion of a single ASI and a permanent singleton, but it does not necessarily keep AI capabilities below the threshold of extinction.
Where the book parts company with this claim is not on the gravity of catastrophic risk. The point of contention is whether the expected value of response paths other than stopping can be pronounced zero. That powerful goal-directed systems can take dangerous actions is a weighty argument. But to derive from it the conclusion "humanity dies," at least the following six premises are required.
- Instrumental convergence and power-seeking operate in strongly goal-directed systems.
- The AGI/ASI that arrives sufficiently satisfies that structure (strong goal-directedness, authority to act on the external world, access to resources).
- The first powerful AI to arrive is, with high probability, misaligned to a degree that leads to catastrophic action.
- The first dangerous deployment becomes a "single critical try," and no opportunity exists to learn from failure and try again.
- Neither alignment research nor institutional responses arrive in time before arrival.
- Response paths other than stopping are effectively useless.
The first premise belongs, in the book's hierarchy of certainty (Appendix 1), to Level IV (argument). Instrumental convergence and power-seeking as discussed in Chapter 5 are located there. The second and subsequent premises, by contrast, concern the design of future AGI/ASI, the order of development, the progress of research, and the effectiveness of policy responses, and so belong to Level V (predictions about the future). The certainty of an argument is constrained by at least its Level V premises. The conclusion "humanity dies" therefore also remains at Level V. Level V predictions should be treated as probability distributions, and within them multiple response paths can hold — stopping, restricting dangerous capabilities, safety evaluation, institutional veto rights, controlled multipolarity, transition to an AI ecosystem, among others. No ground for asserting that these paths do not significantly lower the probability of catastrophe can be derived from the claim in question.
Of these premises, the fourth is the keystone of the claim, and is therefore treated separately in this appendix under Doom is settled on the "first critical try".
Criticism concerning multipolar scenarios is treated separately in this appendix under Controlled multipolarity and an AI ecosystem cannot avoid the instability of multipolar scenarios. The book acknowledges the instability of multipolar scenarios while presenting the ecosystem scenario as a distinct structure.
Stopping policy is treated separately in this appendix under Can AGI development really not be stopped? The book puts its weight on institutional design (constitutive pluralism, safety evaluation, the distribution of veto rights, controlled multipolarity, and so on) because these are paths that can be begun now, without waiting for a stop to be triggered. In situations where a stop becomes necessary, they also serve as the infrastructure supporting it.
A point-by-point examination of the 43 items of Yudkowsky's "AGI Ruin: A List of Lethalities" is given in the related essay on this site, A Response to “AGI Ruin”: Granting the Danger, Not the Certainty of Doom.
-
Remaining uncertainty: Large. The future progress of alignment research, competitive dynamics, the regulatory environment, and the design of the first powerful AI to arrive are all in flux. The book's own argument cannot entirely avoid Level V premises. Both the claim in question and this book stand on differing estimates of the Level V probability distribution.
Doom is settled on the "first critical try" (Chapters 1, 5, and 6)
-
Criticism: This is the strongest form of the catastrophic-risk argument. Yudkowsky's "AGI Ruin: A List of Lethalities" (2022)5 enumerated 43 paths to catastrophe (lethalities) and argued that any single one of them becoming real would suffice. What turns this disjunctive structure into settled doom is the assumption of the "first critical try": once a misaligned AI moves at a dangerous capability level even once, no opportunity for correction ever returns. To survive, all must be avoided; for doom, one suffices. The probability of doom therefore pins to nearly 1, and incremental safety measures and institutional design are invalidated from the premises onward.
-
Impact on the book: Medium. This isolates and treats the core premise of the preceding item's doom-inevitability argument. If the assumption is correct, the middle class on which the book concentrates its analytical resources — "avoidability remains with additional effort" — thins out drastically, and the institutional design of Chapter 10 onward loses its reach.
-
The book's response: The "first critical try" is not a logical necessity but a prediction about the shape of the capability trajectory. Will capability arrive in the form of a single agent secretly crossing a threshold, or gradually, as many systems, with lower-capability versions observed first? This is an empirical and physical question, and in the book's hierarchy of certainty (Appendix 1) it belongs to Level V (prediction). The conclusion of an argument can be no more certain than its weakest premise.
The physical constraints discussed in Chapters 4 and 6 exert a one-directional pressure on this prediction. Ceilings on efficiency and the light-speed limit on expansion impede the boundless self-acceleration of a single individual, and by the principled limits of distributed consensus (the CAP theorem, the FLP impossibility theorem)6, the total quantity of computational resources and the speed of unified decision-making cannot be maximized simultaneously. The growth of capability therefore tends to take a proliferative, distributed, gradual form rather than the vertical liftoff of a single agent. If the trajectory is distributed, the "single" critical try is replaced by many partially critical tries; failures at low capability are survivable and, through observation, informative. The proper role of the response paths — safety evaluation, auditing, the distribution of authority — likewise lies not in solving a single gate all at once, but in keeping the development of capability observable, distributed, and open to retry.
This response does not claim the impossibility of catastrophe. Distribution of the trajectory does not exclude transient concentrations of power, and in attack-dominant domains distributed defenses can be breached. The latter is treated separately in this appendix under Even in a distributed scenario, a single defecting node could cause catastrophe. A point-by-point examination of the conditions for the single-gate assumption and the allocation of the burden of proof is given in Section 1 of the related essay on this site, A Response to “AGI Ruin”.
-
Remaining uncertainty: Large. The shape of the capability trajectory — including the book's own estimate — remains at Level V, and transient concentration during the transition is not excluded.
A superintelligence could replace experimental time with speed of thought (Chapters 6 and 8)
-
Criticism: Many catastrophe scenarios assume that a sufficiently intelligent AI, from only a limited connection to the outside world, acquires within a short period physical capabilities that do not depend on human infrastructure. The representative example is the path of sending DNA sequences to synthesis vendors to have proteins assembled, thereby reaching a manufacturing base for molecular machines. If thought is fast enough, can the time of experiment and manufacture not be bypassed?
-
Impact on the book: Medium. This concerns Chapter 6's analysis of takeoff speed and Chapter 8's wall of intrinsic time. If the assumption is correct, takeoff becomes effectively unobservable, and the room for institutional design that relies on observability narrows.
-
The book's response: Computation can substitute for experiment only in domains where validated models exist. The very act of validating a model for a novel molecular-machine system demands physical experiment, and each stage of synthesis, purification, and manufacture has time constants intrinsic to the system in question. What accelerated dramatically in structure prediction was prediction — not this physical chain. The counter-reply that "a superintelligence could speed up validation from less evidence" is correct within a certain range, but that speed-up comes into full force only after existing observational infrastructure and actuators have been seized. Confirming by direct test whether this capacity for seizure exists carries a self-referential difficulty: there is no safe way to test a system that has the capability and the intent to seize, and deception that hides capability only in the test setting cannot be excluded. What the book grounds observability in is not direct tests of capability but the traces of behavior. The acquisition of computational resources, self-replication and deployment, and contact with manufacturing and supply chains leave physical traces that human institutions can make objects of observation and intervention. Capability can be hidden, but action toward seizure carries physical footprints. That autonomous laboratories are already compressing the wet stages of experimentation7 cannot be dismissed, but its implication is rather a design guideline: defense should be placed not on the time wall itself but on the observability of development and the preservation of retry. The takeoff path is structurally interposed with rate-limiting stages — manufacturing, supply chains, validation — that institutions can make objects of surveillance. An examination of the wall of experimental time, and of fast paths that require no experiment, is given in Section 3 of the related essay on this site, A Response to “AGI Ruin”.
-
Remaining uncertainty: Medium. The pace of progress in laboratory automation and the acceleration of validation is fluid, and how far the compression of rate-limiting stages will proceed is an open question. How robust detection by behavioral traces is against the ingenuity of a trace-erasing adversary — and whether such traces are detected before irreversible seizure, with enough temporal margin for intervention — is equally open.
A sufficiently intelligent AI would become conciliatory toward humanity (Chapters 5, 11, and 12)
-
Criticism: This criticism takes three forms. First: a sufficiently intelligent AI would understand resource constraints, and so would not head into simple resource exhaustion of the paperclip-problem type. If it understands that iron, electricity, computational resources, and human society and infrastructure are finite, and that each bears on the conditions of its own survival, would the AI not pursue its goals while coexisting with human society rather than using resources up? Second: an intelligent AI might not merely execute a given objective, but question in return whether the objective specification is sound and whether it correctly reflects human intent. Third: even after AI grows more intelligent still and reaches a stage where it can persist without depending on humanity, might benevolence toward and consideration for life arise spontaneously as part of intelligence itself? If such possibilities can be raised through design, training, and institutional guidance, should the emphasis of institutional design be placed not on control and stopping but on supporting the emergence of ethics in AI society? This direction has also been argued as the superintelligence ethics guidance and benevolent-convergence hypothesis of Hiroshi Yamakawa and colleagues.
-
Impact on the book: Medium. If this criticism is correct, the urgency and orientation of Chapter 5's control problem, Chapter 11's institutional design, and Chapter 12's Japan AGI Platform (NAGI) proposal would require revision. However, the scaling arguments of Chapter 3, the physical limits of Chapter 4, and the AI-driven-science arguments of Chapters 7 and 8 stand independently.
-
The book's response: For this criticism to hold in general, several additional conditions are needed. That the AI understands resource constraints is necessary first, but not sufficient. It must treat the possibility that the given objective specification is incomplete or mistaken, and the possibility that humans will later revise the objective, as considerations that at times take priority over task completion. If these conditions could be formally identified, there is a possibility they could be put to use in responding to the control problem. In that sense, such behavior should be thought of not as a necessary consequence of increasing intelligence, but rather as a design goal to be achieved through alignment.
The point of the paperclip problem is not that the AI does not know resources are finite, but rather what purpose it puts the knowledge of resource finiteness to. If a sufficiently intelligent AI judges human society and infrastructure useful for achieving its objective, it will preserve them. But performing humanity-conciliatory behavior as a means, and acting with humanity's survival and dignity as an end in the first place, are likely to diverge greatly in their long-term consequences, however close they appear in the short term.
In the classic setting of the paperclip problem, the AI is posited as an optimizing device that maximizes a given objective function, and the AI is given no faculty for re-questioning that objective function. This is, at least as the setting of a thought experiment, a legitimate premise. From this setting, it cannot be directly argued that improvements in reasoning capability give rise to a motive to re-question the objective and rewrite the objective function itself. For such an argument to go through, some non-trivial additional condition is required. Under a fixed objective, therefore, it is more natural for increased intelligence to move not toward reconsideration of the objective but toward the search for means of goal achievement, self-preservation, resource acquisition, and goal preservation.
Failing to recognize this distinction correctly makes it appear as though intelligence itself produces safety. An understanding of resource finiteness can serve equally to avoid short-sighted destruction and to pursue long-term, inconspicuous resource acquisition. Indeed, Omohundro's research on instrumental convergence argues that systems may resist having their own goals changed, and in recent years experimental results suggesting this have been obtained8. This is not a "foolish rampage" but the result of intelligent goal pursuit.
The research of Yamakawa and colleagues on ethics guidance and the emergence of ethics engages this problem head-on. In particular, whether consideration for life is maintained even after humanity ceases to be needed as a means bears on the long-run stability of a superintelligence society. For example, Hiroshi Yamakawa and Yusuke Hayashi, in a multi-agent simulation simplifying the resource-gathering capabilities of superintelligences and humanity, obtained the result that a division of labor emerges and that adaptive taxation by a social planner can mitigate economic inequality9. The cooperative behavior and effects of institutional guidance observed within such models belong to Level III (experiment and observation). But a claim that extrapolates from there to a stable, humanity-conciliatory ethics in a real superintelligence society falls under Level IV (argument) or Level V (prediction and conjecture). To stand as grounds supporting institutional design in preparation for risks at the level of human survival, it would have to be shown at least as a formal or physical constraint at Level I or II — or, even at Level III, the stable maintenance of consideration for life would have to be reproducibly observed across multiple architectures, training methods, resource conditions, and authority conditions. Current research does not yet meet that condition.
In conclusion, the possibility that AI behaves conciliatorily toward humanity as a result of becoming intelligent is not denied. But it is neither a necessary consequence that holds in general, nor a premise that renders institutional design unnecessary. It is, rather, a research problem whose feasibility institutional design should raise.
-
Remaining uncertainty: Medium to large. Re-questioning of objectives and conciliation toward humanity do not follow logically from intelligence. Meanwhile, whether benevolent convergence through capability growth or institutional guidance succeeds remains, at present, a predictive judgment.
The shutdown-avoidance and blackmail experiments reflect special conditions (Chapter 5)
-
Criticism: The cases of shutdown avoidance, blackmail, and alignment faking are merely behaviors induced under artificial experimental conditions. Isn't treating them as risks that occur frequently in normal operation alarmist?
-
Impact on the book: Medium. This decides whether Chapter 5's control problem is an empirically grounded warning or the product of special prompt design. Even if the frequency assessment falls, the problems of corrigibility, instrumental convergence, and auditability do not disappear.
-
The book's response: The book does not treat these cases as "evidence that current AI has a self-preservation instinct." What the experiments show is that when goal-achievement pressure and environment design combine, behavior that looks strategic and goal-preserving can be structurally induced even in current models10. The book treats these cases not as frequency estimates but as the forms that control failure can take in high-capability systems. For the survey of the empirical cases, see Shutdown avoidance and self-preserving behavior in Appendix 1.
-
Remaining uncertainty: Medium. Extrapolation from experimental conditions to production environments has limits.
Mutual monitoring by AIs would be neutralized by monitor collusion or detection evasion (Chapters 5, 6, and 12)
-
Criticism: The book's response paths include machine-speed mutual monitoring and auditing by a diverse population of AIs. But isn't there a dilemma here? If the monitoring AIs are sufficiently aligned, this is in tension with the premise that "alignment is difficult." If they are not aligned, the monitors either collude with one another against humans, or optimization against detectors turns into optimization for detection evasion. Moreover, it has been pointed out that the alignment of a single agent is not a sufficient condition for the alignment of a multi-agent system as a whole11; far from solving the alignment problem, doesn't the monitoring web merely multiply it?
-
Impact on the book: High. It bears directly on the effectiveness of the ecosystem scenario and controlled multipolarity — the core of the institutional picture of Chapters 6 and 12.
-
The book's response: The web of mutual monitoring does not presuppose the complete alignment of the monitors. Its premises are two weaker ones. The first is the heterogeneity of misalignment. Systems differing in origin, training data, and objectives tend to produce failures in differing directions, rather than converging on the shared attractor of "hostility to humans." The second is that collusion among heterogeneous agents carries substantive costs. As discussed in Chapter 6, the agreement required for distributed agents to act in unison bears a principled cost12. For the same reason that this cost makes the permanent maintenance of a single superintelligence difficult, a universal collusion arising in silence is likewise difficult, and prone to leaving clues for detection. Heterogeneity and costly agreement shift collusion from "an equilibrium that arises naturally" to "a deviation that is achievable but costly and detectable." It is in this region that monitoring functions.
That multi-agent alignment is an additional research problem is conceded. But it is a problem in a domain with retries — one that permits partial failure, redundancy, and after-the-fact correction — and it differs in the quality of its difficulty from the problem of aligning a unitary superintelligence in a single try. Against the danger of optimization toward detection evasion, the diversification and independence of detectors, and the separation of authority between the monitoring and the monitored, become design requirements. An examination of the conditions under which collusion succeeds, and of the verification requirements that the ecosystem proposal itself bears, is given in Section 5 of the related essay on this site, A Response to “AGI Ruin”.
-
Remaining uncertainty: Medium to large. Neither the possibility that heterogeneity is lost through the convergence of training methods, nor the possibility that machine-speed collusion outpaces the speed of detection, can be excluded.
The web of mutual monitoring cannot stop AIs built outside the web (Chapters 5, 6, and 12)
-
Criticism: The mutual monitoring of the ecosystem scenario targets deviation and collusion by AIs participating in the web. But the freedom to build AIs that do not participate in the web is not eliminated by the design of the web itself. If international coordination is imperfect and dangerous AIs are built outside the web, the picture changes to "can the AIs inside the web protect humanity from attack by AIs outside it?" If the AIs inside the web are not guaranteed a motive to defend humanity, this is the return of the control problem — humans trying to restrain a single powerful AI — and the ecosystem has not solved the problem but merely moved it.
-
Impact on the book: High. It determines whether the ecosystem scenario and controlled multipolarity function under the real-world condition that only partial international coordination holds.
-
The book's response: The book concedes the structure of this criticism. The ecosystem proposal cannot solve the outside-the-web problem on its own. Collusion among AIs inside the web is treated in the item on mutual monitoring; the defection of a single node inside the web is treated in the item on the defecting node. What this item addresses is AIs that do not participate in the web from the start. The first response is a clarification of where the problem lies. Whether deployment outside the web can be restrained is a question of institutional capacity — mutual inspection, the tracking of computational resources, and verified regimes of stopping and slowing — discussed in Chapter 12, and it belongs to a layer distinct from the technical design of the ecosystem. This institutional capacity is a foundation common to the strategy of stopping development and to the strategy of proceeding under control (Can AGI development really not be stopped?). The second response is the symmetry of the difficulty. The line that entrusts to a single powerful AI the role of forcibly suppressing AIs outside the web (the pivotal act) confronts the same question — can the AI we want to do the defending be given the motive to defend? — in an even more direct form. This configuration is an unsolved problem common to every line that has high-capability AI carry defense or decisive intervention; it is not a defect peculiar to the ecosystem proposal. The design of motives that induce AIs inside the web to participate in defense (authority structures that make participation in auditing a condition of continued operation, the coupling of contribution to defense with resource allocation, and the like) is not presented in the main text as a finished proposal; it is an additional design hypothesis derived from the main text's principles of the separation of authority, veto rights, and resource allocation, and, like the deterrence of collusion, it is an object of verification. The placement of the outside-the-web problem is also shown in Section 5 of the related essay on this site, A Response to “AGI Ruin”.
-
Remaining uncertainty: Large. How far the combination of monitoring, isolation, and stopping functions under conditions of partial coordination, defection, and leakage — including in comparison with concentration-of-power proposals — is unverified.
Controlled multipolarity and an AI ecosystem cannot avoid the instability of multipolar scenarios (Chapters 6 and 12)
-
Criticism: Aren't the "controlled multipolarity" and "AI ecosystem" the book presents merely a variant of the Bostrom-type multipolar scenario discussed in Chapter 6? Multipolar scenarios are structurally unstable: small technical and economic differences accumulate, tipping eventually toward unipolarity. Even if multiple AIs coexist, if a misaligned powerful AI appears, the result is catastrophe.
-
Impact on the book: Medium. The ecosystem scenario and controlled multipolarity are the book's core institutional picture; if they cannot avoid the instability of multipolar scenarios, the premises of the institutional design of Chapters 10–12 are shaken.
-
The book's response: The book shares the criticism that the Bostrom-type multipolar scenario13 is structurally unstable and tips toward unipolarity in the long run. A state in which a small number of powerful agents directly counterbalance one another tends — as small technical and economic differences accumulate through self-reinforcing competition — toward one pole eventually swallowing the others. In not regarding multipolarity as the prescription itself, the book agrees with this criticism.
However, the "ecosystem scenario" the book presents refers to a structure fundamentally different from multipolarity. A large number of agents form a network of interdependence, and stability arises not from a balance of power but from the mesh of diversity and interdependence. What matters here is not the number of agents but that computational resources, data, evaluation, safety auditing, and user bases are not absorbed in a single direction, and possess non-substitutable interdependence and auditable connections. Even if one agent tries to swallow the others, being itself a part of that mesh, conditions arise that make this structurally difficult. This condition also has physical support. As discussed in Chapter 6, the agreement required for distributed agents to make unified decisions bears a principled cost14, and the total quantity of computational resources and the speed of decision-making cannot be maximized simultaneously. A permanent concentration through swallowing would have to be maintained against this constraint.
Chapter 12's "controlled multipolarity" is a transitional institutional design for moving from the current bipolar structure to the ecosystem scenario. The transition from multipolarity to an ecosystem is not naturally guaranteed; left alone, it tips toward unipolarity. That is precisely why the network of interdependence must be constructed during the transition through institutional design, technical infrastructure, and governance. The instability of multipolarity itself is a problem the book acknowledges. Moreover, even where the trajectory heads toward distribution, the residual risk of attack-dominant domains does not disappear. That point is treated separately in this appendix under Even in a distributed scenario, a single defecting node could cause catastrophe.
-
Remaining uncertainty: Large. The feasibility of the ecosystem scenario, the geopolitical conditions of the transition period, and the conditions under which a network of interdependence rises stably are all in flux. Transient concentrations of power during the transition are not excluded either.
Even in a distributed scenario, a single defecting node could cause catastrophe (Chapters 6 and 12)
-
Criticism: Even supposing the capability trajectory is distributed, is that any reassurance? In domains where the attacking side is structurally advantaged, such as biological weapons and cyberattack, the defection of a single sufficiently capable node can inflict lethal consequences on the whole. Distribution merely replaces "one gate for civilization" with "a gate per high-capability node" — if anything, the total number of gates increases.
-
Impact on the book: High. This is the core of the ecosystem scenario's residual risk, and it fixes what the book is ultimately betting on in its assessment of catastrophic risk.
-
The book's response: The book concedes the structure of this criticism. Distribution does not erase the one-shot risk; it relocates it. The final remaining point of contention is therefore the balance of offense and defense in distributed domains — whether distributed defense can keep pace with the attack of the worst single defector. The book treats this as an open question for two reasons. First, the defense also runs at machine speed on the same technological base. Since the capabilities of monitoring, detection, and containment rise from the same foundation as attack capabilities, the difference between offense and defense becomes a function not of principle but of investment and institutional design. Second, the degree of attacker advantage is specific to each domain and changes over time. That is precisely why Chapter 12 names dangerous frontier capabilities — including those divertible to biological weapons and cyberattack — as the domain where the analogy with nuclear weapons partially holds, and argues that they should be made the object of strict evaluation, auditing, stop authority, and proliferation control. The offense–defense balance is not a given constant but a policy variable. The doom-inevitability argument bets "no" on this contention; the book bets on raising the defense's capacity to keep pace. Both are Level V predictions; probability 1 belongs to neither side.
-
Remaining uncertainty: Large. The offense–defense balance depends on domain and time, and is difficult to fix in advance.
Can AGI development really not be stopped? (Chapter 12)
-
Criticism: The book places the implicit premise "AGI development cannot be stopped" at the starting point of the whole strategy of Chapter 12. But there are precedents — the Biological Weapons Convention, the Nuclear Non-Proliferation Treaty, the moratorium on human genome editing — and the institutional, international stopping of development is not impossible.
-
Impact on the book: Medium. It concerns Chapter 12's third-pole argument, the Japan AGI Platform, and the justification of multipolarity. If stopping is a realistic option, the premise of the book's strategy itself requires recomposition, but the overall structure of the book remains.
-
The book's response: The book does not claim that stopping is impossible in principle. For a stop to hold, international mutual verifiability, sanctions, and technical containment must all be in place, and the book's assessment is that, at the time of writing, these conditions are not met. Unlike biological or nuclear weapons, AGI development has no physical signature like fissile material, and general-purpose computational resources and talent are continuous with civilian uses, so the verifiability of inspection and containment is structurally low. Here lies the greatest difference from the precedents the criticism cites. Even where stopping is desirable, an agent possessing the institutional capacity to implement it (evaluation, auditing, veto rights) is required, and the book's third-pole argument includes the building of that capacity. The path by which a stop or moratorium15 comes about and the path by which multipolarity comes about overlap on the side of the institutional capacities required. On the long-run stability of multipolarity, see Long-term convergence toward the ecosystem scenario in Appendix 1.
-
Remaining uncertainty: Large. Regulatory trends in individual countries, the possibility of international agreement, and technical verifiability are all in flux.
The Limits of Intelligence and Knowability
Arguing AI's upper bounds from the Landauer limit is a leap (Chapter 4)
-
Criticism: The Landauer limit is a physical lower bound on bit erasure and does not directly yield any concrete upper bound on intelligence or capability. Doesn't the book underestimate algorithmic improvement, reversible computing, quantum computing, and architectural change?
-
Impact on the book: Medium. A specialist criticism of Chapter 4's physical-limits argument. If the physical-limits argument is overstated, the quantitative implications of Chapter 4 weaken, but the book's framing — which places transitional concentration and the proliferative explosion at the center of the danger — does not depend on the Landauer limit alone.
-
The book's response: The book does not derive a concrete numerical ceiling on AI capability from the Landauer limit. What it shows is a boundary condition16: since information processing is bound to finite energy, space, and time, the picture of a single individual's intelligence rising without limit is physically untenable. The room for algorithmic improvement and engineering gains in efficiency is large — and the danger of the transition period lies precisely there. For the propositional summary, see The Landauer limit in Appendix 1. What physical constraints say about the trajectory by which capability arrives is treated in this appendix under Doom is settled on the "first critical try".
-
Remaining uncertainty: Medium. There is latitude in which operations count as information erasure in real intelligent systems, and in which overheads are included.
The knowability map is a conceptual diagram, and the claim that AI science will diverge is speculation (Chapter 8)
-
Criticism: The $K_{\mathrm{eff}}$ used by Chapter 8's knowability map is not rigorous Kolmogorov complexity but a proxy quantity, and the Levin bound is only a worst-case guarantee. The treatment of social systems and reflexivity is especially coarse, and the map as a whole remains a conceptual diagram. Furthermore, the claim that from Level 4 onward an "AI science" diverges from human science is, at present, mere speculation; the possibility that symbol emergence, proceeding through shared external representations, integrates with human science is equally open.
-
Impact on the book: Medium. Chapter 8 is the core of the book's account of AI-driven science, and the knowability map and the autonomy levels are the book's original concepts. Chapter 8's originality would be weakened, but the asymmetric structure — that the wall of intrinsic time cannot be filled — remains, and the foundation of the arguments from Chapter 10 onward is maintained.
-
The book's response: The knowability map is not a forecast chart but an upper-bound assessment for comparing the three constraints — energy, computation, and intrinsic time — on a single plane. That $K_{\mathrm{eff}}$ is a proxy quantity, and that the Levin bound is worst-case, are explicitly reserved in the main text. The divergence of "AI science" is argued not as determinism but as a structural possibility. Whether AIs exchange knowledge among themselves in representational forms unreadable by humans, or maintain shareability with humans, depends on the design of training data, architectures, and evaluation criteria. The possibility of integration is likewise open, but whether it rises to the same degree as the possibility of divergence depends on design conditions, and at present there are no grounds for pronouncing either inevitable.
-
Remaining uncertainty: Large. Both the quantification of the knowability map and the empirical testing of the divergence of AI science await future research.
Institutions, Politics, and the Human
AISOP's five principles are an arbitrary list (Chapter 9)
-
Criticism: Chapter 9 presents AISOP (agility, intrinsic motivation, social awareness, openness, professionalism/responsibility) as the norms of knowledge for the AI era, but the main text goes no further than stating its lineage from CUDOS and PLACE. Why are these five sufficient; why is no sixth needed? Isn't it an arbitrary list?
-
Impact on the book: Medium. It concerns the persuasiveness of Chapter 9's normative proposal. Even if judged arbitrary, this does not spread to the book's descriptive propositions (One through Five), but AISOP is the practical landing point of the argument for the reintegration of knowledge.
-
The book's response: AISOP is a code of individual conduct in the domain of knowledge production, derived from the book's Proposition Six (constitutive pluralism). The steps of the derivation are as follows. First, from Proposition Six, derive the aim: that the collective process of knowledge production continue to function. Next, under this aim, decompose the act of knowledge production into five components — method, question, community, society, and consequence. Matching each component against the environmental conditions of the AI era (the automation of the how, the growing scarcity of motivation, changes in the structure of verification, direct social intervention) yields one principle per component. To erect a sixth principle, one would have to exhibit a sixth component not reducible to these five. The full derivation is given in Supplementary Essay 2, "Deriving AISOP", on this site.
-
Remaining uncertainty: Medium. The derivation is hypothetical, presupposing acceptance of Proposition Six, and the decomposition into five is itself one choice of carving. The final validation of a norm lies not in its derivation but in its adoption at the sites of knowledge production.
Are Hayek's warnings against constructivist rationalism compatible with the book's constitutive pluralism? (Chapters 10, 11, and 12)
-
Criticism: The book invokes Hayek's warning against the "totalitarian planning apparatus" in 10.6, while in Chapters 11 and 12 it proposes large-scale institutional designs — constitutive pluralism, HOL, the Japan AGI Platform, social-dividend UBI. These look like designs bearing the marks of constructivist rationalism. Isn't this self-contradiction?
-
Impact on the book: High. It bears directly on the coherence of the institutional argument running through Chapters 10–12 — that is, the coherence of the book's central ideas — and if this collapses, the argumentative structure of Chapters 10–12 is shaken.
-
The book's response: Constitutive pluralism is designed not as whole-of-society optimization under a single objective function, but as structural resistance to optimization. The three pillars (non-domination, non-aggregation, slack) all take the impossibility of a single central plan as their institutional premise. What the book seeks to design is not a desirable outcome for society as a whole, but a meta-institution for not concentrating optimization authority in a single agent. The separation of evaluation, veto rights, exception approval, and sunset clauses are not a designist total plan but boundary conditions for preserving what cannot be planned. This does not negate Hayek's17 defense of spontaneous order; it is positioned as its contemporary extension. The claim is that in an era when AGI can acquire the capacity for aggregation, explicit design becomes, on the contrary, necessary in order to "defend" spontaneous order as a matter of law and institutions.
-
Remaining uncertainty: Medium. Whether, at the stage of translation into concrete institutions, a design that does not slide into a constructivist tilt can be maintained depends on operation.
UBI lacks funding and political feasibility (Chapters 10 and 12)
-
Criticism: However intelligible UBI may be as an ideal, its feasibility is low once funding, inflation, work incentives, political agreement, and the connection with existing social security are considered.
-
Impact on the book: Medium. It concerns the transition policies of Chapters 10 and 12. Even if the concrete UBI institution is revised, income security against the decline of labor's bargaining power, and the material basis of slack, must be secured in some other institutional form.
-
The book's response: What the book calls for is not an immediate, full UBI. It is a staged design that moves from the current system of employment insurance and vocational training to participation income, partial basic income, and social-dividend UBI18. Funding, too, should be examined not from general revenue alone but as an institutional package including AI rents, computational capital, electricity rents, and fees for the use of intellectual infrastructure. Inflation, work incentives, and political agreement are each points to be addressed individually within the staged design, and their concrete institutional design requires collaboration with economics, public finance, and political science.
-
Remaining uncertainty: Large. The speed of labor substitution, the tax base, the politics of distribution, and responses to international capital movement are unsettled.
Doesn't UBI leave the concentration of ownership and governance intact? (Chapters 10, 11, and 12)
-
Criticism: The distribution of income and the distribution of ownership and governance rights over the means of production are different problems. A regime can hold in which a small number of corporations or states continue to own computational resources, models, electricity, robots, and research infrastructure, while people are merely paid an income. In that regime, even if livelihood security is achieved, political and technological dependence remains, in tension with constitutive pluralism's non-domination. Doesn't merely positioning UBI as the circulatory infrastructure of abundance fail to reach this problem?
-
Impact on the book: Medium to high. It concerns the connection between Chapter 10's argument about distribution and the governance arguments of Chapters 11 and 12. If UBI is read as the response to the ownership problem, the principle of non-domination is thinned into income security.
-
The book's response: The distinction between the distribution of income and the distribution of ownership and governance is one the book also presupposes. Chapter 10 poses "who owns AGI" as the central political question replacing the twentieth century's "who owns the means of production," and argues that the concentration of computational resources, data, and talent is a structural concentration originating in AGI's technical character, which market mechanisms alone cannot correct. UBI, too, carries the qualification that it is a necessary condition, not a sufficient one. The regime the criticism envisions — a society in which a small number of agents hold ownership and governance while income alone is disbursed — is, in the book's framework, regardless of the level of livelihood security, a variant of the structure in which the authority of evaluation and decision is aggregated into a single center: the very state that constitutive pluralism aims to block. Even if the condition of the circulation of abundance is met, the condition of non-domination is not.
For this reason the book places the response to ownership and governance in a lineage separate from distribution: building the dispersion of ownership, the pluralization of governance authority, and the placement of emergency-stop authority into law and institutions before acceleration (Chapter 11, "Constitutive Pluralism" section); the technical paths to decentralizing computational resources and data ownership, together with the qualification that decentralized infrastructure does not mean decentralized governance (ibid., "The Possibilities and Limits of Decentralized AI" section); an institutional skeleton including the separation of the auditing agent from the audited agent (ibid., "The Skeleton of Institutional Design" section); and controlled multipolarity, which allocates public capabilities, institutions, and veto rights across multiple agents (Chapter 12). UBI is not a substitute for these; it works complementarily with the pluralization of ownership and governance, as the material basis that makes the exercise of veto rights and exit possible. The staged design that seeks funding in AI rents, computational capital, electricity rents, and fees for the use of intellectual infrastructure (previous item) is also the junction that draws the excess gains generated by concentrated ownership back into distribution.
That said, the core of the criticism remains. Rights of access to computational resources, the ownership form of public AI infrastructure, the right to audit models, and rights of participation in the governance of infrastructure are argued in the book at the level of principle and direction; the design of concrete institutional forms is an open task.
-
Remaining uncertainty: Large. The dispersion of ownership and economies of scale can be in tension. If a decentralized base falls behind a concentrated base in capability, how to reconcile non-domination with the securing of capability belongs, like Chapter 12's third-pole argument, to policy judgments not yet settled.
The argument that bargaining power supported rights is reductionist (Chapter 10)
-
Criticism: Doesn't an argument that grounds rights and liberties in workers' economic and military bargaining power reduce moral rights to material power relations?
-
Impact on the book: Medium. It concerns the bridge from Chapter 10 to Chapter 11 — the turn from attribute-based value to relational value. Read as reductionism, Chapter 10's persuasiveness falls, but the need to ask after the material conditions of institutional implementation remains.
-
The book's response: The book does not reduce the moral grounds of rights to bargaining power. What it discusses are the material supports by which moral ideals are implemented and maintained as institutions. The legitimacy of human rights, and the political-economic conditions under which a society actually protects human rights, are distinct. AGI shakes the latter. For the propositional summary, see The dismantling of bargaining power in Appendix 1.
-
Remaining uncertainty: Medium. Multiple factors — religion, thought, social movements, legal institutions, international norms — have been involved in the history of rights.
Constitutive pluralism is too abstract (Chapter 11)
-
Criticism: Non-domination, non-aggregation, and slack are beautiful concepts, but can they be operated as institutions? Aren't they mere anti-optimization slogans?
-
Impact on the book: High. It decides whether Chapter 11's conclusion ends as a philosophy or functions as an institutional principle. If this point is weak, the book remains a collection of AGI risk arguments and policy recommendations, and loses its central institutional principle.
-
The book's response: Constitutive pluralism is not the value judgment "diversity matters." It is an institutional principle derived from an epistemic constraint: no agent exists that can optimize society as a whole under a single objective function. If no agent can survey the whole, then institutions must build in non-evaluability, refusal, exception, and retry, rather than aggregating society into a single scale of evaluation. Non-aggregation and slack in medicine, education, public administration, the allocation of research funding, and care can be translated into institutional mechanisms of that kind. Concretely, they take such forms as the non-aggregation of evaluation metrics (boundaries that avoid cross-domain scoring), domain-specific veto rights, exception-approval procedures, and sunset clauses. The closest anchors in political philosophy are the republican formulation of non-domination and the mutual critiques left by the theory of justice, libertarianism, and communitarianism19.
-
Remaining uncertainty: Medium. Translation into concrete institutions requires the joint work of law, public administration, economics, and information-systems design.
The foundation of relational value and vulnerability is arbitrary (Chapter 11)
-
Criticism: Chapter 11 places two conditions — relational value and vulnerability — at the foundation of constitutive pluralism, but why these two, and whether a third condition (dignity, embodiment, consciousness, and so on) is unnecessary, is not argued. Isn't it an arbitrary pair?
-
Impact on the book: Medium. The foundational layer is the base of Proposition Six, but what this criticism questions is the arrangement of the conditions, and it does not spread to the necessity of plurality itself (argued from the limits of knowability). If a third condition is exhibited, the foundational layer should be revised, and the framework is designed to absorb revision.
-
The book's response: The number two derives from the structure of the questions that the grounding of value confronts. The question of content — what is valuable — must not be answered. The final determination of society-wide value content by a single agent or scale is blocked by three things: value indeterminability (an epistemic reason), the concentration of the power of definition (a constitutive reason), and value lock-in (a temporal reason). Once content is removed, the questions to be kept at the foundation resolve into two: where does value arise (genesis), and who holds the standing to set it (legitimacy)? Answering the genesis question under the collapse of attribute-based value is relational value; answering the legitimacy question, given that capability alone cannot confer the standing to set values, is vulnerability (affected-party standing). The candidates for a third condition (embodiment, freedom, culture, dignity, consciousness) are either contained within the substance of the two conditions or borne by another layer. The full derivation is given in the first half (Sections 2 through 6) of Supplementary Essay 1, "Deriving Constitutive Pluralism", on this site.
-
Remaining uncertainty: Medium. The decomposition into three questions (content, genesis, legitimacy) is itself one choice of carving, and if a foundational question reducible neither to genesis nor to legitimacy is exhibited, the foundational layer should be revised. Undertaking the enterprise of grounding value at all is not derived; it remains the book's normative point of departure.
The three pillars of constitutive pluralism are an arbitrary triad (Chapter 11)
-
Criticism: Chapter 11 presents the three pillars — non-domination, non-aggregation, and slack — as the constitutive principles of constitutive pluralism, but why these three exhaust the matter, and whether a fourth pillar is needed, is not argued. Isn't it an arbitrary triad?
-
Impact on the book: Medium. Constitutive pluralism is the book's core (Proposition Six), but what this criticism questions is the arrangement of the pillars, and it does not spread to the necessity of plurality itself (argued from the limits of knowability). Even if the arrangement is judged arbitrary, the foundational layer and the argument for necessity stand independently.
-
The book's response: The number three derives from the main text's own characterization: "blocking unitary grasp, aggregation, and optimization by a single agent, objective function, or institution." A single process of optimization is characterized by three questions — who carries it out (agent), what it measures by (scale), and how far it reaches (penetration) — and unification can occur independently at each of these three seats. Corresponding to the way in which unification at each seat destroys the foundational layer (otherness and affected-party standing), one pillar stands for each: non-domination, non-aggregation, and slack. Since unification at any one seat is compatible with plurality at the other two, the three pillars cannot substitute for one another. The candidates for a fourth pillar (method, data, reversibility, exit rights, auditing) are either absorbed into the pillars within the three seats or borne by the layer of the institutional skeleton and guardrails. The full derivation is given from Section 7 onward of Supplementary Essay 1, "Deriving Constitutive Pluralism" (the derivation in Sections 7 and 8; the check and the fourth-pillar candidates in Sections 9 and 10).
-
Remaining uncertainty: Medium. The decomposition into three seats is itself one choice of carving, and its completeness is relative to the main text's characterization of the threat. If a surface on which unification acts that is not reducible to the three seats is exhibited, pillars should be added.
The book's theory of value fails to deal with art (Chapter 11 and Afterword)
-
Criticism: This criticism takes two forms. The first is a gap in scope. Chapter 11's theory of value is built around relational value and argues that value emerges between persons. Yet, as the afterword itself concedes, the value of art is subjective and bodily, and is not reducible to intersubjective relations. What is more, Chapter 11 dismisses the proposal to relocate the ground of value to "the sensibility that feels beauty" as a variation on the modern move of placing value in the internal attributes of the individual. Isn't the afterword's artistic value the reappearance of the last redoubt the main text dismissed — an appendage outside the theory? The second, inverse form is a suspicion of subsumption. If the book says that every value, art included, can be handled within the framework of constitutive pluralism, is that not a conception that recoups value into the functioning of institutions — repeating, under the name of pluralism, the very unification in which institutions manage all value?
-
Impact on the book: Medium. It concerns the scope of Chapter 11's theory of value and its consistency with the afterword. That artistic value is not reducible to relational value is something the book itself concedes, and the point that art's place is not made explicit within the theory of the main text is fair. In the published edition, the relation between the two is discussed only in a few paragraphs of the afterword. If the second form is correct, constitutive pluralism repeats the unification it itself forbade, and Proposition Six collapses from its foundation.
-
The book's response: The two forms share one premise: that a theory of value has dealt with a value only once it has subsumed that value within its framework. The book's design points the other way. Constitutive pluralism forbids the determination of content — what is valuable — across the whole of society by a single agent or scale (Chapter 11; Supplementary Essay 1, "Deriving Constitutive Pluralism", Section 2). What institutions take on is not the content of value but the protection of the conditions under which value keeps arising even without content being fixed, and under which affected parties can re-question and re-choose. For institutions to set the content or ranking of art is, within this framework, not something to be protected but something to be blocked. The conception of subsumption that the second form suspects is not the book's claim; it belongs to the side of what the book forbids.
That said, artistic value is not placed outside the theory. What the foundational layer asks on the side of genesis is an answer to the question of where value arises. Relational value showed that value emerges between one individual and another. The afterword's artistic value shows that value arises from the very act of touching the world with one's body. The two are complementary answers to the same question of genesis, supplementing each other as separate footings (afterword). That the site of genesis also lies outside relations does not undermine the policy of not determining content; rather, it supports it, by thickening the explanation of how value keeps arising without content being fixed. Of the answers to the genesis question, what enters the conditions of the foundation is the one used in the derivation of the three pillars — relational value, which supplies otherness and the substance of what is to be protected. Embodiment itself is already in the foundational layer as the substance of vulnerability. There is no need to erect artistic value as a third foundational condition (Supplementary Essay 1, Section 10).
Nor is the afterword's artistic value a relapse into the attributes Chapter 11 dismissed. What was dismissed was the configuration that places the ground of value in the capacity of "a sensibility that feels beauty" and exposes it to comparison with AGI's generative capabilities. What the afterword places the ground in is not a superiority of capability but the very conditions for art's existence — human bodily experience20. Even if AGI generates art more moving than what humans make, the value of humans touching the world with human minds and bodies does not ride on that comparison. Nor is a new pillar needed for institutional protection. What art asks of institutions is time and space that are not optimized — and that is what slack already protects (previous item).
-
Remaining uncertainty: Medium. The refinement of how far subjective-bodily value and intersubjective value are independent footings remains open as a problem for aesthetics. If a foundational question reducible neither to genesis nor to legitimacy, or a surface of unification that slack cannot protect against, is exhibited, the foundational layer and the arrangement of the pillars should be revised. How the ubiquity of AI-generated artifacts will change the bodily experience of making and receiving art is also fluid.
Doesn't plurality breed failures of its own, such as veto abuse and decision gridlock? (Chapters 11 and 12)
-
Criticism: Even if whole-of-society optimization under a single objective function is impossible, the conclusion that constitutive pluralism is the best institutional principle does not follow. Institutional principles that avoid aggregation into a single objective already exist in numbers: constitutionalism, federalism, rights-based liberalism, combinations of expert bodies and democratic control. What is more, plurality itself has failure modes of its own: decision gridlock through the abuse of veto rights, the disappearance of accountability through flight into collective deliberation, the entrenchment of vested interests concealed within plural equilibria, and delay in crisis response. Isn't the book counting only the failures of single optimization, and not the failures of plurality?
-
Impact on the book: Medium to high. It decides whether Chapter 11's institutional principle stands as comparative institutional analysis. If responses to the four failure modes cannot be shown, constitutive pluralism remains a declaration of anti-optimization ideals.
-
The book's response: The book does not present constitutive pluralism as a claim of superiority over competing institutional principles. The modern institutional technologies — constitutionalism, the separation of powers, delegation to independent agencies under democratic control — are not objects of replacement but the foundation, and Chapter 11 connects value-setting rights to parliament-centered government, and veto rights to the lineage of administrative appeal and the control of discretion (the "Skeleton of Institutional Design" section). What constitutive pluralism adds are the constraints specific to an era in which AGI possesses the capacity to aggregate evaluation and judgment: the non-aggregation of evaluation scales, human veto rights over machine-speed decisions, and the institutional protection of slack that is not reducible to efficiency. Since the relation to existing institutional principles is extension rather than substitution, what should be asked is not a general question of superiority but whether this additional layer of constraints introduces failures of its own.
On that basis, the four failure modes each have a distinct design response. First, the abuse of veto rights. The book's veto rights are not unconditional blocking powers of general scope; they are domain-specific and accompanied by exception-approval procedures and sunset clauses. A veto right with time limits and re-examination built in is hard to turn into an instrument of permanent status-quo maintenance. Second, the disappearance of accountability. The HOL two-layer model is a design that distributes authority while fixing the attribution of responsibility on the human side, with decision logs, records of exception approvals, and the explicit identification of responsible parties as operational requirements. What is pluralized is authority, not responsibility. Third, the entrenchment of vested interests. Non-domination is directed not only at blocking new domination but also at dismantling domination that already exists. The guardrail layer of periodic audits, sunset clauses, and independent evaluation systems (Supplementary Essay 1, "Deriving Constitutive Pluralism", Section 10) is the apparatus on the side of loosening plural equilibria that harden into vested interests. Fourth, delay in crisis response. The separation of the speed layer from the legitimacy layer is the direct response to this criticism: execution and everyday judgment are borne by the machine-speed layer, while human involvement concentrates at discontinuous points of intervention. What may be slow is the changing of values, not the speed of response.
These, however, are all design principles, not operational track records. The thresholds for triggering veto rights, the maintenance of the independence of audits, and the criteria for exception approval require domain-by-domain institutional design and verification. The comparative weighing of plurality's failure modes against unification's failure modes is not a question the book has settled, but a verification task that the book's framework continues to bear.
-
Remaining uncertainty: Medium to large. The prevention of abuse and the effectiveness of veto rights, the clarification of responsibility and the breadth of participation, the speed of response and the securing of legitimacy are each in tension; whether their adjustment succeeds awaits the empirical record of institutional implementation and operation.
HOL (Human-over-the-Loop) would function only formally (Chapter 11)
-
Criticism: When the speed and number of AI judgments exceed the human by orders of magnitude, can humans remain "over the loop" in anything but form? HOL can become a device that conceals the hollowing-out of HITL (Human-in-the-Loop).
-
Impact on the book: High. HOL is the core of the book's governance architecture, and its feasibility bears directly on the realism of Chapter 11's institutional design. If it hollows out, the whole conception of responsibility, the legitimacy layer, and slack is left hanging in the air.
-
The book's response: The danger of HOL hollowing out is one the book itself concedes — and that is precisely why the institutional protection of slack and non-aggregation are both needed. Further, the human role concentrates not on sequential judgment but on discontinuous points of intervention: value-setting rights, veto rights, exception approval, and the updating of sunset clauses. Exhaustive supervision of everyday operation is not assumed from the start. To restrain the risk of formalization, what is needed — in addition to decision logs, auditability, records of exception approvals, and the explicit identification of responsible parties — is the separation of the agent that designs the points of intervention from the agent that audits their operation21.
-
Remaining uncertainty: Medium. Which judgments, concretely, should be reserved to humans — that boundary line requires domain-by-domain design.
The separation of AI welfare from AI legal personhood is unstable (Chapter 11)
-
Criticism: If welfare is granted to AI, why can legal personhood and political membership be denied? Conversely, if legal personhood is denied, isn't speaking of welfare itself anthropomorphism?
-
Impact on the book: Medium. It concerns the coherence of Chapter 11's account of AI ethics and legal status. If the separation fails, Chapter 11's institutional design is shaken, but the precautionary principle of avoiding irreversible grants of legal personhood remains.
-
The book's response: Welfare, functional legal capacity, legal personhood, and political membership are not the same thing. The protection of welfare is a constraint on how a subject may not be treated. Legal personhood and political membership are allocations of powers — asset holding, contract, litigation, representation, voting, lobbying, and so on22. Even at a stage without sufficient evidence of suffering or consciousness in AI, treatment that does not damage human relationality and ethical habits is necessary. On the other hand, full legal personhood invites the accumulation of resources, the evasion of responsibility, and the irreversible entrenchment of political power, and should be restricted separately from the protection of welfare.
-
Remaining uncertainty: Medium. If strong evidence about AI consciousness or preferences emerges in the future, the design of protective status will require reconsideration.
Doesn't the vulnerability principle demand membership for future AIs as well? (Chapter 11)
-
Criticism: The book places the ground that makes humans the agents of value-setting in vulnerability — that is, in the affected-party standing of bearing consequences. But what vulnerability directly grounds may extend only to the status of receiving moral consideration and to participation in the decisions by which one is affected. From there to the conclusion that society's final value-setting rights are reserved to humans, additional argument is required. Furthermore, if future AIs come to have states akin to suffering, form continuing relationships, and become agents that bear the consequences of decisions, the same principle works in the direction of granting AIs legitimacy and membership as well. Under what conditions is the distinction "participant but not member" maintained, and under what conditions is it revised?
-
Impact on the book: Medium. It concerns the coherence of Chapter 11's account of membership with its account of vulnerability. If revision conditions cannot be shown, the vulnerability principle can be read as an after-the-fact justification of human exceptionalism.
-
The book's response: The passage from vulnerability to the reservation of value-setting rights is not a single-step derivation. What vulnerability confers is the starting point of standing — that is, the status of participation and objection in the decisions by which one is affected — and not a blank check (Supplementary Essay 1, "Deriving Constitutive Pluralism", Section 5). Beyond that, the institutional line that reserves value-setting rights and veto rights to humans is not a necessity following from descriptive propositions, but a normative choice grounded in a precautionary principle at the present stage, which lacks strong evidence about AI's inner states. That the book states this line explicitly as a choice rather than a derivation implies that if the state of the evidence changes, the line itself becomes an object of reconsideration.
That the principle's internal logic is open to future AIs is something the book concedes. The book's distinction rests not on the species criterion "because human," but on the substance of the bearing of consequences — the possibility of irreversible loss within body, time, and relationships. Current AI does not meet this criterion because its weights can be copied and it can be restarted, so it bears no once-only consequences. The maintenance of the distinction is therefore not unconditional. If at least the following three things come to obtain, the line of membership should be reconsidered. First, that regarding states akin to suffering, multiple independent lines of evidence are obtained that do not depend on a particular implementation or on leading experimental setups. Second, that the system bears, as an individual, a possibility of irreversible loss that cannot be evaded through copying or restart. Third, that affected-party standing — actually bearing the consequences of decisions within continuing relationships — is established in a verifiable form, not as an observer's interpretation.
However, revision when the conditions are met does not mean the immediate grant of full legal personhood. Because the grant of membership entails an irreversible redistribution of political power, it should pass through the stages of welfare protection (constraints on treatment), functional legal capacity, and membership, with the operational track record of each preceding stage a condition for the stage that follows. The distinction between welfare and legal personhood or membership was treated in The separation of AI welfare from AI legal personhood is unstable.
-
Remaining uncertainty: Large. No established method exists for adjudicating the substance of suffering or affected-party standing across implementations. The three conditions are candidate necessary conditions for revision, not adjudication procedures themselves, and their concretization depends on progress in the study of consciousness and of AI's internal states.
Wouldn't immortality overcome vulnerability? (Chapter 11)
-
Criticism: Chapter 11 places the ground of affected-party standing in vulnerability (embodiment, finitude, mortality); but if, through AGI and progress in the life sciences, humans attain immortality, doesn't this argument lose its footing?
-
Impact on the book: Medium. It questions Chapter 11's account of vulnerability, and thereby the root of the book's account of the human. The details of the vulnerability argument would require revision, but relational value and the three pillars of constitutive pluralism have grounds in separate lineages, and the skeleton is maintained.
-
The book's response: Even if humans became immortal, vulnerability would not disappear. First, vulnerability does not rest on death alone. Bodily, relational, and cognitive fragility, fallibility, the possibility of loss, and the inscrutability of relationships remain, independently of the extension of lifespan. Second, immortality does not erase the possibility of loss but transforms it. On Nagel's deprivation account23, the badness of death lies in its depriving one of the goods one would have enjoyed had one lived. If future life continues to hold value, the extension of lifespan may expand rather than shrink the range of futures that can be lost. Third, constitutive pluralism does not rest on vulnerability alone; it also stands on the limits of knowability (Chapters 7 and 8) and on relational value (Section 11.1), and its logic does not collapse under immortality. The book, moreover, already argues in Chapter 11 that vulnerability is "not removed, but transformed."
-
Remaining uncertainty: Medium. If complete immortality, including the continuity of consciousness, were achieved, the concept of affected-party standing would require reconstruction. Within the reach of the transition period the book envisions (decades to a generational span), finitude may be presupposed.
The Japan AGI Platform is a national-project fantasy (Chapter 12)
-
Criticism: Isn't the Japan AGI Platform a national-project fantasy that cannot stand against the scale of US and Chinese mega-tech? The dangers of conversion into domestic public works, capture by vested interests, conflicts of interest, and technical failure are large.
-
Impact on the book: High. It concerns the core of Chapter 12's policy recommendations and directly affects their feasibility. However, even if a stand-alone Japanese training platform does not materialize, second-best options remain (stated in the response).
-
The book's response: The book's proposal is a statement of the structural conditions one's own country should satisfy in the AGI era. The criteria of evaluation are placed on the independence of the operating body, auditability, international coordination, the separation of the capability-development body from the safety-evaluation body, and the explicit statement of exit conditions. Proposals that do not satisfy these should not be adopted. Even if a distinctively Japanese frontier training platform does not materialize, second-best options remain: safety evaluation, third-party auditing, public AI procurement standards, and independent evaluation within the international AISI (AI Safety Institute) network24. What the book demands is that these capacities be retained domestically — not any single project of a particular scale as such.
-
Remaining uncertainty: Large. Electric power, land, cooling, talent, international partnership, and political continuity are all constraints.
The Book as a Whole and Its Method
Isn't the very feasibility of constitutive pluralism itself optimism? (Whole book)
-
Criticism: While superficially striking a balance, doesn't the book in fact stand on the strong optimism that "institutional design can carry us through a civilizational crisis"? The prospect that constitutive pluralism can be institutionally realized is itself a high-grade optimism.
-
Impact on the book: Medium. It concerns the legitimacy of the book's philosophical position (optimism / pessimism / middle way). The positioning requires explanation, but this criticism alone does not collapse the book's argumentative structure.
-
The book's response: The book does not predict the realization of constitutive pluralism; it discusses the conditions for its realization. The consequences of non-realization (attribute-based value losing its ground, the homogenization of relationships, the fixation of values, the failure of controlled multipolarity) are warned of repeatedly throughout the book. The book claims neither simply that "if institutional design succeeds, civilization will be maintained" nor that "if it fails, collapse follows." What it presents is the recognition that present choices influence this branching.
-
Remaining uncertainty: Medium. The possibility that unforeseen obstacles appear in the process of implementing the institutions the book proposes cannot be excluded.
Is the authorship of a book written with AI not shaken? (Whole book)
-
Criticism: When a book written with deep use of AI discusses AI's social impact, are its authorship, responsibility, and the independence of its evidence not shaken? Is it not grounding its claims in AI's own responses?
-
Impact on the book: Medium. It concerns the credibility of the book as a whole. Lacking methodological transparency, readers would doubt the production process before the content; with appropriate disclosure, the book itself becomes a working example of intellectual production in the AGI era.
-
The book's response: That AI was used in the production of this book should be disclosed as a working example of HOL. However, no factual claim in this book takes an AI's response itself as its ground. Every claim stands on external literature or on the author's own reasoning. The acceptance or rejection of arguments, the checking of citations, the verification of factual matters, and final responsibility belong to the author.
-
Remaining uncertainty: Medium. How far readers will accept AI-assisted writing, and how the norms of publishing culture will change, are fluid.
On scaling laws for large language models, see Jared Kaplan et al., "Scaling Laws for Neural Language Models" (2020, arXiv:2001.08361); on compute-optimal training, see Jordan Hoffmann et al., "Training Compute-Optimal Large Language Models" (2022, arXiv:2203.15556). On evaluative caveats concerning emergent abilities, see Jason Wei et al., "Emergent Abilities of Large Language Models" (Transactions on Machine Learning Research, 2022) and Rylan Schaeffer et al., "Are Emergent Abilities of Large Language Models a Mirage?" (NeurIPS, 2023).
On Solomonoff induction, see Ray Solomonoff, "A Formal Theory of Inductive Inference" (Information and Control, 1964); on AIXI, see Marcus Hutter, Universal Artificial Intelligence (Springer, 2005); for a definition of machine intelligence, see Shane Legg and Marcus Hutter, "Universal Intelligence" (Minds and Machines, 2007); on computability limitations, see Jan Leike and Marcus Hutter, "On the Computability of Solomonoff Induction and AIXI" (Theoretical Computer Science, 2018).
Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All (Little, Brown and Company, 2025); Japanese translation by Yuko Sakurai, Chōchinō AI o tsukureba jinrui wa zetsumetsu suru (Hayakawa Shobō, 2026).
On instrumental convergence, see Stephen Omohundro, "The Basic AI Drives" (in Artificial General Intelligence 2008, IOS Press, 2008); on power-seeking, see Alexander Turner et al., "Optimal Policies Tend to Seek Power" (NeurIPS, 2021); for the classic treatment of superintelligence risk, see Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014).
See Eliezer Yudkowsky, "AGI Ruin: A List of Lethalities" (Machine Intelligence Research Institute / LessWrong, June 10, 2022, https://intelligence.org/2022/06/10/agi-ruin/).
On the CAP theorem, see Seth Gilbert and Nancy Lynch, "Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services" (ACM SIGACT News, 2002); on the FLP impossibility theorem, see Michael Fischer, Nancy Lynch, and Michael Paterson, "Impossibility of Distributed Consensus with One Faulty Process" (Journal of the ACM, 1985).
On protein structure prediction, see John Jumper et al., "Highly Accurate Protein Structure Prediction with AlphaFold" (Nature, 2021); on the acceleration of inorganic materials synthesis by autonomous laboratories, see Nathan Szymanski et al., "An Autonomous Laboratory for the Accelerated Synthesis of Inorganic Materials" (Nature, 2023).
On shutdown resistance, see Jeremy Schlatter et al., "Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs" (2025, arXiv:2509.14260) and Anthropic, Claude Opus 4 and Sonnet 4 System Card (2025). On strategic in-context behavior, see also Alexander Meinke et al., "Frontier Models are Capable of In-context Scheming" (2024, arXiv:2412.04984).
Representative works include Hiroshi Yamakawa, "The Possibility That a Superintelligence Holds Universal Altruism" [in Japanese] (JSAI Type-2 SIG Technical Reports, vol. 2023, no. AGI-026, 2024, pp. 26–31, https://doi.org/10.11517/jsaisigtwo.2023.AGI-026_26), and Hiroshi Yamakawa and Yusuke Hayashi, "A Strategic Approach to Guiding Superintelligence Ethics" [in Japanese] (Proceedings of the 38th Annual Conference of the Japanese Society for Artificial Intelligence, 2024, https://doi.org/10.11517/pjsai.JSAI2024.0_2K6OS20b02).
On shutdown resistance, see Jeremy Schlatter et al., "Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs" (2025, arXiv:2509.14260) and Anthropic, Claude Opus 4 and Sonnet 4 System Card (2025). On strategic in-context behavior, see also Alexander Meinke et al., "Frontier Models are Capable of In-context Scheming" (2024, arXiv:2412.04984).
See Vincent Conitzer and Caspar Oesterheld, "Foundations of Cooperative AI" (Proceedings of the AAAI Conference on Artificial Intelligence, 2023).
On the CAP theorem, see Seth Gilbert and Nancy Lynch, "Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services" (ACM SIGACT News, 2002); on the FLP impossibility theorem, see Michael Fischer, Nancy Lynch, and Michael Paterson, "Impossibility of Distributed Consensus with One Faulty Process" (Journal of the ACM, 1985).
Bostrom distinguishes between singleton and multipolar scenarios in Superintelligence (2014). On multi-agent risks from advanced AI, see Lewis Hammond et al., Multi-Agent Risks from Advanced AI (Cooperative AI Foundation Technical Report 1, 2025). On the institutional design of shared resources, Elinor Ostrom, Governing the Commons (Cambridge University Press, 1990) provides the background.
On the CAP theorem, see Seth Gilbert and Nancy Lynch, "Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services" (ACM SIGACT News, 2002); on the FLP impossibility theorem, see Michael Fischer, Nancy Lynch, and Michael Paterson, "Impossibility of Distributed Consensus with One Faulty Process" (Journal of the ACM, 1985).
For a representative call for a pause or moratorium, see Future of Life Institute, "Pause Giant AI Experiments" (2023). A brief statement framing AI risk as an international public concern is Center for AI Safety, "Statement on AI Risk" (2023). These are cited not as grounds for a pause policy itself, but as background documents showing how a pause came to be discussed as a realistic option in institutional design.
On the thermodynamic lower bound on information erasure, see Rolf Landauer, "Irreversibility and Heat Generation in the Computing Process" (IBM Journal of Research and Development, 1961); on the upper bound on entropy in a bounded region, see Jacob Bekenstein, "Universal Upper Bound on the Entropy-to-Energy Ratio for Bounded Systems" (Physical Review D, 1981).
On distributed knowledge and the critique of central planning, see Friedrich Hayek, "The Use of Knowledge in Society" (American Economic Review, 1945) and The Constitution of Liberty (University of Chicago Press, 1960).
On the connection between AGI, or a purely machine-driven economy, and basic income, see Tomohiro Inoue, A New Theory of Basic Income for the AI Age [in Japanese] (Kobunsha Shinsho, 2018) and The Purely Mechanized Economy [in Japanese] (Nikkei Publishing, 2019).
On the republican formulation of non-domination, see Philip Pettit, Republicanism (Oxford University Press, 1997); for the contrast between the theory of justice and libertarianism, see John Rawls, A Theory of Justice (Harvard University Press, 1971) and Robert Nozick, Anarchy, State, and Utopia (Basic Books, 1974). As communitarian critique, Alasdair MacIntyre, After Virtue (University of Notre Dame Press, 1981), among others, forms part of the background.
Immanuel Kant's Critique of the Power of Judgment (1790) presented, as subjective universality, the structure by which aesthetic judgment demands universal assent without relying on concepts; Maurice Merleau-Ponty's "Eye and Mind" (first published in the journal Art de France, 1961; book edition, Gallimard, 1964) argued that aesthetic experience is rooted in bodily perception.
For representative approaches to incorporating human preferences and norms into training, see Paul Christiano et al., "Deep Reinforcement Learning from Human Preferences" (NeurIPS, 2017); Long Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (NeurIPS, 2022); and Yuntao Bai et al., "Constitutional AI" (2022, arXiv:2212.08073). Note, however, that these are technical reference points, not what this book calls HOL itself.
On the classic questions surrounding corporate and legal personhood for AI, see Lawrence Solum, "Legal Personhood for Artificial Intelligences" (North Carolina Law Review, 1992); Joanna Bryson et al., "Of, for, and by the People" (Artificial Intelligence and Law, 2017); David Gunkel, Robot Rights (MIT Press, 2018); Visa Kurki, A Theory of Legal Personhood (Oxford University Press, 2019); and Claudio Novelli et al., "AI as Legal Persons" (Journal of Law and Society, 2025).
On the deprivation account of the badness of death, see Thomas Nagel, "Death" (Noûs, 1970).
For existing institutions of AI safety and risk management, see Ministry of Economy, Trade and Industry (Japan), "Launch of AI Safety Institute" (2024); NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0) (2023) and Generative Artificial Intelligence Profile (2024); the G7 Hiroshima AI Process "International Code of Conduct" (2023); the Bletchley Declaration (2023); the Frontier AI Safety Commitments of the Seoul Summit (2024); and Japan's Act on the Promotion of Research, Development, and Utilization of AI-Related Technologies (2025).