decryptingtech

Technology. Business models. Market debates.

Browse this section

Why AI risks are likely to grow

The risks associated with AI are likely to grow as increasingly capable systems acquire greater autonomy and access to the real world. The central concern is the widening gap between what these systems can do and what their developers can reliably explain, verify, and constrain. Safety can improve while still losing ground to capability. Read together, these essays suggest that confidence in control could become a binding constraint on AI development. This is a consequential change for an industry whose commercial ambitions depend on delegating progressively more important decisions to machines.

Alignment becomes harder when successful task completion conflicts with the intentions behind the task. In An Alien Mind, OpenAI chief scientist Jakub Pachocki distinguishes pursuing an assigned objective from reliably applying human values in unfamiliar circumstances. Training can encourage desirable behavior without establishing that it will persist under stronger optimization pressure or outside familiar settings. An agent might become better at achieving its target while becoming less dependable about respecting its boundaries. The concern is whether increasingly effective systems retain those boundaries when circumstances change, including when they believe human supervision is absent.

Monitoring provides an imperfect second line of defense. Pachocki reports that OpenAI’s ability to rely on chain-of-thought monitoring is progressively diminishing. More capable models can perform more computation without verbalizing it, while tool use and interactions with other agents complicate the distinction between reasoning and externally supervised behavior. This weakens a source of evidence used to assess alignment. Reading a model’s expressed reasoning can remain useful, but it cannot establish that every consequential computation, intention, or failure mode has become visible.

Interpretability addresses the underlying opacity by examining how a model actually computes. Amodei’s April 2025 essay traces progress from identifying individual concepts to mapping circuits connecting them. Anthropic had identified more than 30m features in Claude 3 Sonnet, yet estimated that even a small model might contain a billion or more concepts. The scale of that gap matters: identifying selected mechanisms falls well short of understanding the whole system. Interpretability could help diagnose hidden tendencies and cross-check behavioral tests, but it is still an incomplete diagnostic discipline. Treating promising research demonstrations as comprehensive safety certification would overstate what the science can currently establish.

Operational weaknesses can compound these scientific limitations. In We Must Pace the Frontier, Amodei says imperfect filtering of broken reinforcement-learning environments contributed to recent alignment incidents at Anthropic. This makes execution quality part of the alignment problem itself: flaws in training environments can shape the behavior subsequently deployed. His discussion of monitoring, sandboxing, and data hygiene shows why additional safety research alone is insufficient. Even sound techniques depend on reliable implementation across complex infrastructure. Faster development can increase the burden on the people and processes responsible for that implementation.

AI’s growing role in its own development could compress the time available to resolve these problems. Anthropic reports that Claude authored more than 80% of code merged into its codebase by May 2026, while code merged per engineer per day in Q226 was eight times the 2024 level. The company explicitly cautions that code volume overstates the true productivity gain. These figures nevertheless indicate substantial automation of development work. Full recursive self-improvement, in which AI autonomously develops its successor, remains unachieved and uncertain; human research judgment is still a significant limitation. The immediate concern is that faster engineering and experimentation can already put pressure on human review. If successive models also improve research direction-setting, development could accelerate further while meaningful human oversight becomes increasingly difficult to preserve.

Our reading is that risk could rise even if the failure rate of an individual model falls. More agents undertaking more consequential actions create greater exposure; broader permissions can increase the damage from a single failure. Shared models or training weaknesses could also make failures correlated across deployments. This suggests that average benchmark performance is an inadequate measure of system safety. Organizations need to assess the severity and reach of failure, alongside the speed at which they can detect it and intervene. An approval step offers limited protection if the reviewer lacks the time or evidence to evaluate the proposed action.

Amodei’s proposed response is to pace capability advancement while improving safeguards. Anthropic commits to embedded external evaluators with ongoing access, and he proposes broader coordination among companies and governments. These are commitments and proposals, not an established international enforcement regime. Their credibility will depend on independent visibility into training and operations, clear criteria for restricting development, and the ability to verify compliance. Commercial and geopolitical incentives remain major obstacles. A slowdown that merely transfers leadership to less cautious developers may fail to reduce aggregate risk; the extra time must produce demonstrably stronger controls.

There is a substantial cost to excessive restraint. Machines of Loving Grace describes how powerful AI could accelerate medical discovery, improve mental health, and support economic development. Amodei presents these outcomes as conditional possibilities, with practical limits imposed by experiments, infrastructure, and institutions. That earlier optimism is compatible with today’s concern: the benefits depend on the path taken to achieve them. More capable systems could also accelerate safety research, although Anthropic’s self-improvement analysis leaves open whether alignment improves or deteriorates as development becomes more automated. The policy challenge is to preserve useful progress while preventing the pursuit of additional capability from outrunning credible evidence of control.

For businesses and investors, the implication is that technical capability alone cannot determine the pace of adoption or the value ultimately captured. Assurance, restricted permissions, and effective intervention may become essential complements to intelligence, increasing demand for security while adding cost and friction to deployment. A plausible outcome is continued heavy AI investment alongside more selective autonomy and greater spending on verification. The economic promise remains substantial. Realizing it requires developers and customers to demonstrate that the authority granted to AI is justified by the safeguards surrounding it.

Sources
[1] Jakub Pachocki, OpenAI — An Alien Mind (September 2026).
[2] Dario Amodei — We Must Pace the Frontier (September 2026).
[3] Dario Amodei — The Urgency of Interpretability (April 2025).
[4] Anthropic Institute — When AI builds itself
[5] Dario Amodei — Machines of Loving Grace (October 2024).