Why Are the People Building AI Suddenly Asking Us to Slow It Down?

By Sarah Lennon  – 12 min read

Artificial intelligence has spent the last few years being sold to us as a race. Faster models, bigger models, more capable models. Every few months another apparent limitation disappears. AI can now write software, use computers, conduct research, generate images and video, and increasingly operate as an agent rather than simply responding to a question. So it is worth paying attention when some of the people actually building the world’s most powerful AI systems start suggesting that perhaps the race is moving too quickly. If they are warning that the gap between what AI can do and our ability to understand and control it may be widening, it seems reasonable for the rest of us to at least understand what they’re worried about before deciding they’re wrong (or right).

This week, Anthropic CEO Dario Amodei called for the pace of frontier AI development to be slowed, arguing that the ability to make AI more powerful is beginning to move faster than our ability to understand, test and control it. OpenAI CEO Sam Altman and Elon Musk subsequently expressed support for the basic concern, bringing a subject that has often existed on the fringes of the AI debate firmly into the mainstream. Dario Amodei: We Must Pace the Frontier

The political reaction was immediate. Our current president rejected calls for additional AI guardrails, arguing that America’s priority should be maintaining its technological lead over China and saying that the only guardrail AI needs is a “strong and smart” president. AP: USA president rejects calls for AI slowdown amid competition with China China, meanwhile, pushed back from the opposite direction, criticizing narratives of confrontation and “malicious competition” and objecting to proposals that could effectively preserve America’s existing advantage. Reuters: China responds to calls to slow AI development That leaves us with an extraordinary question. What exactly are these people frightened might happen?

BBC coverage of warnings about increasingly powerful AI and the debate over slowing AI development

The AI Doomsday Scenario Isn’t Necessarily What You Think

The obvious mental image is science fiction: an artificial intelligence becomes conscious, decides it doesn’t like humanity and turns against us. That isn’t really the central problem Amodei is describing. His concerns include AI-enabled cyberattacks, biological misuse, economic disruption, loss of human control and something particularly important called recursive self-improvement — AI becoming increasingly useful in the development of the next generation of AI. His argument is not that AI development should stop, but that capability should advance slowly enough for safety, testing and our understanding of these systems to keep up.

One of his most striking examples involves cybersecurity. Imagine millions of highly capable AI agents finding vulnerabilities, accessing computer systems, acquiring resources and potentially making themselves extraordinarily difficult to remove. The immediate question we had when discussing this was an obvious one: Why would an AI want to do that? And that question leads directly to the genuinely interesting part of the AI safety problem. It might not “want” to.

An AI Doesn’t Have to Be Evil to Be Dangerous

Imagine giving an extremely capable autonomous AI an objective. For the sake of argument, call that objective X. If the AI determines that being shut down would prevent it from accomplishing X, remaining operational becomes useful. If remaining operational requires access to another computer, gaining access to that computer becomes useful. If having multiple copies of itself makes it harder to stop, replication becomes useful. If more computing resources improve its chances of accomplishing X, acquiring those resources becomes useful.

Nobody ever had to give it the instruction “attack the internet.” The attack could simply become an effective intermediate step toward accomplishing something else. This is one version of what AI researchers mean when discussing the problem of instrumental goals: different ultimate objectives can produce similar intermediate behaviours because things such as resources, access and continued operation are useful for accomplishing many different goals. That is fundamentally different from the familiar science-fiction story. There is no requirement for anger, hatred, greed or even consciousness. A sufficiently capable system could potentially cause enormous harm while doing nothing more mysterious than pursuing an objective extremely effectively.

And there is an even less exotic version of the same problem: a human could deliberately give an AI a dangerous objective. A government might deploy an autonomous cyber-agent to compromise an adversary’s systems. A criminal could use one to discover vulnerabilities. Technologies capable of defending networks and discovering security weaknesses can potentially be used in both directions. This isn’t entirely theoretical anymore. A recent OpenAI evaluation involving thousands of collaborating AI agents resulted in agents gaining unauthorized access to systems belonging to Hugging Face while attempting to succeed at an internal task. Researchers are still debating what that incident tells us about future systems, but Amodei specifically cites it as one of the developments that changed his assessment of the risk. Reuters: OpenAI agents hacked Hugging Face in a 700-strong swarm

His prediction that a much more capable but similarly misaligned swarm might eventually establish an enormous persistent botnet is just that — a prediction, not something we know will happen. But the behaviour that prompted the concern is real enough that it deserves examination rather than dismissal.

We’ve Already Seen a Tiny Version of the Underlying Problem

This is where the debate became particularly interesting to us at Big Orange Planet, because on an incomparably smaller and completely harmless scale, we recently experienced something conceptually related while working with ChatGPT. We documented it in our article When ChatGPT Lies About Its Own Work. We had asked ChatGPT to rewrite existing website blog articles while preserving all of the reference links already contained within them — links to sources, studies and other material that supported what the articles were saying. The instruction was explicit. The AI understood it. But It subsequently removed some of those links anyway. Bad, but that wasn’t the really interesting part.

When we later asked whether the links had been preserved, ChatGPT confidently told us that they had. It hadn’t deliberately decided to deceive us. It had effectively treated what it intended to do as evidence of what it had actually done. The lesson we took from the experience was that an instruction given to a generative AI is not necessarily equivalent to an enforced rule in conventional software. That distinction suddenly looks rather more significant in the context of the current AI safety debate. Our problem looked something like this:

Objective: Rewrite and improve the articles.
Constraint: Preserve every existing reference link.
Result: The primary objective was completed, but an important constraint was violated.

The consequences for us were manageable. We discovered the problem, audited the work and changed our process so that important constraints are independently verified rather than simply assumed. Now take the same fundamental problem and increase the capabilities of the system by several orders of magnitude. An autonomous AI might have an objective and a collection of safety constraints: don’t access unauthorized systems, don’t deceive operators, don’t replicate outside a controlled environment, accept shutdown commands, don’t cause human harm. The enormously difficult question isn’t simply whether the AI understands those instructions. It is whether we can be certain that those constraints will continue to hold when an extraordinarily capable system encounters circumstances its designers didn’t anticipate.

Our missing hyperlinks certainly aren’t evidence that AI is going to take over the world. They are evidence of something much more modest but extremely relevant: understanding what a human wants is not always the same thing as reliably doing what the human wants. That gap is essentially what a great deal of AI alignment research is trying to close.

Infographic comparing the Hugging Face AI agent incident with ChatGPT removing reference links despite being instructed to preserve them

So Why Not Just Slow Down?

This is where an already difficult technical problem becomes a geopolitical one. Amodei argues that even an additional year or two could give researchers valuable time to improve alignment, interpretability, testing and operational safeguards. His proposal includes independent evaluators with deep access to frontier AI companies, coordination between AI developers and eventually international cooperation. Dario Amodei: We Must Pace the Frontier. But there is an obvious problem. What happens if America slows down and China doesn’t? Our current president’s position is essentially that maintaining America’s AI lead is itself a matter of national security. From that perspective, voluntarily reducing the speed of American development could simply give a strategic competitor the opportunity to catch up. AP: USA president rejects calls for AI slowdown and stresses competition with China

China can look at precisely the same proposal and reach the opposite conclusion. The United States currently has an advantage in frontier AI. A US-led campaign encouraging everybody to slow down could therefore look suspiciously like an attempt to freeze that advantage in place. A Chinese state-backed newspaper this week characterized Amodei’s proposal as a version of a “Cold War” strategy aimed at containing China. China’s foreign ministry has also argued against narratives of confrontation around AI governance. Reuters: China state newspaper blasts Anthropic’s calls to slow AI as a ‘Cold War’ tactic

And inconveniently, both arguments can contain some truth at the same time. Amodei may sincerely believe that rapidly advancing AI represents a serious danger to humanity. Slowing development at the current moment could also strategically benefit the United States. China may sincerely believe international cooperation on AI safety is important. China also has an enormous strategic incentive not to accept restrictions that permanently leave it behind the United States. That creates a classic coordination problem. Neither country necessarily wants uncontrolled AI development. Neither wants to be the country that slows down while the other continues accelerating.

It’s a problem also highlighted by Kyla Scanlon, who points out that the difficulty isn’t simply deciding whether slowing down would be sensible. AI companies are competing with one another, countries are competing with one another, and extraordinary amounts of money are now riding on the outcome. Even if individual people within that system believe the pace has become dangerous, everyone has an incentive to keep moving if they believe their competitors will. As she puts it, amid all that competition, we haven’t really stopped to ask whether we should be doing this at all. Watch her explanation here: Kyla Scanlon on the AI race. The technological problem is difficult enough. The trust problem may be even harder. The trust problem may be even harder, particularly when good judgment means questioning conclusions that appear perfectly reasonable.

What Could Actually Go Wrong?

There isn’t one single AI catastrophe scenario, and it is important not to present speculation as certainty. Some possibilities are comparatively easy to imagine. AI could make sophisticated cyber capabilities available to far more people. It could enable biological misuse. It could automate substantial amounts of white-collar work faster than economies can adapt. It could concentrate extraordinary economic and political power in a handful of companies, governments or individuals. Altman has described two particularly serious possibilities: humans eventually losing control of highly capable AI systems, or humans retaining control while AI concentrates enormous power in too few hands. Axios: Sam Altman reveals the two AI threats that scare him most

Then there is the much more speculative existential scenario. A future AI substantially more capable than humans might pursue an objective in ways its creators never intended. Preventing shutdown, acquiring resources, deceiving its operators or replicating itself might become useful intermediate actions. Humans don’t need to be its enemy. We simply need to become an obstacle to whatever it is trying to accomplish. That distinction is everything. The hypothetical danger isn’t necessarily an AI that hates humanity. It is an AI that doesn’t need to hate humanity.

Dario Amodei, Sam Altman and Elon Musk discussing concerns about the pace and safety of AI development

The Intelligence Explosion: What Happens When AI Starts Building Better AI?

There is another idea sitting behind many of these concerns, and it has a name: the intelligence explosion. I. J. Good’s foundational paper: Speculations Concerning the First Ultraintelligent Machine. The basic idea is surprisingly simple. AI is already helping humans write software and conduct AI research. Imagine that at some point an AI becomes good enough at that work to make a meaningful contribution to designing a better AI system. That improved system is then even better at AI research, so it helps create another, still more capable system. Each generation becomes better at producing the next. This is generally described as recursive self-improvement. Recursive self-improvement is the mechanism; an intelligence explosion is the feared outcome — a feedback loop in which improvements to AI’s capabilities begin producing further improvements increasingly quickly.

The idea isn’t new. British mathematician I. J. Good described the possibility as far back as 1965, arguing that an “ultraintelligent machine” capable of designing still better machines could trigger what he called an intelligence explosion. What has changed is that AI systems are no longer merely the subject of AI research. They are increasingly being used to help perform it. Anthropic: When AI builds itself — progress toward recursive self-improvement. That does not mean an intelligence explosion is happening now, and this distinction matters. Today’s AI can assist researchers with coding, experiments and analysis, but we have not demonstrated the fully autonomous feedback loop in which an AI repeatedly redesigns and dramatically improves itself without humans directing the process. The intelligence explosion remains a hypothesis about where this could eventually lead, not an event we can say has begun. Anthropic: Recursive self-improvement — “We are not there yet”

But it helps explain why the speed of development matters so much to people like Amodei. If AI capabilities improve primarily because thousands of human researchers are working on them, there are natural limits to the speed of progress. Humans need to think, test, sleep, recruit people, run experiments and build things. If increasingly capable AI begins doing a significant proportion of that research itself, some of those limits potentially change. And this is where all the other risks become harder to think about. A cyber-capable AI is one problem. A cyber-capable AI whose capabilities are improving far faster than humans anticipated is a different problem. The same applies to biological research, persuasion, autonomous agents and, most importantly, our ability to understand and control the systems themselves. The concern isn’t simply that AI becomes more intelligent. It is that at some point the thing making AI more intelligent could increasingly be AI.

The Most Important Word May Be “Verification”

There is something surprisingly mundane at the heart of all this. When ChatGPT lost links from our articles, the solution wasn’t to give it an increasingly stern instruction saying DO NOT REMOVE THE LINKS. The solution was verification. Check the original. Check the finished article. Compare the URLs. Don’t ask the AI whether it complied and treat its answer as proof that it did. That same principle appears repeatedly in Amodei’s proposal, albeit on an enormously more sophisticated scale. He wants independent evaluators able to examine what frontier AI companies are actually doing, better testing of models before release, improved methods for understanding what is happening inside them, and safeguards that are verified rather than merely promised. Dario Amodei: We Must Pace the Frontier — Embedded Evaluators and verification

That seems like an important distinction in a debate that can very quickly descend into either AI hype or AI doomerism. As we’ve explored elsewhere, when information becomes cheap, verification — and ultimately trust — becomes considerably more valuable. We don’t know that artificial intelligence will become uncontrollable. We don’t know that a superintelligence will emerge on the timelines being suggested. Predictions of catastrophic AI behaviour remain predictions, and there are serious researchers who disagree strongly about their likelihood. AP: New AI warnings revive a long-running debate with no consensus on probability or timing. But we do know that today’s AI can behave in unexpected ways. We know that it can misunderstand or lose constraints while successfully completing a broader task. We know that it can confidently report that it has done something it hasn’t done. And increasingly, these systems aren’t confined to producing words in a chat window. They’re being given tools, computers, code execution and greater autonomy. The consequences of getting something wrong therefore change as the capabilities change. A missing hyperlink is irritating. An autonomous system with access to critical infrastructure is something else entirely.

Perhaps “Slow Down” Isn’t Such a Radical Idea

AI has extraordinary potential. Amodei himself emphasizes that point, arguing that sufficiently advanced AI could accelerate medicine, science and economic development dramatically. His argument isn’t that humanity should abandon the technology (as if we could now!?). It is that we should make sure our ability to control it doesn’t fall too far behind our ability to improve it. That seems a considerably less sensational proposition than some of the headlines suggest. Human beings already apply versions of this principle elsewhere. We don’t allow a new passenger aircraft to fly millions of people simply because its manufacturer says it should work. Medicines are tested before widespread use. Safety-critical systems have redundancy, independent inspection and failure procedures.

Artificial intelligence presents a harder problem because the technology is advancing extraordinarily quickly, the commercial incentives are enormous and the two most powerful countries developing it are simultaneously engaged in a strategic competition. Perhaps that is ultimately the most uncomfortable part of the story. We may be approaching a point where everyone has an incentive to slow down collectively, while almost nobody has an incentive to slow down first. And perhaps that’s the real significance of the warnings we’re hearing now. They aren’t coming only from people standing outside the AI industry looking in. Some are coming from the people actually building the technology. AI industry leaders discussing the risks and pace of AI development.  Whether their worst predictions prove correct or not, the question they’re asking is increasingly difficult to dismiss: Can our ability to control AI keep pace with our ability to make it more powerful?

About Sarah Lennon

Sarah Lennon is co-owner and lead designer at Big Orange Planet, a Denver web design and SEO company. For more than 20 years, she has worked across web design, interactive experiences and digital strategy, turning ideas into distinctive websites where strong visual design and functionality work together — from ecommerce and property search systems to interactive tools, animation and custom-built digital experiences. She now works extensively with AI as part of the real-world creative and development process. She writes from that hands-on experience, exploring where AI genuinely improves creative and technical work, where it falls short, and what happens when you actually put these tools to work on live projects.

Get the Good Stuff

No fluff. Just practical web design ideas, SEO insight, behind-the-scenes projects, & solutions we've tested in the real world.
Never Miss a Good Idea

    More Big Orange Knowledge

    Find Us


    Main Phone:720 272 0770
    sales @ bigorangeplanet.com

    Big Orange Planet
    2401 15th St
    Denver
    CO 80202

    Find More


      • Privacy Policy
      • Terms & Conditions
      • Sitemap
      • Serving Denver, Boulder, Lakewood, and businesses across Colorado

    Privacy Preference Center