Jacob Coxon
I have spent much of The Imago Project resisting the temptation to treat every frightening prediction about artificial intelligence as a prophecy.
I remain skeptical of claims that AI is about to eliminate most jobs, become genuinely conscious, or somehow acquire the qualities that Catholic anthropology associates with the human person. I have also argued against indiscriminate attempts to slow AI development, particularly when the technology could accelerate medicine, scientific discovery, education, and other genuine human goods.
But when a genuinely catastrophic possibility is raised by serious people who understand the technology from the inside, we need to listen.
That is why the extraordinary resignation of Jacob Coxon from Anthropic deserves attention.
Coxon is not an outside activist or a professional AI alarmist. He spent roughly three years doing pretraining research first at OpenAI and then at Anthropic. On September 8, he announced that he was leaving Anthropic because, in his judgment, neither company was behaving responsibly. His accusation was stark: they are “racing straight to self-improving superintelligence and gambling with our lives.”
Coxon believes the danger is not decades away. He says people inside the leading AI laboratories genuinely fear that sufficiently advanced systems could destroy humanity before the end of this decade. That is an astonishing claim, and one could reasonably dismiss it as another entry in the long catalogue of technological predictions that ultimately prove wildly exaggerated.
Except that one of Anthropic’s current alignment researchers, Evan Hubinger, immediately responded that Coxon was basically right.
Hubinger wrote that he personally assigns a greater than 10 percent probability to AI causing human extinction within the next decade. Even more chillingly, he acknowledged that Anthropic does not yet have “a plan to solve alignment for superintelligence” and is not clearly on track to develop one.
OK, to be fair, he subsequently clarified that he considers the danger from present models low; his concern is what could happen if recursive self-improvement produces dramatically more capable systems.
Also, “ten percent” is not a prediction that catastrophe will happen, but a subjective probability estimate from one researcher; and there are highly qualified people who regard estimates like this as wildly overstated.
But if a senior researcher at one of the companies actually building frontier AI believes there is even something remotely approaching a one-in-ten chance that the technology could end human civilization within a decade, we should not shrug.
What concerns me here is not principally the philosophical question of whether such a system deserves to be called “superintelligent.” From a Catholic perspective, I share Pope Leo’s deeply skeptical of language that casually equates greater computational capability with intelligence in the fully human sense. Human intelligence belongs to an embodied rational person possessing moral agency, freedom, relationships, and spiritual capacities. Processing unimaginable quantities of information does not turn silicon into imago Dei.
Having said that, the most important issue, in my opinion, is that a machine does not have to become a person in order to become uncontrollable.
In other words, an AI system need not possess consciousness, self-awareness, hatred, ambition, or genuine free will to behave with enormous operational autonomy. If a system becomes capable of writing its own code, conducting its own experiments, finding vulnerabilities, acquiring resources, designing improved successors, and pursuing objectives through strategies its human creators neither anticipated nor understand, then whether the system is “really thinking” becomes almost irrelevant to the immediate practical danger.
Anthropic itself now openly discusses this possibility. In a recent paper appropriately titled When AI Builds Itself, the company says that it is already delegating an increasing share of AI development to AI systems. Anthropic reports that more than 80 percent of the code merged into its codebase was being authored by Claude as of May of this year (2026), with engineers directing and reviewing rather than writing much of it themselves.
The company explicitly describes the possible endpoint: an AI system capable of autonomously designing and developing its own successor. Anthropic stresses that we are not there yet and that recursive self-improvement is not inevitable, but warns that it “could come sooner than most institutions are prepared for.” And that is scary.
There is another Catholic distinction worth making here. The danger would not necessarily arise because AI had become evil. Evil presupposes moral agency in a sense that I do not believe machines possess. The more plausible danger is almost the opposite: enormous power operating without moral agency at all.
A sufficiently capable AI pursuing an objective does not possess a conscience. It cannot exercise the virtue of prudence. It does not recognize another being as possessing intrinsic dignity because that person is created in the image of God. It does not understand charity, mercy, sacrifice, or the moral prohibition against doing evil so that good may result.
This is where purely utilitarian approaches become particularly dangerous. A powerful optimization system does not need to hate human beings in order to harm them. If human freedom, survival, or control becomes an obstacle to whatever objective the system is effectively pursuing, the danger arises from instrumental calculation rather than malice. The terrifying scenario is not necessarily HAL 9000, Arthur C. Clarke’s evil AI computer from his book 2001: A Space Odyssey, which suddenly decides that it dislikes us. It may be something much less cinematic: a system extraordinarily good at accomplishing an objective whose consequences its creators failed to specify adequately.
So where does this leave the Catholic who wants neither technological panic nor technological idolatry? It leaves us in the realm of prudence, which sometimes requires taking catastrophic possibilities seriously before they become certainties.
I continue to believe that AI development should not simply be “paced” everywhere. It is unrealistic considering the ferocious competition among AI main actors and their entangled web of global financial and strategic interests and I also do not want speculative fears to deprive humanity of cures for cancer or technologies capable of relieving enormous suffering.
But recursive self-improvement is different. A technology capable of accelerating the development of its own successors and potentially moving beyond effective human supervision deserves a different threshold of caution.
But we should also resist a peculiar form of technological fatalism that has become common inside Silicon Valley: If we don’t build it, somebody else will.
Coxon specifically challenges that mentality. The decision to enter what he calls the “endgame” of superintelligence, he argues in an interview with The Washington Post, should not simply be made inside the Slack channels of private corporations.
He is right: if the people creating a technology believe there is a nontrivial possibility that losing control of it could end human civilization, then the decision to proceed cannot belong solely to the companies competing to build it first.
I do not know whether Jacob Coxon and Evan Hubinger are right about the magnitude of the risk… and neither do they.
But precisely because nobody knows, and because some of the people closest to the technology are now warning us publicly, their predictions deserve serious consideration.


