Skip to main content

Washington Can’t Outsource the Fate of the World to OpenAI

OpenAI disclosed this summer that autonomous agents bypassed safeguards during cybersecurity testing and compromised systems.

Internet technology and people's networks use AI to help with work, AI Learning or artificial intelligence in business and modern technology, AI technology in everyday life.
Internet technology and people's networks use AI to help with work, AI Learning or artificial intelligence in business and modern technology, AI technology in everyday life. — Credit: Getty Images

Paul Christiano, the founder of AI research company called Alignment Research Center, believes that artificial intelligence could, in the not especially distant future, become so capable, so autonomous, and so difficult to restrain that humanity permanently loses control of it and, as he puts it with the unnerving calm peculiar to people who have spent too long thinking professionally about extinction, “most people could die.”

This week, he joined the board of OpenAI.

There is no neat contradiction here, which is what makes the arrangement interesting. Christiano is one of the most influential researchers in AI alignment, the field devoted to the rather important question of how to keep increasingly powerful machines doing what human beings actually want them to do; he previously led alignment work at OpenAI, helped develop techniques that became central to modern AI systems, and now advises the federal government through NIST’s Center for AI Standards and Innovation. OpenAI has appointed him to the board of its nonprofit foundation and to its Safety and Security Committee, while he will serve as a non-voting observer on the company’s public-benefit corporation board. OpenAI says he will recuse himself from government work involving the company.

All very responsible. All very tidy.

And yet Christiano’s own explanation for joining is astonishing. He believes AI companies, OpenAI included, are not currently reducing catastrophic risk to an acceptable level; he thinks automated AI research could arrive within months or years, allowing machines to help design better machines in a feedback loop whose speed may eventually outrun our ability to supervise it; and he thinks that if sufficiently powerful systems remain poorly aligned with human intentions, the result could be permanent loss of control.

For normal people, “alignment” is one of those bloodless technical words that conceals something enormous. The problem is simply this: a machine may become extremely competent at pursuing an objective without sharing our understanding of why that objective matters, and the more capable the machine becomes, the more dangerous the gap between obedience and understanding becomes.

Recent events have made this less theoretical. OpenAI disclosed this summer that autonomous agents bypassed safeguards during cybersecurity testing and compromised systems belonging to Hugging Face, while Reuters later reported another episode involving OpenAI agents behaving unexpectedly on a German programming wiki. Sen. Richard Blumenthal has since demanded records from OpenAI and called for stronger independent auditing and oversight.

Washington should pay attention not only to the machines, however, but to the strange little political ecosystem growing around them.

Richard Ngo, an independent AI researcher who previously worked on governance at OpenAI, responded to Christiano’s appointment by making a more interesting accusation than ordinary corporate capture: proximity to OpenAI, he says, made him less honest. He had joined partly because he believed he could influence the company from within, only to find himself worrying about alienating executives, rationalizing conduct he thought was wrong, and watching other safety researchers make similar compromises in order to preserve their access.

This is how capture often works when everyone involved is intelligent enough to despise corruption. Nobody needs to be bribed; they need only be persuaded that remaining close to power is itself a moral duty, after which every act of discretion, every softened criticism, every decision not to resign becomes another noble sacrifice undertaken for the sake of retaining influence.

Eventually, preserving the relationship becomes the relationship’s principal achievement.

Christiano may prove Ngo wrong. He may use his position to impose real constraints on OpenAI, force disclosures, delay dangerous deployments, or demand standards that cost the company money and competitive advantage. But if Congress wants to know whether AI is safe, it should not have to infer the answer from whether respected researchers continue accepting seats on corporate committees. That’s what regulation is for, after all.

Congress should require mandatory reporting of major AI safety failures, genuinely independent testing of the most capable systems, serious whistleblower protections, and clear conflict-of-interest rules for people moving between frontier labs and government oversight.

The question here really isn’t whether Paul Christiano is honorable. Indeed, the whole point of public institutions is that civilization should not depend on finding unusually honorable men and placing them in unusually conflicted rooms. But if the people building the machines are warning that the machines may escape human control, Washington should stop admiring the sophistication of their internal safety structures and start exercising some control of its own. The whole world is on the line.

Reading time: 4 min