# The Two Prophets

LLMS index: [llms.txt](/en/llms.txt)

---

In May 2023, Altman arrived in Washington, DC, to testify before Congress. It was a remarkable performance. He reiterated the promise of AGI solving climate change and curing cancer, gave a compelling argument for why OpenAI’s technologies would improve and create “fantastic” new jobs, dodged questions about copyright issues and the lack of transparency and privacy guarantees around its training data, and delivered a sincere call for regulation— that is, regulation with OpenAI’s blessing, evoking the specter of China to urge lawmakers not to slow down its innovation.

Senators loved him. In a telling exchange that captured Altman’s nimble rhetoric and the trust and enthusiasm he was garnering, he offered three policy recommendations that shifted the conversation away from existing issues like labor, environment, and intellectual property and toward regulating future AI systems and extreme risks: First, create an agency that would develop and administer a licensing regime for models above a certain threshold of capabilities; second, create a set of AI safety standards for measuring “dangerous” capabilities; third, require independent audits on those standards to check for compliance. He later elaborated that capability thresholds could be approximated with compute thresholds if necessary, and that “dangerous capabilities” could include a model’s ability to manipulate and persuade, and to generate recipes for novel biological agents.

“Would you be qualified to, if we promulgated those rules, to administer those rules?” Louisiana senator John Kennedy asked.

“I love my current job,” Altman said to laughter. “Are there people out there that would be qualified?” Kennedy asked.

“We’d be happy to send you recommendations for people out there, yes.” “Okay. You make a lot of money, do you?” “I make—no. I’m paid enough for health insurance. I have no equity in OpenAI.”

Sitting next to Gary Marcus, Altman won over even one of his most vocal critics. “Let me just add for the record that I’m sitting next to Sam, closer than I’ve ever sat to him except once before in my life,” Marcus said, “and his sincerity…is very apparent physically in a way that just doesn’t communicate on a television screen.” (Marcus would later backtrack his rare show of approval: “I realized that I, the Senate, and ultimately the American people, had probably been played.”)

Altman’s prep team considered it a resounding success. The hearing was the cherry on top of a long campaign. After the launch of ChatGPT, nearly everyone in Washington had desperately sought meetings with OpenAI. The small policy team under Anna Makanju, after operating in relative obscurity, had received an avalanche of requests. For months, with or without Altman, they had been dining with, giving demos to, fielding questions from, and delivering the legislative proposal that Altman gave during his testimony to as many policymakers as possible—from senators, House representatives, staffers, and cabinet members to visiting diplomats, agency heads, and Vice President Kamala Harris—morning to night, practically nonstop.

By early June, Altman had personally met with at least one hundred US lawmakers, according to The New York Times, some of whom proudly referenced those private conversations during the hearing.

On the day of Altman’s testimony, a small band of Hollywood concept artists, who specialize in the conceptual design of characters, props, and other visual elements in movies, had also been scheduled to meet with several congressional offices. They had crowdsourced funding for their airfare and accommodation and had in the process been threatened by online trolls and doxed for speaking out against the AI industry. Just as Hollywood writers were —and soon Hollywood actors would be—in the midst of historic strikes, to bargain in part for better protections against AI, the artists, too, had planned to speak candidly about the devastating effects that generative AI was already having on their profession. Generative AI developers had trained on millions of artists’ work without their consent in order to produce billion-dollar businesses and products that now effectively replaced them. Those jobs that were being erased were solid middle-class jobs—as many as hundreds of thousands of them. “Artists are in so much pain right now. No one is getting work,” says Karla Ortiz, a concept artist known for her work on Marvel Studios’ Doctor Strange, who was part of the group and filed the first artist lawsuit against several generative AI companies.

As they arrived in Washington, several of their meetings were bumped by Altman’s testimony to the following day, scrambling their schedules and leaving them to walk around the halls of the hearing—quite literally, outside the room where it was happening. That second day, they were once again competing for attention. Altman was attending an exclusive dinner with sixty House members at the Capitol, feasting on an expertly prepared buffet with roast chicken. At the same time, the artists were hosting an interactive cocktail hour and trying to attract as many staffers with the best their budget could buy: wine and Chick-fil-A.

It was a small but darkly comedic illustration of who commanded power and influence in the AI policy conversation and who didn’t.

---

The same narrative Altman had long used within OpenAI to justify hiding its research and moving as fast as possible was now being expertly wielded to steer the US AI regulatory discussion toward proposals that would avoid holding OpenAI accountable, and in some cases entrench its monopoly. Silicon Valley’s tried-and-true “What about China?” card had consistently done wonders to ward off regulation. Now it was punchier than ever, with Washington’s fears about China, fueled by TikTok’s stunning rise, reaching new heights.

That fear could be typified by the mood at the Department of Commerce, which had become the leading edge of an aggressive US government offensive to throttle China’s AI development. The previous year, on October 7, 2022, Commerce had released a directive that it said was meant to undercut Chinese AI military advancements. Using a mechanism called export controls—a way to limit the sale of certain technologies to foreign countries on the grounds of national security—it clamped down without warning on the export of cuttingedge American-designed AI chips, primarily Nvidia’s, to China. The blast radius of this move was far wider than the Chinese military; it pulled the rug out from under the Chinese scientific community and AI industry working on everything from AI health care and education applications to Chinese ChatGPT equivalents. “If you’d told me about these rules five years ago, I would’ve told you that’s an act of war—we’d have to be at war,” a semiconductor analyst said.

Taking stock of the aftermath, Commerce seemed frustrated by China’s seeming resilience. The country’s pace of AI development and adoption had slowed down some but not nearly enough to give the department comfort. Part of this was due to Nvidia’s own maneuvering: China represented a massive market for the American chipmaker, and just as quickly as Commerce had laid down its constraints, Nvidia had designed new chips that fell neatly within them in order to keep selling to Chinese customers. The ban was also a lift to China’s own chipmaking industry, which had long struggled to produce chips as good as Nvidia’s. The US government’s actions had generated a surge of interest for Chinese domestic alternatives, giving the industry a big funding and feedback boost to advance.

But the biggest challenger to its efforts was the vibrant cross-border open-source AI movement, which was rapidly replicating closed corporate generative AI models and putting them out on the internet for anyone to download and use. After vigorously playing catch-up, Meta had become a dominant player in freely putting out its large language models. The company had long been a champion of open-source development; chief scientist Yann LeCun believed in the importance of open science. It was also smart business. Meta didn’t need to sell generative AI models to make money, but unleashing free ones, while integrating them into its core products, could help it establish its AI leadership, attract top scientific talent, and taunt its competitors for that talent who did depend on selling their models.

Having amassed around the level of compute resources that OpenAI had through Microsoft, Meta was now full steam ahead on producing its equivalent of the GPT series, called Llama. Llama didn’t technically clear the true definition of open source, which would have required releasing both the model weights and its training data. But Meta’s follow-through on just the first aspect, publishing its model weights for free, had been enough to turn Llama—despite its policy to not make Llama directly available in China—into a critical building block for the Chinese AI industry.

Amid the climate of frustration and fear in Washington, a policy white paper echoing Altman’s recommendations arrived two months after his hearing in July 2023. Written by a consortium of researchers, including from OpenAI’s Safety clan, Microsoft, and Google DeepMind as well as more than a dozen think tanks, many tied to the Doomer community, it pushed once again for a new licensing regime for AI models using compute thresholds, and the development of AI safety evaluations for dangerous capabilities including the ability to manipulate and persuade and the creation of novel biological weapon recipes. The fifty-one-page document also gave a name to the category of models that needed this government intervention: “frontier” AI models. Per the authors, frontier models —models that might exhibit these dangerous capabilities—did not yet exist, but by scaling existing models from companies like OpenAI and Anthropic with ever more compute, they could arise suddenly and unpredictably at any moment.

Within weeks of the white paper, OpenAI and Microsoft formed a strategic alliance with Google and Anthropic to launch the Frontier Model Forum, a group for advancing relevant research and influencing the policy agenda on AI safety risks. It was a rare issue in which the interests of Doomers, Boomers, and profit-motivated corporates aligned: Keeping frontier models front and center in government regulatory discussions was ideologically imperative to Doomers and convenient to Boomers and corporates for shifting attention away from regulating existing AI models and their problems. Everyone also advanced their causes by arguing against opening up the weights of cutting-edge models. In the Forum’s first year, Meta was conspicuously absent. (It would join a year later to gain a seat at the table.)

---

Core to Altman’s recommendations and the idea of the frontier model was the association of a model’s scale with emergent, and thus possibly dangerous, capabilities. Such an argument was rooted in the philosophy of scaling laws— that more training compute should predictably result in more powerful models— as well as the belief within Doomer circles that highly advanced AI could go rogue. The policy proposals that flowed from this argument centered on regulators basing their interventions on how much compute was being used to train a deep learning model: Models that crossed a certain compute threshold should automatically be viewed with more caution and restricted more tightly.

The July 2023 policy white paper suggested a number for that threshold: 1026 floating point operations, referring to the minimum total number of calculations—1 with twenty-six zeros after it—that a model needed to be trained with to be designated as a frontier model. The authors had admitted that the threshold was somewhat arbitrary, stating simply that it was a level of compute that existing AI models likely hadn’t yet surpassed. Sara Hooker, the VP of research at large language model developer Cohere and one of the coauthors of the paper, says speaking with her collaborators who proposed the number led her to believe they had picked it to be slightly higher than the amount of compute that OpenAI had reportedly used to train GPT-4.

But Hooker and many other researchers, including Deborah Raji, disagree with the compute-threshold approach for regulating models. While scale can lead to more advanced capabilities, the inverse is not true: Advanced capabilities do not require scale. A deep learning model trained only on high-quality biological data, for example, can be a very powerful generator of biological recipes at very small scale. Through distillation, one of the techniques that OpenAI referenced in its 2021 research road map, large models can also be transformed into small models with similar capabilities. Scaling models doesn’t guarantee advancements in certain capabilities either; that depends once again on what’s in the model’s training data as well as which type of neural network is being trained. In the end, not all models are built on Transformers. Compute is thus not much correlated with risk at all, says Hooker, let alone with specific kinds of risks, such as the ones laid out in the white paper.

Without consensus among the white paper’s coauthors about either the compute-centered regulatory framework or the specific threshold, the number, 1026 floating point operations, was placed in a footnote and the appendix without justification, alongside significant caveats for why thresholds were a highly imperfect approach. What shocked Hooker was how quickly not just the framework but also the exact threshold rapidly turned into one of the most popular policy proposals. It captured significant mind share in Washington, after the white paper tapped straight into fears of China. Frontier models sounded scary, and even more so if Beijing got ahold of them. “Parts of the administration are grasping onto whatever they can because they want to do something,” Emily Weinstein, then a research fellow at CSET, told me in late 2023.

The white paper’s ideas found a receptive audience at Commerce. Staff mobilized to meet with experts to hash out what controlling frontier models could look like and whether it would be feasible to keep them out of the reach of Beijing. Soon it was considering an unprecedented proposal to expand its AI export controls to focus on not just hardware but the software itself by banning the export of AI models above a compute threshold. In other words, it was evaluating whether it could block model weights from being posted on the internet and made widely available.

Notably, a key recommendation from Marcus and IBM VP Christina Montgomery, who also testified alongside Altman, did not gain nearly as much traction, despite their repeating it throughout the hearing: compelling companies to disclose what exactly is in the training data they feed into their models. This would have little impact on handing over more advanced capabilities to Beijing per Washington’s concerns but would give real teeth to corporate accountability on a broad range of issues, including company use of copyrighted materials, user data privacy, and rigorous scientific evaluations of model capabilities. “If we don’t know what’s in them, then we don’t know exactly how well they’re doing,” Marcus had said. We’d simply have to take a company’s word for it.

Such an approach would also significantly ameliorate the uncertainty of if and how dangerous capabilities might emerge, for the same reason why compute is a poor risk predictor: A deep learning model’s behavior first and foremost derives from its data. If an AI developer produces a large language model that is able to create recipes for bioweapons, “it’s because they trained it on a dataset that included information on bioweapons,” says Sarah Myers West, the co–executive director of AI Now Institute and former senior adviser on AI to the FTC. As always, the neural network is surfacing patterns within its training data. Opening up that data would be the first step to establishing scientific clarity on what kinds of inputs could lead to dangerous outputs.

---

As Commerce consulted various experts on its proposal to clamp down on model weights, news of its deliberations, which it would announce in early 2024 with a public request for comment, cleaved the AI development community and the rapidly expanding AI policy community into two. This clash was about Closed versus Open, techno-nationalism versus borderless science. In addition to the Frontier Model Forum participants and broader Doomer community, the Closed side quickly won over the US national security and intelligence apparatus. Facing off against them was Meta, open-source AI developers, startups, civil society groups, and independent academics.

Where the Closed side continued to emphasize many of the same points that OpenAI executives had used for years internally, the Open side argued that sequestering models would do far more harm than good. The bottleneck for producing novel biological weapons, for example, is not about finding a recipe, Weinstein noted. Such recipes already abound online and are easily found via Google. It is about obtaining the materials and equipment to actually make the armaments. Restricting access to so-called frontier models would thus do little to fix this. But the collateral damage of suppressing the publication of AI models would risk weakening the foundations of US AI innovation. Open source— sharing and building on code and software released to the broader community— has long been the bedrock upon which the wealth of US-based startups flourish. Restricting model weights from being published would give smaller developers fewer pathways than ever to create their own AI products and services. It would further entrench the dominance of the giants represented in the Frontier Model Forum.

AI models would also become ever harder to scrutinize, such as in the work of Sasha Luccioni, Yacine Jernite, and Emma Strubell, who have relied heavily on open generative AI models to quantify the carbon and environmental costs of continuing to scale them.

In critical ways, contrary to it being a national security risk, a great deal of open collaboration across borders had also strengthened American AI leadership. As the two countries that produce the most AI talent and research in the world, the US and China have long been each other’s number one collaborator in AI development. For more than a decade, scientists and entrepreneurs in both countries have riffed off one another’s work to advance the field and a wide array of applications far faster than either group would have alone, benefiting not just each country but many others globally. One of the most famous examples: ResNet, among the most widely used neural networks in the world, was published by Chinese researchers in Microsoft’s Beijing office. ResNet not only underpins major computer-vision, speech-recognition, and language systems but also was a core ingredient of the first version of DeepMind’s AlphaFold, an AI system released in 2018 that could predict a protein’s 3D structure from its amino acid sequence, crucial for accelerating drug development and understanding disease. (DeepMind’s subsequent advancements in AlphaFold, using a different neural network, would earn Demis Hassabis and another senior research scientist at DeepMind a 2024 Nobel Prize in Chemistry.)

And yet, in October 2023, the ideas championed by the Closed side would gain their greatest endorsement yet when they surfaced in the Biden administration’s AI executive order. The order, one of the longest in history, would read like smashed-together documents written by completely different groups—because it was. One of those documents was rooted in the administration’s 2022 Blueprint for an AI Bill of Rights, which the White House had carefully assembled over time through consultations with civil society groups to outline how AI could be advanced, used, and reined in in ways that bolstered civil rights, racial justice, and privacy protections. Among other things, it emphasized developing AI with broad participation from communities and experts, addressing the discriminatory impact of AI in contexts such as health care and hiring, and protecting people from data collection without their consent.

The other document was a surprisingly faithful reproduction of Altman’s recommendations and the framing of the frontier model white paper, which had been stapled on at the last minute after the paper gripped the attention of a few people sympathetic to the Doomer ideology in the White House. The white paper had emphasized the need to focus on future AI models that didn’t yet exist; the executive order would subsequently focus half of its real estate on such models. The white paper had outlined four examples of dangerous capabilities. The executive order would keep three of them: the generation of novel CBRN (chemical, biological, radiological, and nuclear) weapon recipes, automated cyberattacks, and, in a straight copy and paste, the evasion of human control “through means of deception and obfuscation.”

To the alarm of Hooker, Raji, and many other AI researchers, the white paper’s exact compute threshold, 1026 floating point operations, would also show up in the executive order as the threshold above which models would need to be reported to the US government.

With such an endorsement, the compute-threshold approach would quickly metastasize. By the end of the year, it would get picked up in Europe, which settled on 1025 for something slightly more restrictive, as lawmakers pushing through the long-gestating EU AI Act felt steamrolled by the sudden generative AI developments and hurriedly searched for ways to account for them. At the start of 2024, the approach would then spread to California with the introduction of a new AI safety bill called SB 1047, which would return to 1026 as its threshold. California governor Gavin Newsom would subsequently veto the bill, which, in an ironic twist, OpenAI and other model developers heavily lobbied against for its wide array of other accountability proposals. Many would criticize Newsom for bowing to industry interests. But to several researchers, including Hooker and Raji, the veto was a welcome development. “It was a step in the right direction to make sure we’re anchored to scientific consensus,” Hooker says. “There are big questions that remain about why that number and what risks are you hoping to prevent.”

The whole sequence of events—Altman’s testimony, the white paper, the all-out policy influence campaign, Washington’s hyperreactivity to fears of China, and the hasty enshrining of compute thresholds into consequential policy documents within the US and abroad—was a stark illustration among other things of how much independent AI expertise had atrophied. The prior month, in September 2023, Raji had found herself the singular academic, with financial ties neither to the industry nor Doomer community, testifying to Congress next to Altman, Musk, Nadella, Gates, Zuckerberg, Pichai, and Jack Clark, among other tech executives. They were all present for the very first of Senator Chuck Schumer’s AI Insight Forums, among the hottest and most consequential series of policy convenings that year to set in motion AI legislation. As her fellow witnesses spouted spectacular, unbacked claims about the promises and perils of AI, peppered with well-timed references to beating China that straightened the backs of attending senators, what shocked Raji the most was how much many in the audience appeared to buy into everything.

It dawned on her that the people sitting next to her, and their massive policy teams, had monopolized the message in Washington for so long that many policymakers now viewed it as gospel. A Schumer spokesperson would later note in the press that the senator was personally consulting with Altman and other OpenAI executives as he moved closer to regulation. “That for me was a huge realization,” Raji says. “Wow, we need more people just debunking—just looking at what people are saying and being like, ‘Actually, reality is more complicated.’ ”

---

Washington was only the climax of the US leg of Altman’s policy charm offensive. In March of that year, after tweeting that he planned to travel abroad to meet with users, his trip had evolved into a multicity, multi-continent odyssey to sit for photo ops with seemingly every president in the G20. It now had new branding: Sam Altman’s World Tour.

There was no grand strategy from OpenAI’s communications or policy teams behind the World Tour. Altman had just selected his initial stops and blasted them off to his more than 1.5 million Twitter followers. After the first few appearances, interest had snowballed out of control, and the comms and policy teams were roped in. With each new leg, the teams scrambled to arrange the logistics, bracing for the difficulties guaranteed to arise from the lack of preparation.

It was a manifestation of a dynamic that had always been present: Altman going his own way. Sometimes that way lined up with the company; sometimes it did not. During the release of GPT-4, OpenAI had carefully crafted all of its announcements and publicity to present the project as the company-wide effort that it was. The model had involved over a hundred employees. The author of the company’s announcement was simply “OpenAI.” Altman had then tweeted credit to a single person: Jakub Pachocki. Pachocki had indeed played an important role, but he had been one of eighteen leads on the project. Was his contribution really singular? Some employees wondered. Altman had then leaned in further, tapping Pachocki to sit behind him during his Senate hearing in Washington.

The dynamic also showed up in other ways: There were official executives at the company, but it wasn’t always clear that they were the ones whom Altman was listening to, nor whether official processes and decisions or his relationships and whims were the ones guiding company strategy.

As OpenAI was rapidly professionalizing and gaining more exposure and scrutiny, this incoherence at the top was becoming more consequential. The company was no longer just the Applied and Research divisions. Now there were several public-facing departments: In addition to the communications team, a legal team was writing legal opinions and dealing with a growing number of lawsuits. The policy team was stretching out across continents. Increasingly, OpenAI needed to communicate with one narrative and voice to its constituents, and it needed to determine its positions to articulate them. But on numerous occasions, the lack of strategic clarity was leading to confused public messaging.

At the end of 2023, The New York Times would sue OpenAI and Microsoft for copyright infringement for training on millions of its articles. OpenAI’s response in early January, written by the legal team, delivered an unusually feisty hit back, accusing the Times of “intentionally manipulating our models” to generate evidence for its argument. That same week, OpenAI’s policy team delivered a submission to the UK House of Lords communications and digital select committee, saying that it would be “impossible” for OpenAI to train its cutting-edge models without copyrighted materials. After the media zeroed in on the word impossible, OpenAI hastily walked away from the language.

“There’s just so much confusion all the time,” says an employee in a public-facing department. While some of that reflects the typical growing pains of startups, OpenAI’s profile and reach have well outpaced the relatively early stage of the company, the employee adds. “I don’t know if there is a strategic priority in the C suite. I honestly think people just make their own decisions. And then suddenly it starts to look like a strategic decision but it’s actually just an accident. Sometimes there isn’t a plan as much as there is just chaos.”

---

As Altman zipped around, flying to Europe, Latin America, the Middle East, Asia, and Africa, dazzling—and only on a few occasions offending—carefully curated audiences of students, tech investors, and fans, the lack of strategic clarity was inflaming OpenAI’s age-old rift lines and accelerating the company toward more opposite extremes than ever before.

On one side, the Applied division was still leading the charge, racing against an unprecedented number of competitors to deploy OpenAI’s technologies faster than ever. It was now also bolstered by the other ballooning divisions as well as many people in Research, invigorated by their belief from the dramatic increase in hype and expectation that advancing and releasing OpenAI’s models was the best way to achieve the company’s mission.

On the other side, the Safety clan, spread out across Research, many still concentrated within Miles Brundage’s policy research team and Jan Leike’s alignment team, were now a far smaller minority. They were compensating for their relative size disadvantage by sounding the alarm louder than ever on the dangerous capabilities and existential risks that they believed could become imminently possible. As OpenAI’s models continued to advance, some within Research who didn’t previously identify with the Safety clan were also joining its ranks as the accelerating capabilities converted them to the belief that AI could reach a point of intelligence that would allow it to subvert human control and go rogue. The same dramatic increase in hype and expectation on this side meant OpenAI had a moral imperative to act with maximum caution, in order to fulfill its mission.

It was the Boomers and Doomers incarnate—within OpenAI’s walls. The split reached all the way to leadership. After DALL-E and ChatGPT, most executives and senior managers had grown increasingly comfortable with models as beneficial tools to be put into the world through “iterative deployment,” a phrase that OpenAI had coined between releases to describe its new approach. Unlike the staged release of GPT-2, or the controlled API release of GPT-3, iterative deployment was about going all in—putting models in the hands of users early and often. And as with all of its deployment strategies in the past, OpenAI had a new argument for why this one was the safest approach possible. Iterative deployment, Altman and other executives argued, would give people and institutions time to adjust while allowing it to test its models on real people, collect real feedback, and improve its products. “Going off to build a superpowerful AI system in secret and then dropping it on the world all at once I think would not go well,” Altman had said during his Senate testimony.

But if there was one leader at OpenAI not moving in lockstep, it was Sutskever. As OpenAI’s models advanced and the impact of their deployments accelerated, he believed the company needed to raise, not lower, its guard against their potential to produce devastating consequences. After GPT-4, Sutskever, who had previously dedicated most of his time to advancing model capabilities, had made a hard pivot toward focusing on AI safety. He began to split his time half and half. To people around him, he seemed at times to be at war with himself. He was both Boomer and Doomer: more excited and afraid than ever before of AGI arriving and rapidly surpassing humans to become superintelligence.

Sutskever now spoke in increasingly messianic overtones, leaving even his longtime friends scratching their heads and other employees apprehensive. During one meeting with a new group of researchers, Sutskever laid out his plans for how to prepare for AGI.

“Once we all get into the bunker—” he began. “I’m sorry,” a researcher interrupted, “the bunker?” “We’re definitely going to build a bunker before we release AGI,” Sutskever replied matter-of-factly. Such a powerful technology would surely become an object of intense desire for governments globally. It could escalate geopolitical tensions; the core scientists working on the technology would need to be protected. “Of course,” he added, “it’s going to be optional whether you want to get into the bunker.”

The researcher would in equal parts continue to hold Sutskever in high regard and keep himself at arm’s length. “There is a group of people—Ilya being one of them—who believe that building AGI will bring about a rapture. Literally, a rapture,” he says.

As Sutskever continued splitting his time on alignment, a new idea began to percolate between him and Altman: a team laser focused on developing new alignment methods for superintelligence, in anticipation of methods like reinforcement learning from human feedback no longer being sufficient once systems could, in their view, outsmart humans. Altman called it the Alignment Manhattan Project. At first, the two discussed spinning it out as a different organization: its own independent nonprofit with a starting endowment of $1 billion. In part because Sutskever didn’t want to leave OpenAI and in part due to model access issues, they decided to keep the project within the company. OpenAI subsequently announced the formation of a new team to oversee the effort in a blog post along with its new name, Superalignment. The post also announced a flashy commitment to dedicate 20 percent of the computing power OpenAI had secured to date to the team. Sutskever and Leike would colead the new effort.

At another meeting, Sutskever stepped up in front of employees to introduce the new team and its goals. He grabbed the microphone and began to thump it. Boom. Boom. Boom.

“Alignment is a burning fire,” he said. “Superalignment is a blazing inferno.”

Not long thereafter, with the company outgrowing Mayo, the plant-filled, fountain-adorned office it had moved into after the pandemic, executives shifted most of the Research division back to the old Pioneer Building. The move largely divided the two halves of the company—Applied and Safety—into their own worlds.

---

In July 2023, shortly after OpenAI made news of the Superalignment team public, it rented out a theater at the Metreon in downtown San Francisco for employees to see the movie Oppenheimer, the story of physicist J. Robert Oppenheimer as he led America’s Manhattan Project to create the world’s first nuclear weapon.

“i was hoping that the oppenheimer movie would inspire a generation of kids to be physicists but it really missed the mark on that,” Altman tweeted. “let’s get that movie made! (i think the social network managed to do this for startup founders.)”

For nearly eight years, the analogy that Altman had made in his very first emails to Musk between OpenAI and the Manhattan Project had been a persistent motif within the company, used even during new-hire orientations. Altman was fond of it. He shared a birthday with Oppenheimer, which he’d point out to reporters. He also liked to paraphrase the bomb maker’s belief that “technology happens because it’s possible.” He never seemed to add that Oppenheimer spent the second half of his life plagued by regret and campaigning against the spread of his own creation.

Different employees ascribed different significance to the analogy. Most saw the Manhattan Project as a heroic feat; it represented the ability to pull off a world-saving, history-changing technological breakthrough before dangerous adversaries with a significant concentration of talent and resources. Among the Safety clan, it emphasized the gravity of OpenAI’s burden to usher in a technology that risked the existential demise of humanity.

To Altman, it represented a PR lesson. “The way the world was introduced to nuclear power is an image that no one will ever forget, of a mushroom cloud over Japan,” he had once said, years earlier at an event. “I’ve thought a lot about why the world turned against science, and one answer of many that I am willing to believe is that image, and that we learned that maybe some technology is too powerful for people to have. People are more convinced by imagery than facts.”

Those at the extreme end of the AI safety spectrum with the highest levels of p(doom) grew increasingly unsettled by the seeming uncomplicated optimism with which Altman and the rest of the company viewed this history. OpenAI had at one point also organized a screening of Apollo 11, a documentary about the US Apollo program to launch the first man to the moon, which was one of Silicon Valley’s other favorite analogies. “Why would you ever talk about the Manhattan Project if you could just say ‘Apollo program’? Why bring along that baggage?” an extreme Doomer says.

One scene from Oppenheimer stuck with him in particular: the moment before the Trinity test, the first ever detonation of an atomic bomb, when Oppenheimer, played by Cillian Murphy, calculates that the chances of it blowing up the world are “near zero.”

“Near zero?” responds Major General Leslie R. Groves, played by Matt Damon, incredulously.

“What do you want from theory alone?” Oppenheimer says. “Zero would be nice,” Groves says. “That’s analogous to the situation we’re in,” the extreme Doomer says. “No way can we calculate anything about these AIs yet.”

---

As OpenAI continued to push on its research, the Manhattan Project analogy for some began to take on a new meaning.

With the rate of advancement in large language models slowing with the exhaustion of data and compute, the Research division had pivoted more heavily toward developing AI agents. The idea, gaining traction across the field, was a return to the debate between the “pure language” and “grounding” hypotheses. Pure language was reaching the end of its rope, as was combining language and vision. To many in the AI community, it seemed like the next stage of advancement would likely need to come from agents that could take actions in the real world and collect feedback from its environment. Within OpenAI, such a capability was also seen as a way to gain a competitive advantage. An AI assistant that could chat with you was nice, but one that could automate complex tasks, such as sending emails or coding websites, was even better. Not only was this highly commercially relevant, it could also accelerate the company’s own progress.

The most ambitious of these efforts in the Research division was AI Scientist, an attempt to build an agent for autonomously performing scientific research. With limited new “knowledge,” or data to scrape from the internet or textbooks, researchers on the project had high hopes that an autonomous “scientist” would be able to generate its own knowledge by running experiments. The team had formed from a merger of the previous code-generation team working on Codex with another team that had been trying to crack the reasoning challenge by using large repositories of math problems and their solutions as a structured dataset to teach its model step-by-step logic. Both capabilities— coding and solving math problems—seemed like good building blocks for conducting experiments and analyzing data.

Leading the project were Jakub Pachocki and Szymon Sidor, who were focused on creating not only an autonomous scientist but, specifically, as Altman desired, an autonomous AI researcher—one that would help OpenAI supercharge its AI advancements.

The project made AGI believers both extremely excited and extremely nervous: If AI Scientist succeeded, AGI would surely arrive faster. This shortened the timeline either to utopia or to humanity’s obliteration. Like Sutskever, some researchers began to reference “a bunker” in casual conversations, even imagining a setup similar to Los Alamos: Somewhere out in a remote patch of American desert, an elite team of AI researchers would live and work in secure facilities to protect them from outside threats.

In 2023, they believed those threats now included targeted attacks and rogue AGI itself. On Slack, the security team posted a draft of its threat model for OpenAI, significantly matured from the days when executives had debated how much to heighten the security of the organization. The draft included three categories of threats: foreign state actors, competitors, and ideologically motivated people.

In the third category, the draft linked to an article by Eliezer Yudkowsky, an extreme Doomer and leader in the AI safety community who had coined and popularized the phrase friendly AI to refer to well-aligned systems and wrote a beloved work of fan fiction called “Harry Potter and the Methods of Rationality.” The serial novel, which spans 122 chapters and over 660,000 words, reimagines Harry engaging in the wizarding world as a well-trained rationalist. It had served for many as a gateway into effective altruism and, in turn, to broader Doomer ideology. Yudkowksy had also cofounded the blog LessWrong, a central hub for AI safety researchers to foster community and propagate AI safety ideas, where he’d advocated with increasing alarmism to put a full pause on AI development as his p(doom) shot up to 95 percent. In March 2023, he’d written an article for Time magazine where he discussed the grief of watching his daughter lose her first tooth and wondering if she would have a chance to grow up. He proposed a plan to enforce the halting of AI advancement by shutting down all large GPU clusters, tracking sales of GPUs, and, if necessary, targeting “a rogue data center” with air strikes. OpenAI’s threat model draft linked to this piece as an example of an ideologue advocating for violence.

A small faction of OpenAI employees who were fans of Yudkowsky’s less violent opinions found the fact that the draft singled him out but didn’t mention unfriendly AI itself deeply frustrating. After getting feedback, the security team added a fourth threat category: misaligned AGI systems.

---

As OpenAI skyrocketed to new prominence, the board was shedding members without replacing them. Since the start of 2023, it had lost three independent directors in rapid succession, in part due to the feverish race that ChatGPT had sparked to build and commercialize generative AI technologies across the industry.

Reid Hoffman had been the first to step down in February after five years on the board, due to conflicts of interest. The previous year, he had cofounded a startup, Inflection, with the now-departed DeepMind cofounder Mustafa Suleyman, which was fast evolving into a direct OpenAI competitor.

A month later, Hoffman was followed by Shivon Zilis, Musk’s trusted deputy and Neuralink director, who, after Musk disaffiliated, had continued to oversee OpenAI on his behalf and officially joined the board in 2020. Some in leadership had long worried that Zilis would feed sensitive company information to Musk, but she had pledged to uphold her confidentiality to OpenAI over her loyalty to its spurned former cochairman. That position became highly questionable once news broke in July 2022 that she had had twins with Musk without disclosing it to her fellow directors. Still, Altman had sought to keep her on the board for reasons that eluded the other board members, at one point seeking her advice in October 2022 on how to handle Musk’s apparent irritation over OpenAI’s escalating valuation, by then reportedly nearing $20 billion. “This is a bait and switch,” Musk had texted Altman after noting his substantial contributions to OpenAI’s initial funding. In March 2023, Zilis’s continued board role finally turned untenable as Musk incorporated a new AI venture, xAI, to be another direct competitor to OpenAI.

Third to depart was Will Hurd, a former Republican Texas representative and former CIA officer, who had joined the board in 2021. During the announcement of Hurd’s appointment, Altman had told employees that it was important to have someone that balanced out the liberal bias of Silicon Valley. He then organized a meeting for anyone who had reservations to ask Hurd any questions. Employees didn’t hold back, grilling Hurd about his views on different issues, including Donald Trump. In June 2023, Hurd parted ways with OpenAI to focus full time on his US presidential campaign. He would withdraw from the race by October.

Alongside Altman, Brockman, and Sutskever, only three independent board members remained: Quora cofounder and CEO Adam D’Angelo, roboticist Tasha McCauley, and CSET researcher Helen Toner.

Among the trio, D’Angelo had joined the board first. D’Angelo, a high school classmate of Zuckerberg’s at the boarding school Phillips Exeter, had served for two years as the CTO of Facebook before starting Quora. In 2014, Quora had joined YC in the first batch under Altman’s presidency; in 2017, a year before D’Angelo’s board appointment, Altman had topped up YC’s investment into Quora, coleading an $85 million round of funding. In the announcement, Altman praised D’Angelo as one of “the smartest CEOs in Silicon Valley.” “And he has a very long-term focus, which has become a rare commodity in tech companies these days,” Altman said.

McCauley had joined the board later in 2018. An entrepreneur who had cofounded a telepresence robotics startup and was running a 3D urban simulation company, she had connected with Altman through her mentor, Alan Kay, and also knew Holden Karnofsky. McCauley was well-respected in the AI safety community and would serve on the board of the AI safety research nonprofit Centre for the Governance of AI and for a time on the board of the Effective Ventures Foundation, a UK-based organization that oversees the popular EA podcast 80,000 Hours. Karnofsky, who had been on the OpenAI board at the time, nominated her in his effort to find more independent directors as OpenAI prepared to transition into a capped-profit structure.

Toner had been the last addition to the board in mid-2021, also via a nomination from Karnofsky. This time he was recommending candidates to replace himself as his three-year term came to a close and with the formation of Anthropic. Before establishing herself as an expert on China and emerging technologies, Toner had worked with Karnofsky at GiveWell and Open Philanthropy. She had come to AI safety issues through the EA movement but had slowly pulled back from the latter over time. In her most popular EA Forum post before the FTX crash, she had observed that the movement was growing increasingly dogmatic and socially insular, and noted that she was “leaning into EA disillusionment.” She continued to be highly regarded in its circles and to dedicate herself to the broader AI safety community, serving with McCauley on the board of the Centre for the Governance of AI.

---

By the late summer of 2023, the board had been in a monthslong deadlock over whom to appoint as new independent directors. As part of their effort to increase oversight after the GPT-4 demo, and even more after the launch of ChatGPT, McCauley had engaged in a roughly yearlong process, including interviewing employees and stakeholders outside the company, to articulate what a revamped board and more professionalized oversight mechanisms should look like.

During the process, the board, including Altman, had all agreed that based on what they saw as the rising stakes of the company’s capabilities, the next independent director needed to have a deep background in AI safety. They’d subsequently spent months compiling a list of candidates and interviewing five of them, including Dan Hendrycks, a central figure in the Doomer community running the Berkeley-based Center for AI Safety and serving as the only adviser at Musk’s xAI. But as the independent board members sought to move forward to select one of the five, Brockman and Sutskever each raised various issues with the vetted candidates. Altman was demure as always, not outright disagreeing with any option but not moving any of them forward either. Several times, he also suggested new candidates, who shared the same characteristic: They were all embedded in his network and, financially or otherwise, within his sphere of influence.

When it came to establishing the new oversight mechanisms, which included different channels for increasing the board’s visibility into the company’s safety and security practices, the independent directors were also left with a similar feeling that they weren’t a priority for Altman. Early in McCauley’s tenure as a director, Altman had designated her the board’s employee liaison and advocate; she subsequently met with employees regularly by holding office hours. Once she had also brought her husband, actor Joseph Gordon-Levitt, to a company off-site, where he’d listened intently to technical presentations. But during the pandemic, those meetings had petered out. Afterward, McCauley continued to keep some regular meetings, but the open office hours never restarted.

Without a systematic way of connecting with employees, information about the company’s happenings was instead filtering up to the independent directors through their own personal relationships from the broader AI safety and tech communities. They also relied on Altman himself as a conduit for keeping tabs on important information. What worried them with growing intensity was how much Altman’s rhetoric often differed from other accounts they were hearing. Where Altman regularly portrayed a rosy picture, the directors increasingly received reports from their own sources about various problems, including the company’s lack of preparation before and significant tumult after ChatGPT, the continued AI safety concerns surrounding GPT-4’s release, and the unprecedented pace with which OpenAI was sprinting to launch new products before it had resolved many of its issues.

One incident felt particularly glaring. In late 2022, the board had had an on-site—the first of what was meant to be an annual meeting—during which Altman had highlighted the strong safety and testing protocols that OpenAI had put in place with the Deployment Safety Board to evaluate GPT-4’s deployment. After the meeting, one of the independent directors was catching up with an employee when the employee noted that a breach of the DSB protocols had already happened. Microsoft had done a limited rollout of GPT-4 to users in India, without the DSB’s approval. Despite spending a full day holed up in a room with the board for the on-site, Altman had not once notified them of the violation.

While the independent board directors didn’t have reason to believe that anything unsafe had been released to the public, they were unsettled by the seeming disregard with which Microsoft had broken protocol and Altman had passed over it. OpenAI’s models were on a rapid advancement trajectory and, from an AI safety perspective, they believed, could soon pass a point where such a violation could result in potentially catastrophic, if not existential, consequences. In their view, Altman’s laxness with Microsoft’s breach set a dangerous precedent for how he might treat AI safety processes around model releases once the stakes went up.

Meanwhile, there were other concerning examples of Altman’s behavior. In March 2023, he had emailed the board without D’Angelo and announced that he believed it was time for D’Angelo to step down. The fact that Quora was designing its own chatbot, Poe, Altman argued, posed a conflict of interest. The assertion felt sudden and dubiously motivated. Toner, McCauley, and D’Angelo had each at times asked Altman inconvenient questions, whether about OpenAI’s safety practices, the strength of the nonprofit, or other topics. With his allies on the board dwindling, Altman seemed to the three to be fishing for an excuse to push one of them out. In response to Altman’s email, an independent director pushed back. Poe was a moderate conflict of interest compared with what Altman had long allowed to stand from Hoffman and Zilis. Altman’s motion failed; D’Angelo stayed on.

Shortly thereafter, D’Angelo was at a dinner party when he heard that OpenAI’s Startup Fund was structured weirdly. It was giving those who invested in the fund early access to OpenAI’s products, a kind of preferential treatment that should have been reserved for OpenAI’s own investors. After he heard it come up a second time, the independent directors pressed Altman for documents about the fund’s structure. When Altman finally handed them over, the directors discovered that the structure of the fund wasn’t just weird; Altman legally owned it when it should have been owned by OpenAI.

For the independent directors, every instance added up to a single troubling picture: Bit by bit, Altman was trying to cloud their visibility and maneuver in ways that prevented the board from ever being able to check him. For years, Altman had advertised the board’s ability to counterbalance and even fire him as OpenAI’s most important governance mechanism. Indeed, he was now trumpeting this fact around the world to secure public and government trust.

By the fall, the morale of the independent directors hit new lows as they struggled to make any meaningful progress in the negotiation for a new board member. With every passing month, these gaps in governance were gaining urgency. OpenAI was getting ready to train GPT-5, and it was making progress on AI Scientist. Then, out of the blue, in early October, Toner received an unlikely email: Ilya Sutskever wanted to talk.
