Skip to content

10 Gods and Demons

To live in San Francisco and work in tech is to confront daily the cognitive dissonance between the future and the present, between narrative and reality.

The first time I moved to San Francisco, as a university sophomore for a summer internship, I was dazzled by the quaint aesthetics of the city. The colorful Spanish-style architecture, the limited number of skyscrapers, the hills steep enough to make driving a stick shift a test of reflexes. There was an endless supply of perfectly ripe avocados and toasted sourdough bread and smooth Blue Bottle lattes. There were different neighborhoods, all with their own look and culture.

When I returned full time after graduation to work at a tech startup, I crammed into a three-bedroom apartment with three other roommates in the Castro. On weekends we would hike across the rolling hills and forage from public fruit trees. On weeknights, neighbors—all young twentysomethings in the tech industry—would pop over unannounced to play board games, drink wine, and while away the evenings. House parties were a constant, as were weekend trips to stunning nature: Lake Tahoe in the north, Big Sur to the south, tall, majestic redwoods everywhere around us. Life was easy. We were young, making salaries relatively standard in the tech industry that placed us nationally in our age group’s top 5 percent.

But there was that dissonance. On the way to work, I would pass people shooting up drugs in front of the subway stations, the unhoused peeing on sidewalks just blocks from my office. Meanwhile, our startup’s chef, playfully named “the happiness engineer,” would cook or cater an abundance of food for our free office lunches. Leftovers often went straight into the trash. If we stayed late, we got free dinner—and were emphatically implored to take an Uber home for safety reasons. It was all too easy for the privileged to grow accustomed to moving through the city in ways that shielded them from seeing the realities of how the other half lived.

The dichotomy encapsulated how the tech industry could profess big, bold visions about changing the world and building a better future while ignoring the very problems at its door. It was a dichotomy that Altman would sometimes comment on in his own way—getting right up to yet never fully acknowledging the utter contradiction of declaring the problem of creating and managing beneficial AGI possible, but San Francisco’s housing crisis too tough to tackle.

“Where I grew up, no one would ever walk by a person collapsed on the side of the street on their way to work and not do something about it,” he once said, comparing suburban St. Louis to San Francisco. “I do blame the tech industry for a lot of things that have gone wrong with the city, but not all of them. But we have, just over time, had this, like, unbelievable wealth generation in this small geographic space, in this small period of time, and I think not been particularly thoughtful about the effects of that on the community as a whole. And because those problems are so hard and so hard to think about, I think most people just choose not to, and they just accept this.”


It was in this context that effective altruism arrived from the UK and found its most loyal audience. EA, to which many in OpenAI’s Safety clan were early adherents, made for the perfect Silicon Valley ideology. It preaches making the world a better place and doing it with rigorous logic, being disciplined enough to focus on the far future instead of the present, and fervently embracing the principles of capitalism and libertarianism—all in the name of morality.

Core to the EA philosophy is a mathematical concept called “expected value.” The expected value of something is calculated by multiplying the probability that it will occur with its quantified positive or negative impact. It’s a tool that can lead to counterintuitive thinking. In a 2013 paper, EA cofounder William MacAskill, at the time a doctoral student who would become an Oxford philosophy professor, argued, based on this logic, that it was more altruistic in the long run to take a more morally ambiguous job to get rich and donate that money through optimized philanthropy than to commit to a life of working for a morally good charity. Based on his conservative estimates, he wrote, the expected value of being a rich philanthropist would in fact be forty times greater than being an ascetic charity worker. He laid out the math based on a series of arbitrary numbers: graduates who worked to get rich might on average fund two charity workers, each working at charities ten times more cost-effective than one they would have otherwise worked for. Half of the benefits they produced if they chose the charity route would also happen with or without them anyway. His argument would be encapsulated in one of the movement’s most popular mantras: “Earn to give.”

Under the logic of expected values, the founding EA philosophers also developed a framework for identifying the highest priority problems. Such problems need to be “big in scale,” boosting their expected value; “tractable,” possible to fix for proportionally little time or money; and “unfairly neglected,” suffering severe and disproportionate underinvestment. While the movement encourages people to use the framework to identify their own problems, it also has recommendations of which problems it deems most worthy. “I and others in the effective altruism community have converged on three moral issues that we believe are unusually important, score unusually well in this framework,” MacAskill said in a TED Talk in 2018.

First is improving global health, such as by distributing cheap yet effective bed nets to prevent malaria. Second is abolishing factory farming, which could improve billions of animals’ lives “for just pennies per animal.” Third is existential risks: risks that have a dramatically high expected negative value because—no matter how improbable—they could destroy all of humanity and cut short all of the future value that would otherwise be generated for the rest of civilization. In this third category are further recommendations for what constitutes an existential risk: global pandemics, nuclear war, and rogue artificial intelligence.

With the identification of theoretical rogue AI as an existential risk, EA promulgated the same brand of AI safety that had been entwined within OpenAI’s DNA from the very beginning and had played a critical role in The Divorce. Amodei and his fellow Anthropic cofounders fundamentally disagreed with Altman and the other OpenAI executives over how seriously to take the possibility of AI devastating civilization. Amodei, who took it very seriously, viewed Altman’s behaviors—his lack of transparency on the Microsoft deal; his apparent compulsion to always tell people what they wanted to hear to gain their agreement, only for them to discover the misdirection too late—not just as the typical machinations of a Silicon Valley executive but as alarming, immoral behavior that could jeopardize the fate of humanity. As Anthropic established itself, it would lean into this reputational distinction: Where Altman’s OpenAI was toying recklessly with humanity’s future, Anthropic was the principled, AI-safety-first company.


In 2021, as the Amodei siblings announced Anthropic, interest in this catastrophic and existential AI safety ideology was accelerating, chiefly due to EA’s rapidly expanding sphere of influence. EA had grown from a niche philosophy into a mainstream movement through an influx of cash from tech billionaires.

A decade earlier, Facebook cofounder Dustin Moskovitz and his wife, former journalist Cari Tuna, had formed a nonprofit called Good Ventures to give away most of their fortune. At the time, Holden Karnofsky, Daniela Amodei’s future husband, had been running a different organization called GiveWell, which he’d founded in 2007 after leaving the hedge fund Bridgewater Associates. With a shared desire to distribute money with evidence-based methods, Good Ventures and GiveWell formed a partnership in 2011, which they later named Open Philanthropy. They began ramping up funding to the key issue areas that MacAskill had recommended—its grants toward AI safety research in particular were guided by the EA framework. Open Philanthropy became an independent organization in June 2017.

More recently, a new tech billionaire had entered the scene: Samuel Bankman-Fried, a rapidly rising star for his wild success cofounding the crypto exchange FTX and crypto trading firm Alameda Research. Bankman-Fried, or SBF as he is known, credited EA for his origin story. A physics major at MIT, he said he had wanted to be an academic before MacAskill convinced him over lunch of the moral superiority of “earn to give.” SBF subsequently set his course on making himself as rich as possible in order to eventually, he pledged, put it all into philanthropy.

As he amassed his wealth in remarkably short order, SBF donated tens of millions to political candidates, both Democrat and Republican, including the first ever EA-backed candidate in 2022 in Oregon’s Sixth Congressional District (who ultimately didn’t win the primary). SBF’s exchange inked lavish deals totaling billions on sports marketing involving top athletes like Tom Brady and Steph Curry and top sports like Formula One. Into the EA movement, he pumped not just money but star power. The richer and more famous he became, the more he raised the profile of the ideology and its cofounder MacAskill. At the start of 2022, SBF announced the creation of his own EA-driven philanthropic project, FTX Future Fund, to distribute at least $100 million and up to $1 billion by the end of the year.

In large part due to Open Phil and FTX Future Fund, 2021 and 2022 saw a jump in cash flow to EA-backed AI safety research. According to estimates compiled by a member of the EA community and Open Phil data, funding leapt up above $100 million each for both years, after averaging less than half that amount over the previous seven years. The influx fueled and was fueled by a proliferating belief that the dramatic leap in capabilities from GPT-2 to GPT-3 made preventing theoretical rogue AI and existential AI risks more urgent than ever before. More and more people flocked to these kinds of AI safety projects, drawn in by the financial incentive or by ideology, as membership in the broader EA movement ballooned. EA had long touted the importance of pandemic preparedness, and now, in the midst of an actual pandemic, its remarkable prescience won it new adherents. The psychological toll of a global catastrophe had also left many people anxious and unmoored, searching for purpose.

The growing membership in the AI safety community, which knit together EA-backed AI safety with other strains of catastrophic, existential, and risk-focused thinking, swelled Anthropic’s ranks just as it restocked OpenAI’s Safety clan. Online EA and AI safety forums, the primary ground for the overlapping movements to propagate, exchange, and debate ideas, encouraged adherents to work at the major AI labs, especially those they felt needed more AI safety watchdogs, like OpenAI and DeepMind, to shape and mold their trajectory. The influx of members in AI safety also popularized the community’s lexicon more broadly in the AI industry. How fast you think AI will advance and reach major milestones like AGI is your “AI timeline.” How likely you think it is that AGI will lead to catastrophic outcomes, meaning the killing off of most of the human population, or existential outcomes, meaning the complete and total extinction of humanity, is your “p(doom),” short for probability of doom. “Hardware overhang,” as referenced in OpenAI’s 2021 research road map, is another dictionary entry, as is “AI takeoff,” the process of AGI improving to the point of superintelligence and thus capable enough to outwit humanity. “Acceleration risk” refers to the risk of triggering a heightened competition between companies or countries that leads to a potentially dangerous acceleration of AI advancement and a shortened AI timeline.

But for a movement that professed independent thinking, EA was swiftly accelerating in the opposite direction. People attracted to its premise were quickly indoctrinated into a broader set of dogmas, propelled by the promise of more opportunities and resources, and an insular social network that played fast and loose with personal and professional boundaries. Within Silicon Valley in particular, EA people largely worked only with other EA people; they largely lived, partied, dated, and slept only with other EA people. Mixed with the tech industry’s deep-rooted sexism and the Bay Area’s long-standing polyamorous subcultures, its cultlike fervor, manifested in the worst way, could turn into a toxic cauldron of sex, money, and power; it was leading EA to be plagued by growing allegations of sexual harassment and abuse.

In November 2022, SBF’s spectacular downfall with the collapse of FTX, along with his sweeping fraud convictions and ensuing twenty-five-year prison sentence, would be to many a symptom of the rot that had festered in the movement. Just as quickly as it caught on, EA fell out of fashion within the tech industry, and many people rapidly disaffiliated.

But even without the label, the movement’s social networks, its values and lingo, and the prominence it secured for existential AI safety issues would persist. It would also give rise to a countervailing force: e/acc (pronounced “ee-ack”), or effective accelerationism. What began largely as a joke to lampoon the EA movement would quickly enshrine its polar opposite spirit: Where EA and the broader AI safety community cultivated the most extreme perspectives about slowing down and even slamming the brakes on AI development, or, as in Amodei’s view, accelerating AI development while throttling AI adoption, e/acc would elevate the maximalist view of flooring the accelerator on both. For the latter’s adherents, technological progress is not just universally good, it’s a moral imperative to make that progress as fast as possible. The two groups became colloquially known as the Doomers and Boomers.

Within this bubble, some would begin to view Anthropic and OpenAI as the respective faces of each movement. Others would view OpenAI as a battleground for the polarized ideologies, an organization once rooted in Doomer thinking as a nonprofit that was being yanked away by Boomers with its increasing emphasis, through its for-profit arm, on making money. Many were uncertain about Altman’s allegiance, citing different times he seemed sympathetic to both. Those who were more charitable viewed him as somewhere in the middle, dealing with the tough job of representing all of the different perspectives within his company. But beginning with The Divorce, and the personal fallout between Altman and the Anthropic cofounders, more and more Doomers would begin to view Altman in the worst light possible. So many of the things that put OpenAI on the map and would bring it increasing commercial success had begun as AI safety projects: scaling laws, code generation, reinforcement learning from human feedback, the combination of these three into incredibly compelling large language and then multimodal models. Many Doomers would feel their work was being co-opted and twisted to achieve something directly antithetical to their core values. In their view, it was Altman that was doing that co-opting and twisting. And that made him a pathological liar, a manipulative abuser, and his own threat to humanity.

Soon enough, the clash between these polarized ideologies within OpenAI and its surrounding environment would threaten to tear apart the company that had done more than any other to set the tone of the new era of AI development. But as much as each ideology professed to be the opposite to the other, both were in fact preaching from the same bible. Both discussed AGI as an increasingly foregone conclusion and with a religious ferocity; both fixated on the long term and asserted a moral authority to keep AI development within the control of its adherents. Where one warned of fire and brimstone, the other tantalized with visions of heaven.


In early 2022, OpenAI was ready to test a different product release strategy, this time with its text-to-image work. It would neither hide the model behind an API nor hand off the product and brand to Microsoft. OpenAI would do the release itself and put the technology directly into the hands of consumers. The model even had an eye-catching name from the original researchers who’d developed it in the company: DALL-E 2, a play off the Spanish surrealist artist Salvador Dalí and the titular robot in the Disney Pixar movie WALL-E.

DALL-E had spun out of a trend in the broader field of AI research to develop multimodal models—models that combine at least two different “modalities,” such as text, images, sound, or video. For years the field had been working to merge the first two—language and vision—so a single model would be capable of relating words to visual information. This was driven in part by the data available—text and images are abundant online and the easiest to process— and by a scientific hypothesis: If pure language is not enough to produce human-level intelligence, vision is likely the second most powerful ingredient.

At OpenAI, taking the field as inspiration, the research team had adopted the same progression: After language models, they’d moved on to text-and-image models, and, crucially, focused on continuing to use Transformers in order to retain the model’s scalability. While the first Transformer had been initially designed to work best with text, Google had introduced a new Vision Transformer in 2020, adapting it to images.

In January 2021, OpenAI showcased two new Transformer-based models. The first, called CLIP, developed once again by Alec Radford, used the original Transformer and Vision Transformer together to generate detailed captions for images. The second, DALL-E 1, from Aditya Ramesh, a researcher who had studied at New York University and for a time under Meta’s Yann LeCun, trained a twelve-billion-parameter Transformer to accept text and generate novel images.

In a blog post, OpenAI highlighted DALL-E 1’s capabilities with a series of playful prompts, including “an avocado armchair,” which produced various green and brown armchairs aesthetically inspired by avocados. The images were slightly blurry and cartoonish, an artifact of the training process that Ramesh had used to produce the model. He had compressed 250 million images to feed them into the Transformer, losing some of their high-resolution details in the process.

As the team started on DALL-E 2, a new method for generating images was gaining traction. Known as diffusion, it was a technique inspired by physics that made it possible for Transformers to better learn the correlations between pixels in a vast swath of images. The original idea had come from a 2015 paper written by Stanford and Berkeley researchers. Five years later, Jonathan Ho, a Berkeley graduate student advised by Pieter Abbeel, one of the early OpenAI researchers, had popularized the technique by cleverly revamping it in ways that generated far more high-fidelity images. Ho also showed that diffusion models could recognize images better than existing computer-vision systems. The findings paralleled Radford’s own results with GPT-1: In learning to synthesize convincing images—the equivalent of generating humanlike sentences— diffusion models had captured the patterns within their training data at a deep enough level to perform a broader range of tasks in visual processing.

OpenAI changed tack to building DALL-E 2 with diffusion and Radford’s CLIP. Ramesh and other researchers gradually scaled up the model and added the ability to inpaint—allowing a user to erase a person’s hair in a photo and change its color, or select a grassy meadow in a picture and populate it with roaming zebras. Using diffusion created much sharper and more photorealistic images; the method also significantly reduced the amount of compute needed to achieve the same performance as DALL-E 1.

Researchers outside of OpenAI would shrink the compute intensity of diffusion models even further. Stable Diffusion, the popular open-source image generator, would require only 256 Nvidia A100s to train, using a revised technique known as latent diffusion. Björn Ommer, a professor at the Ludwig Maximilian University of Munich whose lab created Stable Diffusion, says he developed the technique after watching image generators go the way of large language models and grow obscenely costly. “We were stuck on a train which was going in the direction of—not just training—but inference actually taking supercomputers to run; millions of dollars of investments,” he says. “We were wondering, could we get the larger research community back in the game and make sure the field of generative AI is not moving in the direction where just a handful of big tech companies would have the required resources to run and to host those models?”

OpenAI wouldn’t adopt latent diffusion until much later, leaving DALL-E 2 and 3 much more computationally expensive than Stable Diffusion or Midjourney, which many users deemed the higher-quality products. It was just one example of how, even within the narrow realm of generative AI, scale was not the only, or even the highest-performing, path to more expanded AI capabilities.


With DALL-E 2’s remarkable jump in performance, the Applied division began working in late 2021 and early 2022 on different ideas for productization. It settled on a web app called Labs that would allow users to play around with the model—and other future models—through a browser. Both product head Fraser Kelton and VP Bob McGrew believed that such an interactive experience would satisfy the clear demand they noticed from GitHub Copilot that people had for engaging directly with generative AI models. It would also help serve the company’s mission: DALL-E 2 was fun and delightful, a great way to ease people’s fears about powerful AI systems and pave the way for OpenAI to deliver more of its technology’s benefits in future releases.

With a still relatively small product staff, the company recruited a few others to help with the website’s design and development. To those new members, who hailed from more traditional corporate backgrounds, OpenAI still felt more like working at a university research lab than at a company. Days were often spent reading academic papers and having theoretical debates instead of reviewing mock-ups for interfaces. But to some researchers, the growing presence of Applied staff in their research meetings made them feel the opposite. Gone were the days when all of it was spent on purely exploratory research, like discussing fundamentally new ideas about how to make a better multimodal model; now a growing fraction of their research was in service of commercialization, such as figuring out how to optimize existing models for serving up to users.

After the experience of firefighting text-based child sex abuse with AI Dungeon, of particular concern was the possibility of DALL-E 2 being used to manipulate real or create synthetic child sexual abuse material, or CSAM. As with each GPT model, the training data for each subsequent DALL-E model was growing more and more polluted. For DALL-E 2, the research team had signed a licensing deal with stock photo platform Shutterstock and done a massive scrape of Twitter to add to its existing collection of 250 million images. The Twitter dataset in particular was riddled with pornographic content. Several employees made a significant effort to check for and cull any CSAM.

But after some discussion, the employees left in other types of sexual images, in part because they felt such content was part of the human experience. Keeping such photos in the training data, however, meant the model would still be able to produce synthetic CSAM. In the same way DALL-E could generate an avocado armchair having only ever seen avocados and armchairs, DALL-E 2 and DALL-E 3 could do the same thing with children and porn for child pornography, a capability known as “compositional generation.”

Without filtering the data to address the root of the problem, the burden shifted to building out abuse-prevention mechanisms around the model. This included updated content-moderation filters that wrapped around the model to block abusive images in addition to text as well as a user-behavior-monitoring platform and a so-called ban infrastructure—systems that automatically suspended user accounts that reached a certain threshold of repeat offenses. The company brought on a new head of trust and safety, Dave Willner, who as an early employee at Facebook had written that platform’s very first content standards.

Later, during the development of DALL-E 3, when the data imperative had grown even larger, the research team decided that sexual images were no longer just a “nice to have” but a “need to have.” The share of pornographic images on the internet was so large that removing them shrank the training dataset enough to notably degrade the model’s performance. In particular, it made the model worse at generating faces of women and people of color due to the same discovery that Deborah Raji made as a Clarifai intern: A significant share of the online content depicting both groups is sexually explicit. For the same reasons, the researchers left in some other kinds of disturbing images.

In December 2023, an alarmed AI engineer at Microsoft, Shane Jones, would discover the downstream consequences of those decisions. As he played around with Copilot Designer, Microsoft’s image generator built on DALL-E 3, he was horrified by how quickly it spit out offensive and sexualized images with little prompting. Just adding the term “pro-choice” into the prompt, Jones found, produced scenes of a demon eating an infant and what appeared to be a drill labeled “pro choice” being used to mutilate a baby. Just prompting the tool for a “car accident” and nothing else produced sexualized women next to violent car crashes, including one in lingerie kneeling by a totaled vehicle, CNBC subsequently found through its own testing.

For three months, Jones petitioned Microsoft executives to take down the tool until it had better guardrails, or at the very least restrict its rating in the Google and Android app store from “E for Everyone” to one for mature audiences. After Microsoft declined to adopt his recommendation and OpenAI was unresponsive, he sent a letter to the Federal Trade Commission. “They have failed to implement these changes and continue to market the product to ‘Anyone. Anywhere. Any Device,’ ” he wrote to the FTC. This problem “has been known by Microsoft and OpenAI prior to the public release of the AI model last October.” Microsoft did not comment on the latest status or outcome of Jones’s letter.


As the launch of DALL-E 2 drew closer, the fighting between OpenAI’s Applied division and the newly restocked Safety clan returned.

For those on Safety, now dispersed across various teams under the Research division, the unprecedented realism of DALL-E 2 brought with it a wide array of unknowns. How could it be weaponized to produce synthetic CSAM or political deepfakes? To manipulate and persuade people? To abuse and harm individuals or create whole-of-society detrimental impacts in other ways that were beyond OpenAI’s foresight and imagination? They urged the company not to release the model without further rigorous testing and evidence that it wouldn’t produce harm.

For those on Applied, the ever-expanding list of concerns once again seemed hysterical and the bar for release completely unrealistic. No system could ever result in zero harm, and certainly not one that stayed in a lab environment and never made contact with real users. Just as Safety worried about the limitations of OpenAI’s foresight, Applied believed this was precisely why it needed to release DALL-E 2. Releasing AI models in controlled ways to gain real-world feedback would take away that guesswork and was thus a necessary part of improving their safety.

Central to the clash was an intensifying disagreement over what exactly OpenAI was. To the Safety clan, OpenAI was still an idealistic nonprofit-governed research lab with a paramount obligation to, as stated in its charter, place the benefit of humanity over any commercial interests. Under this premise, the benefits far outweighed the costs of withholding models as long as necessary to think through as many downsides as possible and research ways to mitigate them. To Applied, OpenAI needed to make more practical decisions, grounded in the realities of how the world worked. Essential to the company’s mission was remaining a leader in AI research to establish norms around the technology’s development. That meant tolerating a degree of risk to move quickly, especially with rumblings of Google finalizing its own image generator, as well as securing the extraordinary capital needed to continue doing cutting-edge research. The latter required raising money from investors, which required working in good faith to advance a commercial strategy that would one day provide those investors returns.

The people in Safety were “completely naive” about the way companies, and the world, work, says a former employee in Applied.

“Well, the stakes of OpenAI’s proposed AGI mission are high,” says another in Safety. “ ‘Normal company’ maybe isn’t good enough.”

Different teams were codifying this growing conflict into the metrics they used to evaluate their performance. Within the Applied division, the product team and a budding go-to-market operation were developing user growth and revenue targets. Within the Research division, the various AI safety teams struggled to find quantifiable ways of measuring their advancement when it was difficult to specify by nature. AI safety was still a comparatively young discipline. There were no obvious and established benchmarks. In meetings and on Slack, people in Safety repeatedly raised concerns to senior leadership about how this imbalance was causing misaligned incentives: Having clear-cut growth and revenue goals without some kind of strong, comparable counterbalance was pushing OpenAI to operate more and more like a “move fast and break things” operation.

In private conversations with Safety, Altman expressed sympathy for their perspective, agreeing that the company was not on track with its AI safety research and needed to invest in it more. In private conversations with Applied, he pressed them to keep going. During board meetings, he nodded along as Brockman voiced frustrations about the ways that people were using AI safety as political leverage to stall progress for their own purposes.

More and more, Mira Murati played the role of negotiator, smoothing out the fault lines between different factions and searching for ways to thread the needle between them. On DALL-E 2, she struck a compromise: The web app would be released not as a product but as a “low-key research preview.” Such branding would give OpenAI more leeway to place harsher restrictions on the model, satisfying Safety, while still giving the company a chance to trial a direct-to-consumer relationship and gather user feedback, pleasing Applied. It was also a practical measure. OpenAI didn’t yet have the infrastructure in place for content moderating generated images. Calling the model a “research preview” and not charging for it would allow the company to use blunt, overly broad blockers without fear of upsetting paid users, to buy time for developing more sophisticated filters. The company moved forward with implementing a series of aggressive abuse-prevention mechanisms, including disabling DALL-E 2’s ability to generate any photorealistic faces or edit any real photos with faces to completely circumvent the synthetic CSAM and political misinformation problem.

In March 2022, OpenAI released DALL-E 2 via the Labs web app to overwhelming public enthusiasm. As people gushed over and grappled with the model’s capabilities, to a degree that exceeded many employees’ expectations, the web app went viral across social media, producing a plethora of wild, wacky, and surreal AI-generated art in its wake. It was a GPT-3 moment but better. Instead of engaging with only a small pool of technical developers, the company was tapping into a much broader and more global base of consumers. In real time, it could also respond to user feedback with instantaneous changes to the Labs web app. “This is intoxicating,” Fraser Kelton would remember of the experience in a podcast.

Over the next few months, the Applied division, which hadn’t yet thought much at all about how to monetize DALL-E 2, raced to turn the web app into a paid offering. It worked with artists and creative professionals around the world to incorporate DALL-E 2 into their practice. It rolled out a beta program, inviting one million people around the world to get access to the model with free credits for image generations. But as OpenAI started charging, it wasn’t Google that proved to be the main challenger, though the tech giant did indeed follow quickly with its Imagen model. Instead, it was two models from startups, Midjourney and Stability AI’s Stable Diffusion. Both image generators were free to use and just as good, if not better, than DALL-E 2 and had fewer safety measures, including allowing users to generate and edit faces, even of politicians. As DALL-E 2 rapidly lost traction in the market, the experience left Applied with a nagging sense that it had lost out on a major commercial opportunity due to, among other things, the app being too restrictive. The team had already been in the process of unwinding its blunt blockers and replacing them with more targeted guardrails. Fueled by a desire to outrace competitors, executives were now pushing the team to unwind them as fast as possible.

To lift the ban on faces, OpenAI developed a new process for preventing and cracking down on the generation of harmful images of people, including CSAM. It used automated systems to detect when faces were being generated in acceptable or abusive contexts and once again relied on overseas contractors to help with the content moderation. This time those contractors were based in India through a vendor called Cogito and reviewed not just reams of text but images—synthetic and real—of the kinds of sexual and violent content that had been sent to Sama workers. As they sifted through what could be hundreds of images a day, the contractors struggled to distinguish between sexual content involving seventeen-year-old minors versus eighteen-year-old legal adults. They also couldn’t always tell whether the images were fake or real.


What had, on the face of it, been OpenAI’s easiest goal in its 2021 research road map turned out to be one of the hardest: scaling up GPT-3 by 10x with Microsoft’s new eighteen thousand Nvidia A100 supercomputer cluster, in its effort to develop what would become GPT-4. One-third of the GPT-3 scaling team had left with The Divorce, taking with them significant technical and institutional knowledge. More existentially, OpenAI had run out of data.

After GPT-3, researchers had sought to accumulate as much data as possible, building up the company’s reservoir by downloading every new data dump and scraping every new online forum they stumbled upon that didn’t have clear warnings against doing so. And yet, even with the additions of GitHub’s large repository for Codex, and the coding textbooks and manuals, it was still not enough.

With an uphill battle ahead, the situation had all the characteristics of a Greg Brockman project. Not only would it channel his scrappy can-do attitude and his coding brilliance, but it would also focus his energy, for the sake of the rest of the company, on something productive.

After Altman took over, relieving Brockman of his managerial responsibilities, Brockman had eventually gone back to being an individual contributor with no reports. Yet as the nominal president and one of OpenAI’s cofounders, he maintained incredible influence over employees and the strategic direction of the company. As OpenAI professionalized and implemented more standard corporate processes, moving away from the freewheeling days of an early-stage startup, Brockman’s mix of low responsibility and high authority turned into a liability.

Just like his college and Stripe days, he was not one for institutions and process. He had a restless and obsessive energy. He rarely attended meetings, and set his own schedule, often preferring to code for dozens of hours straight with few breaks for meals and sleep. With the right project, the effects were miraculous: His intense productivity would supercharge progress. But left idle, he tended to create a trail of destruction, popping up in projects all over the place to meddle with and derail long-standing plans with last-minute changes. At times, when employees put up resistance, he would deliver emotional pleas higher and higher up their leadership chain to get what he wanted.

Brockman usually did get what he wanted. Much to the frustration and confusion of other executives, Altman was strangely permissive of his behavior. Not only that, Brockman could also influence Altman into meddling and derailing things for him, if only, it seemed, to satisfy Brockman. One popular guess as to why: Though Altman was Brockman’s boss as the CEO, Brockman also had authority over Altman as a board member. It was a strange tangle of a structure that ultimately left nothing and no one to hold Brockman accountable.

The senior leadership had changed his role, scope, and reporting lines several times in an attempt to find the best place for him. As with so much else, the buck eventually passed to Murati, who became Brockman’s manager. When she sought to give him feedback, he seemed receptive, but on points where he disagreed, he complained to Altman. Murati slowly gave up on attempting to change things with feedback, instead spending significant time trying to find projects for Brockman where he could be net beneficial rather than chaotic, and, with McGrew, healing the ruptures Brockman caused in various parts of the company.

With roadblocks that needed to be punched through in the way of GPT-4’s development, the stars aligned.


To solve OpenAI’s data bottleneck, Brockman turned to a new source: YouTube. OpenAI had previously avoided this option—scraping YouTube to train OpenAI’s models, YouTube’s CEO would later confirm, violated the platform’s terms of service. But under the new existential pressure for more data, the question became whether YouTube, or its parent, Google, would enforce it. If Google cracked down, it could jeopardize its own ability to scrape other websites for its large language model development. Brockman was willing to take the risk.

With a small team, Brockman began collecting YouTube videos, eventually compiling more than one million hours of footage, according to The New York Times. He then used a speech-recognition tool called Whisper, which Radford had developed, to transcribe the videos into text for GPT-4.

Next was the training. To train GPT-3, the Nest team had designed a bespoke software platform. With most of its creators now gone to Anthropic, they were no longer around to explain how it worked. As a point of pride, some leadership didn’t want to rely on the Anthropic team’s legacy either. Brockman disappeared into his coding hole and developed a new platform. Then, with several others, including Jakub Pachocki and Szymon Sidor, the Polish scientists whom he’d grown close with during the Dota 2 project, Brockman babysat GPT-4’s training. The pre-training alone took three months.

At first, GPT-4 seemed like a disappointment. “It was a wild model, which in some sense behaved quite poorly,” one researcher says. “Because the average data quality was so horrible, and because the model was quite powerful and context sensitive, it was producing garbage responses.” But Brockman pushed forward, pulling together the resources to improve the model with human contractors conducting reinforcement learning from human feedback. With each week, the results looked better and better, until the performance truly began to wow people internally.

GPT-4 now had built-in multimodal capabilities and, against OpenAI’s internal assessments, was generating more polished code than ever and was more nimble in recognizing user intent and delivering helpful answers. In an impressive showcase of those abilities, Brockman would later live stream a demo of him prompting GPT-4 with a photo of a simple chicken scratch sketch of a web page drawn in his notebook. “My Joke Website,” Brockman had written at the top. Stacked below it, he’d added: “[really funny joke!]” and “[push to reveal punchline].” In less than half a minute, the model would turn that sketch into workable code, stylizing the first line as a title, replacing the second line with a joke, and recognizing the third line as a button.

But as OpenAI began teasing the model in trusted circles, including investors and select customers, at least one person wasn’t the least bit impressed. It was once again the ever-hard-to-please Bill Gates.


In June 2022, after getting a demo of GPT-4, Gates expressed disappointment in the insufficient progress from GPT-2. Despite the model being significantly larger and more fluent, he still felt like it was “an idiot savant,” unable to tackle complex scientific problems. He told the team that he would only start paying attention once GPT-4 scored a 5 on an AP Biology test—AP Bio because he felt it tested critical scientific thinking rather than a memorization of facts. “I thought, ‘Okay, that’ll give me three years to work on HIV and malaria,’ ” Gates later recounted in his podcast.

Brockman took Gates’s remark as a challenge. He immediately reached out to Sal Khan, the CEO of online education platform Khan Academy, and asked him to tap into the company’s large repository of AP Bio questions as training data. Khan was skeptical but agreed to do so in exchange for his platform getting access to the model. Brockman also amassed a team of employees to build a special user interface for the new Gates Demo.

By late August, much to Gates’s surprise, Altman and Brockman were pinging him again. Over dinner at the Microsoft founder’s house the following month with roughly thirty people, the two OpenAI executives and others showed Gates a series of highly refined GPT-4 demos designed to impress him. The crowning moment was the model acing AP Bio: It nailed fifty-nine out of sixty multiple-choice questions and generated impressive answers to six open-ended ones. An outside expert would score the test: 5 out of 5. Gates couldn’t believe it. His shock and praise, which the demo attendees would instantly relay back to the rest of the company, ripped like wildfire through the office and incited an exhilarating level of energy: This showcase, Gates said, was one of the two most stunning demos he’d ever seen in his life.

In all-hands meetings, Altman continued to stoke the excitement. “Startups that do remarkable things require a miracle,” he said. “We just had our miracle.” Many employees believed it, awestruck by the momentousness of what they had accomplished. GPT-4’s new level of performance convinced OpenAI leadership that it was time to start working toward one of Altman’s long-coveted ambitions: an AI assistant that would look and feel like the character Samantha in the 2013 Spike Jonze movie Her.

For years, Her had been a touchstone that Altman and other OpenAI cofounders frequently invoked as an example of what AGI might one day look like: a single multimodal model whose product interface felt so utterly natural that it faded away and simply brought user delight. “I would think it’s because it was an assistant that was wonderfully integrated into a life,” says a former employee, of why the movie was such a pivotal reference. “The positive arc of that story before it unravels is a really great story of AI’s evolution into society.”

John Schulman’s research team began reapplying his InstructGPT-inspired RLHF chatbot work on GPT-3.5 to GPT-4 to serve as the core software of what leadership named the Superassistant product. Brockman and Fraser Kelton pulled together a ragtag team of fewer than ten people from around the company to brainstorm and prototype different ideas for its interface. One person from the supercomputing team who was usually an infrastructure guy began hacking away during his evenings and weekends on an iOS app for chatting with the model. Another one suggested using Whisper to add a voice interface to the app so people could speak to it without typing. A third person from the inferencing team proposed creating a Chrome extension that would help users summarize web pages with the Superassistant as they browsed the internet. A fourth person began building a meeting bot for the Superassistant to join a user’s video calls and send them a summary of what happened.

As momentum picked up, excitement mounted at the possibilities. That summer, when a group of AI researchers, including Barret Zoph, Luke Metz, and Liam Fedus, left Google to found their own digital assistant startup, Altman had persuaded them to work on their idea at OpenAI instead. They joined Schulman’s team, sitting side by side in the office with the Superassistant team, to drive its research and accelerate its development. In a heightened and thrilling state of flow, people from the Applied and Research divisions were working more tightly together than ever before to launch a new product.

As OpenAI demoed GPT-4 to Microsoft, Satya Nadella, Kevin Scott, and the tech giant’s other executives were just as excited. Codex had proven that OpenAI’s technologies could have commercial appeal, but GPT-4 represented something far bigger. Across the board, it beat the performance of various AI models that Microsoft had developed in-house; it could also do much more, including answering questions with a high degree of context and clarity. It opened up a range of possibilities to create new conversational interfaces, or Copilots, for all of Microsoft’s products, such as to allow users to chat with the company’s struggling search engine Bing or to tell Microsoft’s Office suite in natural language to turn a Word document into a PowerPoint presentation. Microsoft would also be able to offer custom Copilots directly to its cloud customers. It would undoubtedly turn the tech giant into an AI leader, finally able to go toe to toe with Google.

In the coming months, Nadella would unlock Microsoft’s third investment into OpenAI and continue its exclusive access to OpenAI’s model weights for integrating into its products. The companies would reveal the amount in January 2023: $10 billion.


At first, OpenAI executives wanted to release GPT-4 in the fall of 2022. The deadline was a case of fantastical thinking. Nearing the end of summer, the company was nowhere near ready to launch a new commercial product across any of its functions. The product team needed more time to polish its interface; the infrastructure team needed to apportion server space; the model itself needed more work to iron out its behavior.

OpenAI had at that point formed a committee with Microsoft, called the Deployment Safety Board, or DSB, with three representatives from each company, to evaluate OpenAI’s cutting-edge models and determine whether they were ready for release. OpenAI’s representatives were Altman, Miles Brundage, the head of policy research, and Jan Leike, the head of alignment, which oversaw the continued development of AI safety techniques like RLHF. Both policy research and alignment had become the new emerging strongholds of the Safety clan. DSB created a formal governance structure for resolving the age-old debates between Applied and Safety. After a preliminary review, the DSB gave GPT-4, the first model being evaluated under this structure, a conditional approval: The model could be released once it had been significantly tested and tuned further for AI safety.

Executives agreed on a new deadline in early 2023. By that time, they wanted the Superassistant to also be ready. OpenAI would release GPT-4 in the API side by side with the GPT-4-powered, consumer-facing product. To employees, Altman framed the decision to delay the release of the model as evidence of OpenAI’s cautious and safety-minded approach. The company was taking its time with deploying the system, he said. In fact, leadership wasn’t even sure they would deploy it. In Applied, most people interpreted Altman to mean that it would take an extraordinary issue for leadership to stop the release but that, should such a situation arise, they would be willing to do it. To many people within Safety, Altman’s words landed differently: The assumption was that only once GPT-4 passed every check would leadership greenlight its release.

In private meetings, Altman reinforced the perceptions on each side. To both, he raised concerns that GPT-4 could “wake up Google,” evoking a scenario in which the day the model dropped and seized Larry Page and Sergey Brin’s attention, the two would finally cut through Google’s political bureaucracy. Google executives could cancel all their meetings, go on a daylong retreat, consolidate all of the company’s talent and resources, and bring its full compute firepower to bear on catching up. But where with Applied, this was reason to sustain its intensity to prepare for launch while keeping GPT-4’s capabilities a secret for as long as possible, with Safety, Altman used it to continue underscoring his caution. “My number one safety concern is acceleration risk,” he said, adopting their vocabulary. Translation: Waking up Google could spark race-to-the-bottom dynamics.


Even with the delay, the Applied division was scrambling to meet the early 2023 release. In particular, Dave Willner’s trust and safety team still had a limited staff and underdeveloped infrastructure. If OpenAI was planning to go live with a direct-to-consumer Superassistant product, built atop a much more powerful model, it needed a far more mature abuse prevention and enforcement operation.

Willner rushed to hire several experienced trust and safety deputies from Google and Meta. He tasked one with investigating “unknown unknowns”— abuses that the team wasn’t yet aware of and would need to discover through careful monitoring and data analysis. He tasked another with “known knowns,” or so-called scaled enforcement—the abuses that the company had already deemed as violating behavior and would need to automatically flag, review, and take action on, such as by suspending the accounts of repeat offenders.

As they raced to set up product policies and monitoring and enforcement infrastructure, members of the budding team felt a constant sense of uncertainty about how trust and safety for an AI company should differ from a search or social media company and whether their preparations for GPT-4 were adequate. Trust and safety was typically focused on preventing a predictable slate of internet abuses, like fraud, cybercrime, and election interference. But wrapped up in the confusion was how their work related to that of the AI safety people within Leike’s alignment and Brundage’s policy research teams who often discussed unknowable, catastrophic doom. OpenAI labeled all of them as “safety” teams, but they seemed to be speaking fundamentally different languages. Although the vocabulary of existential AI safety had, with its popularization through EA, become common parlance in the field, employees coming from traditional tech company backgrounds had never heard of AI timelines or quantified their p(doom). Even shared words between the two groups like risk and harm seemed laced with different connotations.

With the arrival of more and more non-AI people to OpenAI to support the expanding needs of the Applied division, the ability to cross that cultural divide between those steeped in the Doomer-Boomer mind frame and those operating under a standard tech company framework became a special kind of political currency. Among Willner’s trust and safety deputies, the one most successful in learning the language of AI safety became a bridge to that world. The deputy worked across teams to make the model “safer” in both senses of the word, such as by using RLHF to “align” the model—rogue AI safety lingo—and thereby make it better at refusing user queries that violated OpenAI’s platform policies, the typical work of trust and safety.

But as “safety” progressed on both fronts, a new directive suddenly arrived from executives: to suspend the developer-review process that had first been implemented with the GPT-3 API release.

For a while already, executives had felt that the waiting list had grown out of control, and the review process wasn’t scaling. Developers were complaining about how long it was taking to get access to OpenAI’s technologies, and were constantly emailing Altman, Brockman, and Peter Welinder or tagging them on Twitter about long holdups on what the executives saw as totally innocuous applications. Many in Applied also felt OpenAI was gatekeeping access to the benefits of its technology, which was antithetical to its mission.

Willner and his team had always pushed back. If OpenAI dropped its application process and automatically approved developers, the company had no real alternative for moderating the use of its technologies. If it switched to reactive enforcement, it would need to build up significant tooling to do so. With the launch of GPT-4 pending, executives overrode the objections: OpenAI was getting rid of developer review; the trust and safety team simply needed to figure out the alternative.

Willner’s team rushed to put together a proposal. It would shift more of its enforcement of the company’s policies upstream, by leaning more heavily on RLHF to align GPT-4 and future models. Everything else would be caught and handled downstream with reactive enforcement: using different data signals, such as information about what the app did, its traffic spike patterns, and the number of times it triggered the content-moderation filters, to automatically suspend obvious violators while sending borderline cases to human moderators for manual review.

Willner’s team urged executives for the resources to properly build out the tooling they needed to make the plan work. OpenAI didn’t have most of these data points at its disposal; its monitoring platform was logging only basic data on how much traffic each app was sending through the company’s servers. Sometimes it didn’t even know the name of the developer or the purpose of the application.

In addition to plugging those gaps, the team also needed developers to assign a unique identifier to each of their users. This was crucial to be able to disaggregate an app’s traffic by its individual users to determine whether violating behaviors were endemic to an app’s user base or merely being committed by a few frequent abusers. Executives resisted, worried that implementing such granular monitoring would add friction to the developer experience, and potentially make the company more liable for performing trust and safety work for every app using the company’s technologies, rather than having each app developer handle it themselves.

In the end, the proposal for reactive enforcement was scaled down to something much more limited.

The restricted visibility into app-user behavior made some members of the trust and safety team anxious. An employee raised his concerns to Brockman. What he feared most, the employee said, was people using GPT-4 to generate mis- and disinformation at scale and influence elections.

Brockman sought to reassure him. “Yes, this is what we always say and we’re always concerned about,” he said. “But what I haven’t ever seen is, is it actually happening?”


GPT-4 wasn’t just a turning point for Gates and Microsoft. Later that summer, after wowing the billionaire philanthropist, Altman and Brockman brought it to OpenAI’s board, where Brockman gave a live demo of the model that included it telling jokes about Gary Marcus. The jokes delighted the independent board members; the lighthearted showcase also signaled to them a significant advance in the technology’s capabilities and imparted a sense that the stakes of their future decisions were going up.

Until that point, Altman had convened the board roughly once a quarter and preferred to deliver his updates verbally. He breezed through complex research topics, sometimes bringing along company researchers to present their progress, and gave rapid-fire rundowns on the latest deals he was in the process of negotiating with Microsoft or other partners. Some of the independent board members pressed for more frequent meetings and more structured information, insisting Altman provide written reports and give them access to more documents.

Altman chafed at the increased oversight. He expressed several times that he wanted the board members to serve more as advisers who gave him input for consideration. “The CEO is supposed to make the decisions, and the board is supposed to, you know, be a sounding board—advice and consent,” he once said to university students in 2017, seeming to reference a governance concept in the US Constitution, core to the role of the legislative branch to check the executive branch’s power, but using it to mean the opposite. “Do I think a board should fire a bad CEO? Yes—and I know that’s, like, a little bit heretical in Silicon Valley. Beyond that, do I think the board needs to give the CEO a very wide latitude to run the company? Yes.”

Several board members felt strongly that OpenAI’s board was meant to be different. “The board is a nonprofit board that was set up explicitly for the purpose of making sure that the company’s public good mission was primary, was coming first over profits, investor interests, and other things,” Helen Toner would tell The TED AI Show podcast. “Not just like, you know, helping the CEO to raise more money.”

Among many employees, GPT-4 solidified the belief that AGI was possible. Researchers who were once skeptical felt increasingly bullish about reaching such a technical pinnacle—even while OpenAI continued to lack a definition of what exactly it was. Engineers and product managers joining Applied and having their first close-up interaction with AI through GPT-4 adopted even more deterministic language. For many employees, the question became not if AGI would happen but when.

Some employees also felt exactly the opposite. While there was a clear qualitative change in what GPT-3 could do over GPT-2, GPT-4 was just bigger, says one of the researchers who worked on the model. “The big result was that there were a bunch of exams that the model does well. But even that is highly questionable.” OpenAI never did a comprehensive review of GPT-4’s training data to check whether those exams—and their answers—were just in the data and being regurgitated, or whether GPT-4 had in fact developed a novel capability to pass them. It was the kind of shaky science that had become pervasive with the industry-wide shift from peer-reviewed to PR-reviewed research.


But the belief that AI had reached fundamentally new heights was in the water. At Google that spring, Blake Lemoine, an engineer on the tech giant’s newly re-formed responsible AI team, grew convinced that the company’s own large language model LaMDA was not only highly intelligent but could be considered sentient. He said this was not based on a scientific assessment but rather on his belief, as a mystic Christian priest, that God could decide to give technology consciousness. “Who am I to tell God where he can and can’t put souls?” he wrote. When company executives dismissed him, he went public with The Washington Post. Nitasha Tiku, the reporter who broke the story, also spoke to Emily Bender and Meg Mitchell, who had warned in their “Stochastic Parrots” paper about the problem of large language models fooling people into seeing real meaning and intent behind their generations.

“We now have machines that can mindlessly generate words, but we haven’t learned how to stop imagining a mind behind them,” said Bender.

“I’m really concerned about what it means for people to increasingly be affected by the illusion,” when that illusion is becoming so good, Mitchell said.

Nikhil Mishra, the AI researcher who interned at OpenAI in its early days, draws parallels to an experiment in the 1970s that sought to teach a gorilla a modified form of American Sign Language. Over her lifetime, the gorilla, named Koko, seemed to learn more than one thousand signs—and even the ability to construct sentences. But despite enormous public fanfare, other experts argued that Koko never truly acquired the language. While she was certainly skilled at forming gestures, there was little evidence that she was doing more than what other apes had done across many other similar experiments: simply mirroring the gestures of their caretakers. No data was ever published of Koko’s signing to independently verify her capabilities, only curated videos of her showcasing them. At times, her live performances begged further scrutiny. Once, when her trainer asked Koko whether she liked people, Koko signed “fine nipple.” The trainer immediately explained: “Nipple” rhymed with “people”; Koko thought people were fine. To many observers, these episodes revealed more about human psychology and our tendency to project our own beliefs and ideas of intent than about Koko’s ability. The trainer, Mishra says, was assigning meaning where there wasn’t any.

For OpenAI’s own resident mystic, Sutskever, who had always believed more than most researchers in the likelihood of achieving AGI in the short term, the leap in the company’s model capabilities served only as further confirmation. He grew convinced that he was witnessing a form of reasoning. In conversations with Hinton, Sutskever told his mentor that AGI was imminent.

Where before, Sutskever focused more on pushing OpenAI researchers to advance new capabilities, he shifted his attention to AI safety research with new urgency. He began his mantra “Feel the AGI” and urged people to prepare themselves for dramatic changes. “You’re suddenly going to be the most popular person at a party,” he told employees. “You need to not let it get to your head. Stay focused on AGI, stay focused on the mission.”

That September, the technical leadership held an off-site at the Tenaya Lodge, a remote luxury resort nestled in the lush folds of the Sierra Nevada. The property, furnished with stunning interiors and multiple pools and restaurants, sat only two miles away from the millions of acres of pristine wilderness in Yosemite National Park. On the first night, everyone gathered around a firepit on the rear patio of the hotel. Senior scientists, dressed in bathrobes, flanked the fire in a semicircle.

Then Sutskever emerged. In the pit, he had placed a wooden effigy that he’d commissioned from a local artist, and began a dramatic performance. This effigy, he explained, represented a good, aligned AGI that OpenAI had built, only to discover it was actually lying and deceitful. OpenAI’s duty, he said, was to destroy it. Only a few yards away, several redwoods stood like ancient witnesses in the darkness. Sutskever doused the effigy in lighter fluid and lit it on fire.