Summary
From the article:
For a group trying as hard as possible to make sure AI does not kill everyone (another shibboleth, this one significantly clearer), I expected them to be a lot more depressed and gloomy. This is far from the case. If you were to construct a world from first principles that provides for all the physiological, safety, belonging, esteem, and self-actualization needs of its inhabitants in a kind of Maslovian utopia, the solution would be roughly Berkeley shaped and centered around Constellation. If you believe you’re making progress on the most important problem in the world in a tight-knit and well-resourced community of intelligent coworkers who double as your closest friends, what more could you possibly want in life?
In some ways, Constellation seems less like an office and more like a strange idealized version of college. On an average Saturday morning, I read Borges and ’90s sci-fi about the philosophy of randomness and mathematical Platonism at Jacobian Juggler’s book club with nearly two dozen Stellies. Almost everyone seemed to have studied a STEM subject in college yet knew enough philosophy that “what’s your favorite ontology, and why?” was a perfectly acceptable icebreaker. The book club had nothing to do with AI, yet there was little desire to leave the social orbit of Constellation outside work. After the seminar-like discussion of a few short stories, most of them drifted back toward Constellation to keep on working.
This is not to say Constellation is stress-free. Numerous Stellies speak of the long hours and burnout involved in trying to keep up with the exponential pace of AI progress. And in the same breath, they’ll express how much they love it and feed off the energy of late nights and the meaning it brings. If taking more time off means, however marginally, that models deployed at massive scale are more dangerous, it becomes difficult to justify relaxing much at all.
[...]
From the outside, there is an obvious and unflattering comparison: the community looks millenarian and cult-like. They’re convinced that history is approaching a discontinuity after which ordinary human life will be either incomprehensibly better or permanently over and reorganize their careers and social lives around preparing for the shift. Indeed, coworkers become friends, friends become housemates and romantic partners, and the whole world becomes highly endogamous. This is not necessarily because anyone objects to dating outsiders; it’s more that the arguments permeate ordinary life so much that being with someone who is not “pilled” can become difficult. After all, if the most important follow-up question to “do you want to have kids?” is not “how many?” but “before or after ASI?”, you’re going to have issues dating the general populace.
[...]
The people and organizations of AI safety have had an impact on the development of AI that can be hard to appreciate if you aren’t steeped in the culture. Machine Intelligence Research Institute (MIRI) founder Eliezer Yudkowsky introduced future DeepMind cofounders Demis Hassabis and Shane Legg to Peter Thiel, who subsequently became the startup’s first major investor. ARC founder Paul Christiano helped architect modern reinforcement learning from human feedback (RLHF) while at OpenAI, work that turned pretrained language models into useful chatbots like ChatGPT. Anthropic itself was founded to be a safer and more aligned version of OpenAI, with much of its early funding coming from effective altruism-associated investors like longtime Coefficient Giving funder Dustin Moskovitz.
When the history of AGI is written, the organizations on those floors will be discussed alongside Anthropic and OpenAI in the way Oak Ridge, the Berkeley Rad Lab, and Los Alamos feature in the history of the Manhattan Project. Like Los Alamos, the population skews young, elite, and imbued with a sense that if they fail, everyone is thoroughly fucked.
[...]
Each organization is shaped around what is called the “theory of change”, or ToC. These are always AI safety flavored but often sharply diverging on what they believe is the most important way to prevent disaster. MIRI is the most doomerish of the lot, with its co-founder Eliezer Yudkowsky and president Nate Soares publishing the book If Anyone Builds It, Everyone Dies last year. They believe short-term attempts to make AI more aligned are futile and campaign for building “off-switches” and enacting policies to block ASI. The AI Futures Project instead tries to make possible futures legible enough for governments and other decision-makers to prepare for them, while groups like METR and Redwood focus more narrowly on measuring dangerous capabilities and controlling advanced systems.
If AI safety is a cult, it does an awful job of disseminating and enforcing orthodoxy because anything is up for debate. There is enough heterogeneity among ToCs in AI safety that some believe others’ projects in the movement are truly net negative (“in expectation”, I feel obligated to add), either by advancing AI in a way that improves dangerous capabilities more than safety, generating bad optics, or simply wasting money.
[...]
Even though I did not find myself believing in their strongest visions of doom, it’s not too hard to understand the appeal if you psychologically inhabit their world for enough time. There’s a sense you’re right in the middle of world events, with near weekly articles from major outlets featuring METR and Redwood reports. There’s a world of pleasure, community, and meaning in what most believe to be the “probable pre-apocalypse”. There’s a shared map of the many possible roads to hell and a common goal of ensuring we never arrive. There’s a heaven at the last stop before hell.