How Dario Amodei’s “Big Blob Of Compute” Essay At OpenAI In 2017 Had Predicted Some Parts Of The AI Revolution

It turns out that Dario Amodei has been publishing essays long before he published “Machines of Loving Grace”.

Ahead of the release of his book The AGI Chronicles, tech journalist Kevin Roose has published an excerpt from “Big Blob of Compute” (BBOC), a 2017 internal research document written by Dario Amodei when he was a researcher at OpenAI. The document has never been published before. Roose calls it, in his words, the essay that started the AI race.

According to Roose, the essay laid out Amodei’s hypothesis that scaling up AI models was by far the best way to make them smarter. He says GPT-2 was an experiment to test that hypothesis, and that when it worked, OpenAI went all-in on scaling, setting off the race the industry is in today. Amodei, who now runs Anthropic,, wrote the document years before most of the world had heard of a large language model.

Here’s what the essay says, and where it looks prescient.

The hypothesis: a big, minimally structured blob

Amodei starts by saying the debate over whether raw compute or new algorithms matters more is badly framed. His own position is “definitely closer to the ‘raw compute’ side.” The BBOC hypothesis, as he states it, is that

creating any given intelligent behavior is mostly about providing a large, minimally structured mass of computational capacity, and then giving it shape and form via interactions with a rich environment and a training process that drives it towards behavior appropriate for the task at hand.

He argues that only seven things really matter: compute, parameters, quantity of experience, distribution of experience, normalization and conditioning, symmetries and shape, and the training target. Everything else, he says, “usually has a minor impact or is operating at too low a level of abstraction to be properly general.”

Clever algorithms, in his telling, mostly help by staying out of the way:

algorithms matter much more in a negative sense than a positive sense — it’s possible to come up with horrible architectures that totally block the flow of compute, but it’s really hard to beat the unstructured blob by all that much

He also makes a pointed claim about the innovations that did work. Many of the “algorithmic” advances of the preceding five years, he writes, were really simple methods for keeping numerical computation well-behaved, naming Adam, BatchNormalization and ResNets.

Snowflakes, and “networks want to learn”

To build intuition, Amodei offers what he calls the “snowflake model” of intelligence. Nobody makes a snowflake by assembling its intricate pieces. You get one by knowing the laws of physics, supplying enough raw material and a big enough chamber, setting the conditions correctly, and waiting. The training target, compute and parameters play the roles of physics, raw material and chamber.

He also credits a colleague with a pithy version of the idea: Ilya Sutskever expresses BBOC as “networks want to learn.”

The evidence he cited

The essay draws on “3 years of training AI systems and many years of studying the brain.” A few examples stand out.

Vision before 2012. Researchers built elaborate pipelines of edge detectors, contour integrators and segmenters. Amodei says those failed because they used too little compute, too few parameters, organized the computation too rigidly, and trained the pieces on the wrong objectives. Convolutional nets fixed all of that.

A Google experiment. He describes a project to replace the standard layer of a neural net with other kinds of machine learning models. The models performed slightly worse, and he eventually found that they were quietly reorganizing themselves into ordinary neural-net layers. “Other models want to become them,” he writes, giving an indication to his thoughts about AI consciousness long before they became mainstream.

Planning and hierarchy in RL. Model-free reinforcement learning, he suggests, may learn to plan internally if the network is simply allowed to iterate for many steps before acting, while explicitly model-based systems perform relatively poorly because computation is forced into two disconnected blocks. Hierarchical RL, meanwhile, keeps collapsing into what a single deeper network would have done, which he describes as compute “routing around the structure imposed on it by a supposedly important algorithm.” That idea of giving a network room to think before it acts has an obvious echo in today’s reasoning models.

Generalization. Amodei argues that poor generalization is usually a data problem, not an algorithmic defect. Drawing on his speech recognition work (he notes that Deep Speech 2 used about 0.3 petaflop-days per trained model in 2015, which he believed was the largest single training use of compute at the time), he writes that a model trained on one accent does poorly elsewhere, but “if you train on 5 or 6 accents you will immediately do well on an unseen accent.” His summary: “it’s not algorithms that lead to generalization, so much as the training setup and environment.”

Adversarial examples. He isn’t surprised that classifiers are fooled by engineered images, since they were never trained on such images. “BBOC says you get what you optimize for; what you don’t get is a human’s idea of what properties the system should have.”

What aged well

The most striking part of the essay is how closely the later history of the field tracks its central bet. GPT-2, the model Roose says was built to test the idea, was a 1.5 billion parameter model described as a direct scale-up of GPT with more than 10x the parameters and data, and Amodei was among its authors. He was also a co-author of OpenAI’s 2020 paper on scaling laws for neural language models, which showed that performance improves predictably as model size, data and compute grow.

The “big blobs” are now being built at a scale the essay’s author could hardly have pictured in 2017. PwC recently estimated that global spending on data centers and the computing equipment inside them could reach $31.6 trillion through 2050.

The essay’s wariness about handcrafted structure and its emphasis on the training target also look familiar. It uses reinforcement learning from human feedback as an example of a training target built from several learning processes, a technique that went on to become central to how chat assistants are tuned.

What it did not claim

It’s worth remembering that Amodei hedged heavily. He called it “not a precise hypothesis or one I’m super confident in,” and he stressed that

none of this implies that scaling up today’s exact algorithms will lead to AGI. It does suggest that any algorithmic changes will likely be simple things that exploit symmetries or sparsity

In other words, the essay predicted that scale would be a huge part of the story, not all of it. Its bet against model-based and hierarchical RL, which it said were starting to look “increasingly like the edge detectors of 2011,” also isn’t fully settled.

The scaling debate itself has moved on, too. Sutskever, the man behind “networks want to learn,” argued in a late-2025 interview that the “age of scaling” is giving way to an “age of research,” while clarifying that scaling current systems can still yield improvements.

For now, the essay survives as a historical artifact: a compact statement of the idea that drove the most expensive technology build-out in history, written before anyone knew whether it would work. Roose’s book arrives tomorrow.

Posted in AI