Zuckerberg Blames Team Structure For Llama 4’s Failure, Says AI Models Needs Small, Tight-Knit Teams

Mark Zuckerberg has offered one of his most candid assessments yet of what went wrong with Llama 4, saying the problem wasn’t just the model, but the way Meta had organised the team building it.

Speaking in an interview, the Meta CEO traced the arc of the Llama family, from the model that kicked off the open-source AI movement to the release that left Meta behind the frontier.

“Llama 1 was quite interesting as a model, and it sort of pioneered the whole open source AI movement, which I think has been very powerful and is something that we’re very proud of,” Zuckerberg said. “Then Llama 2 scaled. Llama 3 was a very good model. It was almost at the frontier at the time. And then with Llama 4, we basically fell off the trajectory that we needed to be on.”

mark zuckerberg

“The whole shape of the team I had gotten wrong”

Zuckerberg said his instinct after a disappointment is to dig into the cause. “Whenever something doesn’t go the way that I think it should, I always spend a bunch of time thinking about why did that happen, and what do we need to change to make it better,” he said. In this case, his conclusion was that he had misjudged how the team should be built: “The whole shape of the team I had gotten wrong.”

The mistake, he explained, was borrowing a playbook that had worked elsewhere in the company. “I modeled it off of the way that we’d done our machine learning work for making Instagram feed or our ad system, these teams that have many hundreds or thousands of people who can work on a lot of stuff in parallel,” he said. “Whereas I think for building these language models, what you really want is just a very tight-knit team that views it as a group science project.”

That philosophy shaped what came next. A small team means each seat carries enormous weight, Zuckerberg said. “Every seat on that team is extremely valuable. So we ended up bringing a bunch of the best people from across Meta into the team, but also a lot of other awesome people from around the industry, and building a completely new team, which we called Meta Superintelligence Lab.”

The lab was unveiled in June 2025, headed by Scale AI’s former CEO Alexandr Wang as Meta’s chief AI officer, and accompanied by an aggressive hiring spree that pulled researchers from rival labs.

A rebuild, and a “negative surprise”

Zuckerberg acknowledged that the new organisation couldn’t deliver overnight. “By the time that we were starting MSL, we knew it was going to take some time to reboot and rebuild the infrastructure to build the next set of models,” he said. “But I knew that we’d pulled together a great team, and I knew that if we could gel the team and have that work well, then it would be good.”

The fallout from Llama 4 was not only technical. Questions about how the model was presented have lingered: Yann LeCun, Meta’s former chief scientist, later alleged that the Llama 4 team used different model versions for different benchmarks to make results look stronger than they were.

Zuckerberg described the aftermath of the launch as the most difficult point of the whole episode. “To me, actually, the scariest moment was after the Llama 4 launch, when it was just like, ‘Oh, I thought we were on this trajectory, and we’re not.’ It was a pretty big negative surprise.”

He framed it as a test of leadership. “When you’re an entrepreneur, you’re building something, you get tested in those moments, because inevitably not everything is going to go well,” he said. “The things that define the trajectory are, if something doesn’t go the way that you want, how do you figure out how to move forward?”

What changed after Llama 4

Meta’s answer has reshaped its AI strategy well beyond personnel. The company has since signalled it will be far more selective about which models it open-sources, a notable shift for the lab Zuckerberg credits with pioneering open-source AI. LeCun, a longtime open-source advocate, has gone further, warning that Meta is rethinking its approach to open models while the best open models in the field are now Chinese.

On the model front, the first major output of the rebuilt stack has been Muse Spark, which Meta says reaches Llama 4 Maverick-level capability with over ten times less compute, and which beats some frontier rivals on select benchmarks. Muse Spark 1.3, in particular, is extremely competitive, and has put Meta within the top 5 AI labs on the Artificial Analysis Intelligence Index.

For a company that once set the pace for open models, Zuckerberg’s account is a rare admission that the failure was one of organisation, and that fixing it meant starting over.

Posted in AI