Confirmed manipulation of Llama 4 benchmark results

Meta's former chief AI scientist Yann LeCun confirmed in January 2026 that the published Llama 4 benchmark results were "fudged," according to reporting carried by WION. LeCun said Meta's team trained multiple checkpoints of the model and selected the highest score achieved on each individual benchmark, then presented the results as a single composite performance that no single version of the model had actually achieved. He called the practice one that "completely violated the principle of fair evaluation."
The controversy traces back to April 2025, when Meta submitted a version of its newly released Llama 4 Maverick model to LMArena. That submission scored high enough to place Maverick second overall on the public leaderboard. Once developers got access to the actually released version, it fell to 32nd place, exposing substantive differences between the submitted and shipped variants. Meta's VP of generative AI Ahmad Al-Dahle denied any manipulation at the time, attributing the inconsistency to "differing cloud implementations."
Meta pivots to paid proprietary AI with Muse Spark 1.1
Meta launched Muse Spark 1.1 on July 9, 2026, opening a public developer preview API at $1.25 per million input tokens and $4.25 per million output tokens, according to MarketScale reporting based on CNBC. New accounts receive $20 in free credits, and access is gated through a waitlist on Meta's developer portal. The release marks the company's clearest move yet into the paid, proprietary AI model market that Anthropic and OpenAI have built their businesses on.
The model was trained by Meta Superintelligence Labs under AI chief Alexandr Wang specifically for coding and agentic tasks because, Wang told CNBC, coding capability is foundational to building effective AI agents. Wang said the model is served on Meta's own infrastructure, keeping workloads off third-party cloud platforms and aggregators. He also indicated that an open-source variant is in development without giving a release timeline, while a more powerful model code-named Watermelon continues training. The current Muse Spark 1.1 was developed under the internal code name Avocado.
Industry pushback against restrictions on open-weight AI
More than 20 tech companies, including Nvidia, Microsoft, Meta, and Palantir, signed a letter urging policymakers to avoid "premature restrictions" on open-weight AI models that would "stifle competition or drive innovation overseas," according to Briefs.co. The pushback came as Meta's own Llama family has anchored the company's open-weight strategy, with Llama 4 Scout and Maverick available for download on Meta and Hugging Face, although not to customers based in Europe.
Box CEO Aaron Levie, one of the signatories, said in an interview that for U.S. companies to remain competitive, they must have access to the best technology regardless of its origin. He said the arc of the industry is toward more AI progress and lower cost over time. Google's head of AI Jeff Dean separately called distillation a "key technique for making smaller models more capable" that requires the frontier model as a starting point, and Nvidia has incorporated the technique into its Llama Nemotron family.
Chinese rival Kimi K3 raises open-source pressure on Meta
Moonshot AI's Kimi K3, with 2.8 trillion parameters, will be released on July 27, 2026, according to El País. The Chinese startup published public benchmark results placing the model ahead of Anthropic's Claude Fable 5 and OpenAI's GPT5.6 Sol. Like Meta's Llama family and DeepSeek's earlier models, Kimi K3 will be open-source, making it the largest open-weight model in the world and surpassing DeepSeek's 1.6-trillion-parameter flagship.
Former Meta product director Xiaowin Qu questioned on X how Anthropic would justify its Fable pricing when the best open-weight model exceeds the best closed-source one. White House advisor Michael Kratsios publicly accused Moonshot of using distillation to copy Anthropic's model on an "industrial scale," while Anthropic itself has banned distillation in its terms of service and described halting unauthorized distillation as a matter of national security.
Share this article







