Meta launches Llama 4-powered Seller app for Facebook Marketplace

According to TechTimes, Meta launched a standalone iOS Seller app that uses Llama 4 multimodal AI to generate Facebook Marketplace listings from a single photo. The free app produces a title, description, suggested price, and category without the seller typing, and is available immediately to U.S. users aged 18 and older, with Android and web versions in testing. Tom Alison, Meta's head of Facebook, announced the launch as Facebook Marketplace marked its tenth anniversary. The pricing engine surfaces comparable active listings in the seller's local market in real time rather than pulling historical sold comparables, a structural advantage over third-party crosslisting tools that rely on older sold data. The app's vision encoder is built on Llama 4, Meta's April 2025 multimodal model, which uses an "early fusion" architecture that processes text and image tokens together from the first layer, replacing the separate DeepText and Lumos systems that handled Marketplace listings in 2018.
Multiverse Computing ports CompactifAI Llama 3.3 70B to Intel Xeon 6
Multiverse Computing announced from San Sebastián, Spain, on July 23, 2026, that its CompactifAI-compressed version of the Llama 3.3 70B model now runs on Intel Xeon 6 processors with Performance-cores, using vLLM CPU and Intel Advanced Matrix Extensions (AMX). In benchmarks on an Intel Xeon 6737P processor, the compressed model achieved an output throughput of 3.86 tokens per second versus 2.00 tokens per second for the uncompressed baseline, a 93.6% improvement, while total token throughput rose 94.1% to 7.81 tokens per second. Latency at one concurrent user fell 48.6% to roughly 2,598 seconds, and inter-token latency dropped 48.9% to 252.08 ms. Across BoolQ, GSM8K, HellaSwag, MMLU, and WinoGrande benchmarks, the compressed model retained over 97% of baseline accuracy; WinoGrande scores rose 6.86% after a "healing" retraining phase. The largest gains appeared at the highest concurrency level tested, where 256 concurrent users saw throughput rise 107.0% and latency fall 51.7%.
Benchmark gaming allegations resurface around Llama 4 Maverick
An April 2025 report on Llama 4 Maverick resurfaced in recent coverage, detailing how Meta deployed a specially tuned "experimental chat version" of Maverick to LMArena, where it posted an ELO score of 1417 and the number-two leaderboard spot at the time. LMArena said Meta's interpretation of its policy "did not match what we expect from model providers" and updated its leaderboard policies in response. Meta spokesperson Ashley Gabriel said the company "experiment[s] with all types of custom variants" and released an open-source version for developer customization. Meta generative AI VP Ahmad Al-Dahle denied training on test sets, attributing performance variability to "needing to stabilize implementations," while CEO Mark Zuckerberg, asked about the Saturday release timing, said the model shipped "when it was ready."
Share this article







