Chip performance versus Nvidia Blackwell

Close-up shot of a smartphone screen showing the OpenAI website with greenery in the background.

OpenAI’s first-generation Jalapeño ASIC is reported to draw 700 watts and outperform a 1,400-watt Nvidia Blackwell-class flagship on throughput-per-kilowatt by a factor of 1.5× to 1.9×, with up to 3.6× lower latency. The figures come from the first published benchmarks of the chip and frame Jalapeño as a direct efficiency challenge to Nvidia’s top data-center GPU, the GB300.

Architecture and Broadcom co-development

The Jalapeño package pairs a single reticle-sized compute die built on TSMC’s N3P node with an N3E I/O chiplet and HBM4 memory, delivering 15.4 TB/s of memory bandwidth per package, according to SemiAnalysis reporting cited by Wccftech. Core and memory are partitioned into matching slices with a low-latency local view of HBM and a high-bandwidth collective network. OpenAI co-developed the chip with Broadcom and paired it with its own Gluon kernel programming language rather than relying on Nvidia’s CUDA stack.

Matrix engines and MXFP numerical formats

The matrix engine is built as a weight-stationary systolic array and uses MXFP numerical formats that compress AI math down to 4-bit (MXFP4) by letting grouped numbers share a single exponent. Weight-stationary data flow keeps model weights locked in the processing grid, MXFP shrinks the numbers being moved, and the systolic array passes data cell-to-cell without round-tripping to memory, a combination the report credits with eliminating data-travel bottlenecks. The design avoids speculative decoding tricks to inflate performance numbers, according to Wccftech.

Scalar and vector cores

Around the matrix engine, Jalapeño includes 64-bit out-of-order scalar cores with paired L1 caches to handle control code, memory management, and orchestration, and FP32/INT32 vector cores for high-precision steps such as softmax and normalization where MXFP compression would be unsuitable. The out-of-order scalar cores are described as streaming data into queues so that the in-order matrix and vector units always have operands ready.

Strategic threat to Nvidia’s CUDA moat

By building a custom ASIC and an accompanying kernel language, OpenAI is positioned to reduce its dependence on Nvidia hardware and software, a dynamic framed by Wccftech as a threat to Nvidia’s CUDA ecosystem. The Jalapeño results, drawn from SemiAnalysis, are framed as a first published benchmark rather than an Nvidia-validated figure, leaving room for independent verification of the 1.5×–1.9× per-kilowatt and 3.6× latency claims.

Share this article

FacebookX

3 sources

Sources