Summary

OpenAI and Broadcom have unveiled Jalapeño, OpenAI’s first custom-designed “Intelligence Processor” — an inference accelerator built from scratch specifically for large language model workloads. The chip went from initial design to manufacturing tape-out in just nine months, with OpenAI’s own AI models assisting in the design and optimization process.

Early testing indicates Jalapeño will deliver substantially better performance per watt than current state-of-the-art hardware. Broadcom CEO Hock Tan stated the chip could offer roughly 50% lower inference costs per token compared to current Nvidia GPUs. Engineering samples are already running ML workloads in the lab at production target frequencies, including GPT-5.3-Codex-Spark.

This marks the first step in a multi-generation compute platform the two companies are building together, with initial deployment slated for gigawatt-scale data centers alongside Microsoft by the end of 2026. Celestica is providing board, rack, and system integration, while Broadcom handles silicon implementation and high-performance networking.

Source

OpenAI Official Announcement | Tom’s Hardware | Quartz

Commentary

This is a significant strategic move that signals OpenAI’s intention to control the full AI stack — from models down to silicon. Until now, the AI industry has been almost entirely dependent on Nvidia’s GPU ecosystem for inference workloads. A purpose-built inference chip that achieves even a fraction of the claimed 50% cost reduction per token would fundamentally alter the economics of running AI at scale.

The nine-month development cycle is particularly striking and suggests AI-assisted chip design is no longer theoretical — it’s accelerating the hardware development timeline itself. If Jalapeño delivers on its promises at gigawatt-scale deployment, we’re looking at a potential shift in the AI infrastructure landscape that could affect pricing, availability, and competitive dynamics across the entire industry.

By Allan