Signal

Q&A with OpenAI VP of Hardware Richard Ho on its Jalapeño inference chip co-designed with Broadcom, using internal OpenAI models to design the chip, and more

First reported by Morethanmoore.substack ·

The signal ●●●○ Compiled by AI from Morethanmoore.substack, Techmeme and TechTechPotato on YouTube
Why you might care

If you are building or operating AI inference at scale, the cost and performance benchmarks for custom silicon have just been reset.

What happened

OpenAI has unveiled its custom-designed inference chip, codenamed Jalapeño, developed in partnership with Broadcom. This specialized chip is designed for inference tasks, not training, and features 216 GiB of HBM4 memory with a bandwidth of 15.4 TB/s. Jalapeño boasts a peak power draw of 700W, with sustained usage around 550W. A single domain comprises 128 accelerators, scaling up to a system of 2,048 accelerators, capable of 27 EFLOP/s at four-bit precision. At HotChips 2026, OpenAI presented performance metrics indicating Jalapeño offers 1.5x to 1.9x better performance per watt and 1.7x to 3.6x better latency compared to NVIDIA's GB200/GB300. Notably, OpenAI utilized its own AI models extensively in the chip's design process, from compute architecture to kernel optimization, reportedly achieving significant efficiency gains in a short timeframe.

What it means

OpenAI's entry into custom silicon with Jalapeño signals a significant shift in the AI hardware landscape, moving beyond reliance on established GPU vendors. The company's approach, using its own models to co-design the chip, represents a new paradigm in hardware development, potentially accelerating iteration cycles and optimizing for proprietary workloads. This direct control over hardware allows OpenAI to tailor performance and efficiency specifically for its advanced AI models, challenging the more generalized approach of current market leaders.

This development directly impacts the competitive dynamics between AI model providers and hardware manufacturers. By bringing inference acceleration in-house, OpenAI can reduce dependencies, potentially lower operational costs, and gain a strategic advantage through bespoke performance. Other AI companies may be compelled to follow suit, investing in their own custom silicon or forging deeper partnerships to achieve similar efficiencies, thereby driving further innovation and specialization in the AI hardware market.

AI-written summary. May contain errors.

AI Chips