OpenAI has unveiled Jalapeño, its first custom inference chip designed to slash data center energy costs and boost response times. Slated for deployment by late 2026, the 700-watt processor outperformed Nvidia’s GB300 in recent public benchmarks for power efficiency and latency, according to reporting across the tech industry.
OpenAI and Broadcom Partnership for Jalapeño
The race for artificial intelligence hardware dominance is entering a new phase as major model developers begin designing custom silicon to reduce their reliance on established market leaders. OpenAI has stepped directly into that arena by releasing performance data for its proprietary inference chip, known as Jalapeño (Deraktionaer). Developed in partnership with Broadcom, the custom processor aims to tackle soaring electricity and infrastructure expenses as user demand for generative tools and multi-step AI agents accelerates.
Der Betriebsbeginn des Chips in OpenAI eigener Infrastruktur ist laut Berichten noch für dieses Jahr vorgesehen (Deraktionaer). Richard Ho, chip chief at OpenAI, sagte, Jalapeño verbinde hohen Durchsatz mit niedriger Latenz (Yellow). Kunden könnten dadurch Modelle wählen, die entweder besonders kostengünstig oder besonders schnell antworten (Yellow). OpenAI-Chefentwickler betonen, dass bei KI-Agenten zahlreiche einzelne Schritte nacheinander ausgeführt werden müssen, weshalb sich jede Verzögerung auf die gesamte Aufgabe auswirkt (Deraktionaer).
Benchmark Performance Against Nvidia GB300
InferenceX Benchmark Results Against Nvidia GB300
In public evaluations using the InferenceX benchmark from SemiAnalysis, Jalapeño went head-to-head with competing hardware. According to Yellow, the chip surpassed Nvidia’s GB300 in two primary metrics during tests conducted at a 700-watt power ceiling: throughput per unit of power and response speed. Across examinations involving models such as GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5 1T, OpenAI reported that the processor delivered between 1.5 and 1.9 times more AI work per watt at maximum throughput.
End-to-end latency dropped significantly during the runs, registering between 1.7 and 3.6 times lower than comparison systems (Deraktionaer). Bei interaktiven Anwendungen erreichte der Chip laut OpenAI eine 2,1- bis 4,1-mal höhere Performance (Deraktionaer). The most pronounced efficiency gains surfaced while processing Moonshot AI’s larger Kimi model (Yellow). While the hardware is nominally rated for a 700-watt design envelope, actual continuous power draw during testing registered at a maximum of 550 watts (Deraktionaer).
Architectural Strategy and Broadcom Partnership
Broadcom Collaboration and Software Optimization
The creation of Jalapeño stems from a collaboration that moved rapidly from concept to completion. Following an announcement on June 24 regarding their joint hardware efforts, OpenAI and Broadcom brought the initial design from its drawing board to a manufacturing tape-out in just nine months, utilizing OpenAI’s own software models to optimize the layout (Yellow).

Richard Ho, chip chief at OpenAI, noted that the processor successfully pairs high throughput with minimal latency, giving infrastructure operators flexibility depending on whether their operational priorities favor cost savings or rapid execution (Yellow). Rather than focusing solely on raw single-chip speed, the engineering team designed the hardware stack—incorporating localized memory, specialized networking, and customized software—to minimize energy-wasting data movements, such as keeping the KV-cache local (Deraktionaer).
Evolving Infrastructure and Future Roadmap
Cerebras Systems and Second Generation Jalapeño
Despite the competitive benchmarks, industry analysts emphasize that the comparison comes with strict boundaries. Jalapeño is optimized strictly for inference tasks—the execution of trained models to generate answers—rather than the intensive compute required for model training (Yellow). Furthermore, it was not pitted against Nvidia’s newer Vera-Rubin architecture (Yellow).
OpenAI intends to maintain a diversified hardware supply chain. For specific smaller-model inference workloads, the organization also relies on technology from Cerebras Systems (Yellow).
Work on subsequent iterations is already well underway. Eine zweite Jalapeño-Generation ist bereits weit fortgeschritten, das Tape-out wird in den kommenden Monaten erwartet, und parallel hat OpenAI konzeptionelle Arbeiten an einer dritten Generation begonnen (Yellow).
Продолжение темы

