OpenAI revealed initial benchmark results for its custom Jalapeño chip at the Hot Chips conference on Tuesday, demonstrating what the company describes as a significant performance advantage for AI inference workloads. Richard Ho, OpenAI's hardware vice president, stated that the chip provides both higher throughput and lower latency, a combination often requiring trade-offs in current AI hardware systems.

The Jalapeño chip, an Application-Specific Integrated Circuit (ASIC) developed in partnership with Broadcom, was first introduced in June. It is designed specifically for AI inference, the process by which a trained AI model processes input and generates a response. OpenAI intends to deploy the chip within its own data centers beginning in small volumes by the end of 2026, with a larger rollout planned for 2027.

During testing, Jalapeño demonstrated 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency compared to Nvidia's GB200 and GB300 rack systems. These comparisons were conducted using SemiAnalysis' public InferenceX suite across three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5. For highly interactive workloads, Jalapeño delivered 2.1 to 4.1 times higher performance. The chip operates at 700W, while the Nvidia accelerators it was compared against are rated at 1,200W and 1,400W. OpenAI noted that Jalapeño's measured sustained power remained at or below 550W during testing.

OpenAI's decision to develop its own chip stems from a strategy to optimize its entire AI stack, from models to hardware. The company aims to design products, models, chips, and memory as a cohesive system rather than integrating disparate components. This approach targets specific bottlenecks in AI inference, such as "prefill," where the chip processes a user's prompt, and communication, which involves data transfer between system components. Jalapeño keeps important data, including the KV cache, a short-term memory used during response generation, close to the processor to minimize data movement and reduce delays.

The development of Jalapeño was accelerated by OpenAI's own AI models, which assisted in the chip's design and optimization. This vertical integration allows OpenAI to leverage insights from real-world workloads to enhance every layer of its technology stack. The company indicated that a second-generation chip is already in deep development, with a third generation also taking shape.

Despite these advancements, OpenAI stated it will continue to deploy accelerators from Nvidia and other partners for both training and inference workloads. Richard Ho clarified that OpenAI is currently constrained by data center power availability, emphasizing that "tokens per megawatt" is a key metric. OpenAI has no plans to sell the Jalapeño chip or offer it on a rental market, citing its own need for compute capacity.

The introduction of Jalapeño marks a notable moment in the semiconductor industry, particularly as AI model creators increasingly develop their own chips to manage costs, meet demand, and lessen reliance on established hardware providers. The chip's performance in efficiency tests positions it as a competitor to existing products in the AI processing market.