OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks
OpenAI unveiled "Jalapeño," its first custom inference-only chip developed with Broadcom, which reportedly outperforms Nvidia's Blackwell and Rubin platforms in throughput per watt and token latency Benchmarks show 1.5x–1.9x more AI work per watt and 1.7x–3.6x lower end-to-end latency than the best commercially available systems, with interactive workloads achieving 2.1x–4.1x higher performance The chip was designed in just nine months using OpenAI's own AI models, achieving 54x–104x the token t
Analysis
TL;DR
- OpenAI unveiled "Jalapeño," its first custom inference-only chip developed with Broadcom, which reportedly outperforms Nvidia's Blackwell and Rubin platforms in throughput per watt and token latency
- Benchmarks show 1.5x–1.9x more AI work per watt and 1.7x–3.6x lower end-to-end latency than the best commercially available systems, with interactive workloads achieving 2.1x–4.1x higher performance
- The chip was designed in just nine months using OpenAI's own AI models, achieving 54x–104x the token throughput per kilowatt compared to leading accelerators at matched decoding speed
- SemiAnalysis suggests this could signal the potential end of Nvidia's "CUDA moat," as OpenAI demonstrates rapid silicon bring-up capabilities
- Despite the impressive results, Jalapeño remains at the engineering sample stage and hasn't yet been tested on the latest large models like Deepseek V4 Pro or Kimi K3
Why It Matters
OpenAI's entry into custom chip design represents a significant shift in the AI infrastructure landscape, challenging Nvidia's long-standing dominance in AI accelerators and potentially reshaping the competitive dynamics between cloud providers, chipmakers, and model developers. For AI practitioners and researchers, this development signals a future where leading model companies may increasingly pursue vertical integration of hardware and software, potentially affecting compute costs, availability, and the ecosystem around established platforms like CUDA.
Technical Details
- Chip Architecture & Scope: Jalapeño is a general-purpose LLM inference accelerator designed exclusively for inference, not training. It was developed in partnership with Broadcom, with design work beginning in mid-2024 and fabrication starting in November 2025. The full development cycle took approximately 16 months, with only nine months between initial chip design and the final blueprint.
- Performance Metrics: Tested on SemiAnalysis's public InferenceX benchmark using three models—GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T—Jalapeño achieved approximately 1,400 tokens per second per user on GPT-OSS and over 700 tokens per second on a single concurrent request for Deepseek R1. It delivered 54x to 104x the token throughput per kilowatt compared to the best available accelerator at matched decoding speed.
- Benchmarking Methodology: OpenAI provided the benchmark numbers, and SemiAnalysis verified some runs on-site in their lab. Notably, Jalapeño achieved these results without employing optimizations like multi-token prediction or speculative decoding, while some comparison systems utilized these techniques, leaving room for further performance gains.
- Memory & Comparison Platform: The chip uses HBM4 memory, making Nvidia's Vera Rubin platform a more appropriate comparison than Blackwell. Despite Rubin systems already shipping to customers, Jalapeño still achieved more output tokens per megawatt, though total cost of ownership per token between the two platforms is roughly equivalent.
- AI-Assisted Development: OpenAI leveraged its own AI models during chip development, with older model generations assisting in chip design and newer models accelerating programming and optimization tasks.
Industry Insight
- Erosion of CUDA's Competitive Moat: OpenAI's ability to rapidly design, develop, and bring up custom silicon challenges the notion that Nvidia's CUDA ecosystem creates insurmountable switching costs. If other major model developers follow suit, the AI chip market could fragment, reducing Nvidia's pricing power and ecosystem lock-in.
- Strategic Implications for Partnerships: OpenAI's CFO positioned Jalapeño as complementary to existing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave rather than replacement technology. However, the competitive tension is evident—several of these partners are also building their own AI chips. Organizations should anticipate a more complex hardware landscape and diversify their compute strategies accordingly.
- First-Generation Chip Exception: Historically, first-generation custom chips from tech companies struggle to compete with established offerings. OpenAI's ability to beat both Blackwell and the newer Rubin on the first attempt is unusual and suggests either exceptional engineering talent, significant prior experience in chip design, or favorable design choices tailored specifically to inference workloads. This could accelerate industry trends toward vertical integration among top AI labs.
Disclaimer: The above content is generated by AI and is for reference only.