AI News AI资讯 3h ago Updated 1h ago 更新于 1小时前 50

OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks OpenAI首款定制芯片"Jalapeño"据报在推理基准测试中击败英伟达Blackwell和Rubin

OpenAI unveiled "Jalapeño," its first custom inference-only chip developed with Broadcom, which reportedly outperforms Nvidia's Blackwell and Rubin platforms in throughput per watt and token latency Benchmarks show 1.5x–1.9x more AI work per watt and 1.7x–3.6x lower end-to-end latency than the best commercially available systems, with interactive workloads achieving 2.1x–4.1x higher performance The chip was designed in just nine months using OpenAI's own AI models, achieving 54x–104x the token t OpenAI首款定制推理芯片"Jalapeño"在Hot Chips会议公布基准测试结果,性能超越Nvidia Blackwell和Rubin 芯片专为LLM推理设计,不用于训练,采用HBM4内存,与Nvidia Vera Rubin平台形成直接竞争 在InferenceX基准测试中,Jalapeño实现1.5-1.9倍更高能效比(每瓦AI工作量),端到端延迟降低1.7-3.6倍 芯片由OpenAI与Broadcom合作开发,从设计到流片仅用9个月,全程使用OpenAI自有AI模型辅助开发 SemiAnalysis认为此举可能终结Nvidia的"CUDA护城河",但Jalapeño目前仍处于工

75
Hot 热度
68
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • OpenAI unveiled "Jalapeño," its first custom inference-only chip developed with Broadcom, which reportedly outperforms Nvidia's Blackwell and Rubin platforms in throughput per watt and token latency
  • Benchmarks show 1.5x–1.9x more AI work per watt and 1.7x–3.6x lower end-to-end latency than the best commercially available systems, with interactive workloads achieving 2.1x–4.1x higher performance
  • The chip was designed in just nine months using OpenAI's own AI models, achieving 54x–104x the token throughput per kilowatt compared to leading accelerators at matched decoding speed
  • SemiAnalysis suggests this could signal the potential end of Nvidia's "CUDA moat," as OpenAI demonstrates rapid silicon bring-up capabilities
  • Despite the impressive results, Jalapeño remains at the engineering sample stage and hasn't yet been tested on the latest large models like Deepseek V4 Pro or Kimi K3

Why It Matters

OpenAI's entry into custom chip design represents a significant shift in the AI infrastructure landscape, challenging Nvidia's long-standing dominance in AI accelerators and potentially reshaping the competitive dynamics between cloud providers, chipmakers, and model developers. For AI practitioners and researchers, this development signals a future where leading model companies may increasingly pursue vertical integration of hardware and software, potentially affecting compute costs, availability, and the ecosystem around established platforms like CUDA.

Technical Details

  • Chip Architecture & Scope: Jalapeño is a general-purpose LLM inference accelerator designed exclusively for inference, not training. It was developed in partnership with Broadcom, with design work beginning in mid-2024 and fabrication starting in November 2025. The full development cycle took approximately 16 months, with only nine months between initial chip design and the final blueprint.
  • Performance Metrics: Tested on SemiAnalysis's public InferenceX benchmark using three models—GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T—Jalapeño achieved approximately 1,400 tokens per second per user on GPT-OSS and over 700 tokens per second on a single concurrent request for Deepseek R1. It delivered 54x to 104x the token throughput per kilowatt compared to the best available accelerator at matched decoding speed.
  • Benchmarking Methodology: OpenAI provided the benchmark numbers, and SemiAnalysis verified some runs on-site in their lab. Notably, Jalapeño achieved these results without employing optimizations like multi-token prediction or speculative decoding, while some comparison systems utilized these techniques, leaving room for further performance gains.
  • Memory & Comparison Platform: The chip uses HBM4 memory, making Nvidia's Vera Rubin platform a more appropriate comparison than Blackwell. Despite Rubin systems already shipping to customers, Jalapeño still achieved more output tokens per megawatt, though total cost of ownership per token between the two platforms is roughly equivalent.
  • AI-Assisted Development: OpenAI leveraged its own AI models during chip development, with older model generations assisting in chip design and newer models accelerating programming and optimization tasks.

Industry Insight

  • Erosion of CUDA's Competitive Moat: OpenAI's ability to rapidly design, develop, and bring up custom silicon challenges the notion that Nvidia's CUDA ecosystem creates insurmountable switching costs. If other major model developers follow suit, the AI chip market could fragment, reducing Nvidia's pricing power and ecosystem lock-in.
  • Strategic Implications for Partnerships: OpenAI's CFO positioned Jalapeño as complementary to existing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave rather than replacement technology. However, the competitive tension is evident—several of these partners are also building their own AI chips. Organizations should anticipate a more complex hardware landscape and diversify their compute strategies accordingly.
  • First-Generation Chip Exception: Historically, first-generation custom chips from tech companies struggle to compete with established offerings. OpenAI's ability to beat both Blackwell and the newer Rubin on the first attempt is unusual and suggests either exceptional engineering talent, significant prior experience in chip design, or favorable design choices tailored specifically to inference workloads. This could accelerate industry trends toward vertical integration among top AI labs.

TL;DR

  • OpenAI首款定制推理芯片"Jalapeño"在Hot Chips会议公布基准测试结果,性能超越Nvidia Blackwell和Rubin
  • 芯片专为LLM推理设计,不用于训练,采用HBM4内存,与Nvidia Vera Rubin平台形成直接竞争
  • 在InferenceX基准测试中,Jalapeño实现1.5-1.9倍更高能效比(每瓦AI工作量),端到端延迟降低1.7-3.6倍
  • 芯片由OpenAI与Broadcom合作开发,从设计到流片仅用9个月,全程使用OpenAI自有AI模型辅助开发
  • SemiAnalysis认为此举可能终结Nvidia的"CUDA护城河",但Jalapeño目前仍处于工程样品阶段

为什么值得看

这篇文章揭示了AI基础设施领域的重要转折:垂直整合的推理芯片开始挑战通用GPU的统治地位。对AI从业者而言,这标志着推理成本优化和算力自主可控成为新的战略焦点,可能重塑云服务商与芯片厂商的合作格局。

技术解析

  • 性能指标:在GPT-OSS 120B模型上达到约1,400 tokens/秒/用户,Deepseek R1 670B模型单并发请求超700 tokens/秒;匹配解码速度下,每千瓦token吞吐量是最佳加速器的54-104倍
  • 测试基准:采用SemiAnalysis公开的InferenceX基准,测试模型包括GPT-OSS 120B、Deepseek R1 670B和Kimi K2.5 1T,部分结果在实验室现场验证
  • 开发周期:2024年中启动设计,2025年11月流片,总周期16个月,其中从首版设计到完成蓝图仅9个月,使用多代OpenAI模型辅助芯片设计和编程优化
  • 技术限制:未采用多token预测或投机解码等优化技术,仍有性能提升空间;与Nvidia Vera Rubin对比时,后者使用了优化技术但Jalapeño仍胜出

行业启示

  • CUDA护城河动摇:OpenAI快速在自研硅片上部署模型的能力,可能削弱Nvidia长期依赖的软件生态壁垒,其他云厂商和AI公司或加速自研芯片进程
  • 推理芯片市场格局重塑:Jalapeño证明专用推理芯片在能效和延迟上可超越通用GPU,推动AI基础设施从"训练优先"向"推理优化"转变,Cerebras、Groq等专用芯片公司面临更直接竞争
  • 合作与竞争并存:OpenAI CFO强调Jalapeño是补充而非替代现有Nvidia/AMD/AWS合作,但芯片自研成功将增强OpenAI在算力谈判中的议价能力,可能引发云厂商与AI公司的关系重构

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Chip 芯片 Inference 推理 Benchmark 基准测试 GPU GPU Product Launch 产品发布