AI News AI资讯 3h ago Updated 2h ago 更新于 2小时前 48

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated 英伟达称其Groq 3 LPX比Cerebras快四倍,但计算更为复杂

Nvidia has moved its Groq 3 LPX inference accelerator into full production, targeting ultrafast token generation for agentic AI systems In an Artificial Analysis benchmark, the Groq 3 LPX rack achieved 3,400 tokens/sec on Gemma 4 31B, which Nvidia claims is 4x faster than Cerebras at 882 tokens/sec The comparison is considered skewed: Nvidia requires at least 64 LPUs to reach that result, while Cerebras needs only 1-2 chips, and the CS-4 generation was excluded Groq's SRAM-heavy architecture (50 Nvidia宣布Groq 3 LPX推理加速器进入全面量产,专为Agentic AI的超快token生成设计 在Gemma 4 31B模型基准测试中达到3,400 tokens/秒,Nvidia声称比Cerebras快4倍 专家质疑比较方式不公平:Groq需要至少64个芯片,而Cerebras仅需1-2个加速器 Groq采用SRAM-heavy数据流架构,每个LPU仅500MB内存,远小于Rubin GPU的288GB 对于更大规模的MoE模型(如DeepSeek V3),需要超过1,342个加速器(约5个机架)

72
Hot 热度
68
Quality 质量
67
Impact 影响力

Analysis 深度分析

TL;DR

  • Nvidia has moved its Groq 3 LPX inference accelerator into full production, targeting ultrafast token generation for agentic AI systems
  • In an Artificial Analysis benchmark, the Groq 3 LPX rack achieved 3,400 tokens/sec on Gemma 4 31B, which Nvidia claims is 4x faster than Cerebras at 882 tokens/sec
  • The comparison is considered skewed: Nvidia requires at least 64 LPUs to reach that result, while Cerebras needs only 1-2 chips, and the CS-4 generation was excluded
  • Groq's SRAM-heavy architecture (500 MB per LPU vs. 288 GB per Rubin GPU) splits models across accelerators, making dense models like Gemma 4 31B a best-case scenario
  • Nebius will be the first cloud provider to offer the chip via its Token Factory, with Groq itself as an early adopter

Why It Matters

This highlights the growing importance of inference speed for agentic AI workflows, where token generation latency directly impacts how many reasoning steps and tool calls can occur within acceptable user wait times. It also serves as a cautionary example of how benchmark comparisons can be misleading when hardware configurations and chip counts are not normalized, which is critical for practitioners evaluating inference infrastructure.

Technical Details

  • The Groq 3 LPX is an "interactive AI inference accelerator" extending the Vera Rubin platform, using a mixed GPU-LPU setup where GPUs handle the compute-heavy prefill phase and LPUs handle the bandwidth-heavy decode phase
  • Each LPU contains only 500 MB of SRAM (576x less than a Rubin GPU's 288 GB), requiring models to be split across multiple accelerators over Ethernet, with a single rack holding up to 256 LPUs
  • The benchmark used Gemma 4 31B with a 100,000-token context window across 50 back-to-back requests, achieving steady performance between 10,000 and 100,000 tokens of input length
  • Scaling to larger MoE models like DeepSeek V3 would require approximately 1,342 accelerators (over five racks), raising questions about the architecture's efficiency at scale
  • Nvidia acquired the Groq license for ~$20 billion in December and brought on founder Jonathan Ross and president Sunny Madra

Industry Insight

  • Benchmark claims in the AI hardware space require careful scrutiny of chip counts and system configurations; raw token/sec numbers can be misleading without normalization per chip or per dollar
  • The agentic AI trend is driving demand for specialized inference accelerators that prioritize decode-phase throughput, signaling a shift in hardware design priorities away from training-focused GPUs
  • SRAM-heavy architectures may excel on dense models but face significant scaling challenges with large MoE models, suggesting that hybrid or alternative approaches may be needed for frontier model inference

TL;DR

  • Nvidia宣布Groq 3 LPX推理加速器进入全面量产,专为Agentic AI的超快token生成设计
  • 在Gemma 4 31B模型基准测试中达到3,400 tokens/秒,Nvidia声称比Cerebras快4倍
  • 专家质疑比较方式不公平:Groq需要至少64个芯片,而Cerebras仅需1-2个加速器
  • Groq采用SRAM-heavy数据流架构,每个LPU仅500MB内存,远小于Rubin GPU的288GB
  • 对于更大规模的MoE模型(如DeepSeek V3),需要超过1,342个加速器(约5个机架)

为什么值得看

这篇文章揭示了AI推理硬件市场竞争的关键细节,展示了Nvidia通过收购Groq强化其在推理领域的布局。同时, benchmark数据的解读方式直接影响对技术实力的判断,对从业者评估硬件选型具有重要参考价值。

技术解析

  • 架构设计:Groq 3 LPX采用SRAM-heavy数据流架构,每个LPU仅配备500MB内存(比Rubin GPU少576倍),模型需通过以太网分割到多个加速器上运行。混合架构中GPU负责计算密集的prefill阶段,LPU负责带宽密集的decode阶段。
  • 基准测试表现:在Artificial Analysis的测试中,Gemma 4 31B模型(100K token上下文)在50个连续请求下达到3,400 tokens/秒,性能在10K-100K token输入长度范围内保持稳定。
  • 扩展性挑战:对于大型MoE模型如DeepSeek V3,需要1,342个加速器(约5个机架),而Cerebras仅需1-2个芯片即可运行相同模型,凸显了不同架构在扩展效率上的差异。
  • 商业化进展:Nebius将成为首个通过Token Factory提供该芯片的云服务商,Groq自身也是早期用户之一。

行业启示

  • 硬件benchmark需谨慎解读:厂商宣传的"性能领先"往往基于特定配置和比较方式,从业者应关注芯片数量、模型适配性和实际部署成本等综合指标。
  • 推理加速成为新战场:随着Agentic AI应用对token生成速度要求越来越高,专用推理加速器市场将加速竞争,Nvidia通过收购Groq补齐推理短板。
  • 架构选择影响规模化成本:SRAM-heavy架构在小模型上表现优异,但在处理大规模MoE模型时面临扩展效率挑战,不同架构适合不同应用场景。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Chip 芯片 Inference 推理 Benchmark 基准测试 Product Launch 产品发布 GPU GPU