AI Practices AI实践 6h ago Updated 1h ago 更新于 1小时前 50

NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure 英伟达NVLink Fusion将NVHBM引入下一代AI基础设施

NVIDIA NVLink Fusion enables hyperscalers and AI-native companies to deploy custom XPUs and CPUs into NVIDIA's AI infrastructure platform using the MGX rack-scale architecture and scale-up/scale-out technology stack NVHBM delivers up to 30% more memory bandwidth per stack compared with standard HBM4e, significantly improving accelerator utilization for memory-bound AI workloads NVHBM reduces PHY and support area by up to 67% versus JEDEC HBM4e, freeing up to 30% more main-die silicon for compute NVIDIA NVLink Fusion 允许超大规模企业和AI原生公司通过统一的scale-up/scale-out技术栈和MGX机架架构,将自定义XPU和CPU部署到NVIDIA AI基础设施平台 NVHBM相比标准HBM4e提供高达30%的内存带宽提升,显著改善内存密集型AI工作负载的加速器利用率和吞吐量 NVHBM将PHY和支持区域减少高达67%,释放最多30%的主die硅片用于计算或其他功能 NVHBM实现15%的HBM功耗降低,在1GW数据中心中可为2000W XPU额外支持多达15,000个XPU NVLink Fusion与NVHBM结合通过带宽、面积和功耗改进的叠加效应,实现

72
Hot 热度
68
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • NVIDIA NVLink Fusion enables hyperscalers and AI-native companies to deploy custom XPUs and CPUs into NVIDIA's AI infrastructure platform using the MGX rack-scale architecture and scale-up/scale-out technology stack
  • NVHBM delivers up to 30% more memory bandwidth per stack compared with standard HBM4e, significantly improving accelerator utilization for memory-bound AI workloads
  • NVHBM reduces PHY and support area by up to 67% versus JEDEC HBM4e, freeing up to 30% more main-die silicon for compute or other features
  • NVHBM achieves 15% lower HBM power usage compared with standard HBM4e, creating thermal headroom for up to 15,000 additional 2,000W XPUs in a 1-gigawatt data center
  • Combining NVLink Fusion with NVHBM delivers a 30% overall end-to-end performance increase per XPU through compounded bandwidth, area, and power improvements at rack scale

Why It Matters

This announcement represents a critical inflection point for the custom accelerator market, as NVIDIA opens its infrastructure platform to semi-custom XPU deployments through NVLink Fusion—effectively competing with its own GPUs while enabling hyperscalers to differentiate their hardware. The NVHBM improvements directly address the three most pressing bottlenecks in modern AI accelerator design: memory bandwidth, silicon area allocation, and power efficiency, making it a pivotal technology for anyone building or deploying large-scale AI inference and training systems.

Technical Details

  • NVLink Fusion serves as the connective technology and IP layer that allows custom XPUs and CPUs to integrate into NVIDIA's scale-up and scale-out networking fabric, leveraging the MGX rack-scale architecture to reduce development complexity and accelerate time to market for semi-custom AI factories
  • NVHBM memory architecture is a custom HBM base-die technology designed and validated with leading memory vendors, delivering up to 30% more memory bandwidth per stack than standard HBM4e, with up to 67% reduction in PHY and support area, and 15% lower power consumption
  • Silicon area optimization through more efficient interface connections frees up to 25% more compute die area, giving accelerator designers additional flexibility to allocate silicon toward matrix engines, vector units, on-chip SRAM, cache hierarchy, or workload-specific capabilities
  • Rack-scale performance combines NVLink Fusion's scale-up networking (supporting expert parallelism and WideEP routing techniques) with NVHBM's bandwidth and power advantages, achieving a compounded 30% end-to-end performance increase per XPU across the entire rack
  • Power and density impact at scale: the 15% HBM power reduction translates to enough thermal headroom in a 1-gigawatt data center to support up to 15,000 additional 2,000W XPUs, directly addressing the power constraints limiting AI factory expansion

Industry Insight

  • NVIDIA is strategically positioning itself as the infrastructure platform provider for the emerging semi-custom XPU market rather than solely a GPU vendor, creating a potential ecosystem lock-in where custom accelerator developers become dependent on NVIDIA's networking, packaging, and rack-scale software stack
  • The 30% bandwidth and 15% power improvements in NVHBM set a new performance-per-watt benchmark that will force competitors (AMD, Intel, and custom XPU startups) to accelerate their own HBM integration roadmaps or risk falling behind on the most critical metrics for large-scale inference and training
  • Hyperscalers pursuing custom AI accelerators should evaluate NVLink Fusion integration early in their design cycle, as the combination of validated NVHBM base dies, MGX architecture, and NVLink scale-up fabric significantly reduces qualification bottlenecks and time-to-deployment compared to building a proprietary memory and interconnect stack from scratch

TL;DR

  • NVIDIA NVLink Fusion 允许超大规模企业和AI原生公司通过统一的scale-up/scale-out技术栈和MGX机架架构,将自定义XPU和CPU部署到NVIDIA AI基础设施平台
  • NVHBM相比标准HBM4e提供高达30%的内存带宽提升,显著改善内存密集型AI工作负载的加速器利用率和吞吐量
  • NVHBM将PHY和支持区域减少高达67%,释放最多30%的主die硅片用于计算或其他功能
  • NVHBM实现15%的HBM功耗降低,在1GW数据中心中可为2000W XPU额外支持多达15,000个XPU
  • NVLink Fusion与NVHBM结合通过带宽、面积和功耗改进的叠加效应,实现每个XPU端到端性能提升30%

为什么值得看

本文揭示了NVIDIA如何通过NVLink Fusion和NVHBM两大技术解决AI加速器规模化部署的核心瓶颈——内存带宽、硅片面积和功耗限制。对于AI从业者和基础设施规划者而言,这提供了自定义AI加速器(XPU)规模化部署的关键技术路径和性能预期。

技术解析

  • NVLink Fusion架构:作为连接技术和IP,NVLink Fusion使客户能够利用NVIDIA的scale-up和scale-out技术栈、生态系统及MGX机架级架构,降低自定义AI工厂的开发和部署复杂度,加速半定制AI加速器的上市时间。
  • NVHBM带宽优势:相比标准HBM4e,NVHBM提供高达30%的内存带宽提升,通过更高效的接口连接减少内存访问延迟,使计算核心持续获得数据供应,特别有利于内存密集型训练和推理工作负载。
  • 面积优化机制:NVHBM通过更高效的接口设计减少PHY和支持区域高达67%,释放最多30%的主die硅片面积,使加速器设计师能够将更多面积分配给矩阵引擎、向量单元、片上SRAM和缓存层次结构等计算功能。
  • 功耗与规模效应:NVHBM实现15%的HBM功耗降低,在1GW数据中心中可为2000W XPU额外支持多达15,000个XPU,通过功耗和散热余量释放实现更大规模的加速器部署。
  • 专家并行支持:NVLink Fusion结合NVHBM支持专家并行(EP)和WideEP等高级路由技术,实现不同专家在不同GPU间的无缝高速同步,NVHBM减少数据饥饿,NVLink处理激活值和隐藏状态传输。

行业启示

  • 自定义加速器生态成熟:NVIDIA通过NVLink Fusion开放自定义XPU部署能力,标志着AI基础设施从封闭GPU生态向半定制加速器生态演进, hyperscalers和AI原生公司可获得更多设计灵活性和成本优化空间。
  • 内存带宽成为新竞争焦点:随着AI模型规模和推理复杂度持续增长,内存带宽瓶颈日益突出,NVHBM的30%带宽提升表明下一代AI加速器竞争将从纯计算性能转向内存子系统优化。
  • 规模化部署的经济性突破:NVHBM的功耗优化和面积释放使1GW数据中心可部署更多XPU,为AI工厂的规模化经济性和TCO优化提供了关键技术支撑,影响数据中心基础设施规划策略。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

GPU GPU Chip 芯片 Training 训练 Inference 推理 Deployment 部署