AI Practices AI实践 13h ago Updated 5h ago 更新于 5小时前 49

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage NVIDIA Vera 存储基准测试:为 AI 原生存储提供更快的加密、压缩、完整性检查和恢复

NVIDIA Vera BlueField-4 STX Storage Processor delivers significant throughput advantages over x86 CPUs across encryption, decryption, Reed-Solomon recovery, CRC32C integrity checking, compression, decompression, and multi-stage pipeline operations The processor integrates 88 Olympus Armv9.2 cores with Spatial Multithreading, Scalable Coherency Fabric (SCF), and SOCAMM2 LPDDR5X memory to address both single-thread performance and bandwidth-intensive storage processing demands Vera architecture en NVIDIA Vera BlueField-4 STX Storage Processor采用88核Olympus Armv9.2架构,专为AI原生存储处理优化 相比x86 CPU,在加密/解密、压缩/解压、完整性检查等存储原语操作上实现1.29x-3.67x性能提升 通过SCF可扩展一致性总线和SOCAMM2 LPDDR5X内存提供3.4 TB/s片上带宽和1.2 TB/s内存带宽 统一Vera CPU架构可在更低CPU、功耗和散热预算下扩展代理执行和存储处理,支持更高服务密度 存储已成为agentic AI工作流的核心环节,需持续供应和保留数据以驱动代理推理循环

72
Hot 热度
70
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • NVIDIA Vera BlueField-4 STX Storage Processor delivers significant throughput advantages over x86 CPUs across encryption, decryption, Reed-Solomon recovery, CRC32C integrity checking, compression, decompression, and multi-stage pipeline operations
  • The processor integrates 88 Olympus Armv9.2 cores with Spatial Multithreading, Scalable Coherency Fabric (SCF), and SOCAMM2 LPDDR5X memory to address both single-thread performance and bandwidth-intensive storage processing demands
  • Vera architecture enables AI-native storage platforms to scale agent execution and storage processing within lower CPU, power, and cooling envelopes
  • Storage processing is identified as a critical bottleneck in agentic AI workflows, where each agent step triggers multiple storage operations across thousands of concurrent agents
  • The unified Vera CPU architecture bridges the gap between GPU inference acceleration and CPU-side storage processing, ensuring storage systems can keep pace with accelerated computing requirements

Why It Matters

This development addresses a growing bottleneck in AI infrastructure where storage processing operations—encryption, compression, integrity checking—become performance constraints as agentic AI workloads scale. For AI practitioners and infrastructure engineers, the Vera processor demonstrates that CPU-side storage processing is no longer a secondary concern but a critical component of AI-native data platforms that must match GPU acceleration speeds.

Technical Details

  • Architecture: 88 NVIDIA-designed Olympus Armv9.2 cores with 176 Spatial Multithreading threads, Scalable Coherency Fabric (SCF) providing up to 3.4 TB/s bisection bandwidth and 164 MB unified L3 cache, and SOCAMM2 LPDDR5X memory delivering up to 1.2 TB/s aggregate bandwidth (14 GB/s per core)
  • Benchmark Performance: Vera outperforms x86 CPUs by up to 1.43x (encryption), 1.29x (decryption), 3.26x (Reed-Solomon recovery), 3.67x (CRC32C integrity checking), 3.29x (compression), 1.72x (decompression), and 3.21x (multi-stage pipeline operations)
  • Storage Primitives Optimization: The architecture addresses dual demands of sustained per-core performance for sequential data stream operations and high bandwidth with predictable latency for concurrent multi-stream workloads
  • Implementation: Part of the NVIDIA STX foundation for AI-native data platforms, bringing Vera CPU performance directly into the storage data path to support agentic AI workflows involving enterprise knowledge retrieval, persistent memory access, KV cache reuse, and tool execution

Industry Insight

  • AI infrastructure design must evolve to treat storage processing as a first-class performance concern rather than an afterthought, as agentic AI workloads will increasingly demand storage systems that can process encryption, compression, and integrity operations at accelerated computing speeds
  • The convergence of CPU and storage processing architectures in the Vera processor suggests a strategic direction toward unified compute solutions that reduce infrastructure complexity while improving energy efficiency in AI factories
  • Organizations deploying large-scale agentic AI systems should evaluate storage processing bottlenecks early in their infrastructure planning, as conventional CPU scaling approaches will prove insufficient and cost-prohibitive compared to purpose-built storage processors

TL;DR

  • NVIDIA Vera BlueField-4 STX Storage Processor采用88核Olympus Armv9.2架构,专为AI原生存储处理优化
  • 相比x86 CPU,在加密/解密、压缩/解压、完整性检查等存储原语操作上实现1.29x-3.67x性能提升
  • 通过SCF可扩展一致性总线和SOCAMM2 LPDDR5X内存提供3.4 TB/s片上带宽和1.2 TB/s内存带宽
  • 统一Vera CPU架构可在更低CPU、功耗和散热预算下扩展代理执行和存储处理,支持更高服务密度
  • 存储已成为agentic AI工作流的核心环节,需持续供应和保留数据以驱动代理推理循环

为什么值得看

本文揭示了AI原生存储处理的新范式:存储不再是被动数据层,而是主动参与AI工作流的关键组件。NVIDIA通过专用存储处理器解决AI工厂级基础设施的存储处理瓶颈,为从业者提供了存储加速的技术路线参考。

技术解析

  • 硬件架构:Vera BlueField-4 STX集成88个NVIDIA设计的Olympus Armv9.2核心,支持176个Spatial Multithreading线程,配备164 MB统一L3缓存和SOCAMM2 LPDDR5X内存子系统
  • 带宽设计:SCF提供最高3.4 TB/s的分区带宽,SOCAMM2 LPDDR5X提供1.2 TB/s聚合内存带宽(每核14 GB/s),满足高并发带宽密集型工作负载需求
  • 性能基准:加密/解密提升1.29x-1.43x,Reed-Solomon恢复提升3.26x,CRC32C完整性检查提升3.67x,压缩提升3.29x,解压提升1.72x,多阶段流水线提升3.21x
  • 架构优势:Olympus核心结合宽指令吞吐量、高级分支预测、深度乱序执行和向量/加密资源,同时满足单线程性能和多流并发带宽需求
  • 统一架构:与NVIDIA Rubin GPU相同的Vera CPU架构,实现存储处理与GPU计算的统一平台,降低基础设施复杂度

行业启示

  • 存储处理成为AI性能新瓶颈:随着agentic AI并发度和上下文规模增长,存储路径上的加密、压缩、完整性检查等操作已成为制约系统整体性能的关键环节,需从架构层面重新审视存储处理设计
  • 专用存储处理器价值凸显:传统x86 CPU在存储处理扩展上面临功耗和成本压力,专用存储处理器(如BlueField系列)通过架构优化实现更高吞吐和更低功耗,将成为AI原生基础设施的重要组件
  • 统一CPU架构趋势:NVIDIA将同一Vera CPU架构同时用于GPU数据供给和存储处理,体现了AI基础设施向统一计算平台演进的趋势,有助于降低部署复杂性和运营成本

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Chip 芯片 Benchmark 基准测试 GPU GPU Product Launch 产品发布 Deployment 部署