AI Practices AI实践 6h ago Updated 1h ago 更新于 1小时前 50

NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories 英伟达BlueField-4驱动Agentic AI工厂的新型Scale-In网络基础设施

NVIDIA introduces "Scale-In" as the fifth pillar of its AI networking infrastructure, focusing on north-south access acceleration for agentic AI factories BlueField-4 DPUs provide host-independent acceleration for policy enforcement, storage access, security, and telemetry at up to 800 Gb/s throughput DOCA microservices and Spectrum-X Ethernet enable programmable, policy-driven infrastructure operations including tenant isolation, runtime threat detection, and fleet-wide observability Scale-In a NVIDIA推出Scale-In作为AI网络基础设施第五大支柱,专为Agentic AI工厂的南北向访问提供加速、安全和统一管理 BlueField-4 DPU支持800 Gb/s吞吐量,实现主机独立加速,将策略执行、存储访问、安全和遥测等关键服务从宿主CPU卸载 DOCA微服务与Spectrum-X以太网结合,提供租户隔离、运行时威胁检测、存储虚拟化和全舰队可观测性等可编程基础设施能力 传统软件定义基础设施已无法满足Agentic AI工厂需求,专用DPU处理成为线速网络、存储和安全的必要条件 Scale-In将南北向网络演变为统一的基础设施域,确保AI计算扩展时数据访问、安全和运维不成为瓶

72
Hot 热度
68
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • NVIDIA introduces "Scale-In" as the fifth pillar of its AI networking infrastructure, focusing on north-south access acceleration for agentic AI factories
  • BlueField-4 DPUs provide host-independent acceleration for policy enforcement, storage access, security, and telemetry at up to 800 Gb/s throughput
  • DOCA microservices and Spectrum-X Ethernet enable programmable, policy-driven infrastructure operations including tenant isolation, runtime threat detection, and fleet-wide observability
  • Scale-In addresses the bottleneck where traditional software-defined networking is insufficient for AI factories connecting massive compute with diverse users, agents, and data sources
  • BlueField-4 serves dual roles as both the infrastructure processor for Scale-In and the data/storage processor for CMX (shared KV-cache storage)

Why It Matters

NVIDIA's Scale-In pillar represents a strategic shift from treating infrastructure services as secondary concerns to making them first-class citizens in AI factory design. As agentic AI workloads multiply users, agents, and data sources per server, the north-south network path is becoming a critical bottleneck that traditional CPU-based infrastructure can no longer handle at line rate.

Technical Details

  • BlueField-4 DPU Architecture: Dedicated processing unit offering up to 800 Gb/s throughput, offloading infrastructure services (security, policy enforcement, storage access, telemetry) from host CPUs to prevent bottlenecks at scale
  • DOCA Software Framework: Provides a unified programming model for infrastructure microservices, enabling tenant isolation, runtime threat detection, storage virtualization, and fleet-wide observability as programmable, policy-driven operations
  • Spectrum-X Ethernet Integration: Delivers high-performance Ethernet connectivity across the Scale-In access path, connecting external storage, data sources, and enterprise AI systems to accelerated compute
  • Five-Pillar Infrastructure Model: Scale-Up (NVLink for GPU coherence), Scale-Out (Spectrum-X/Quantum for server interconnect), Scale-Across (Spectrum-XGS for distributed factories), Context Memory (CMX for shared KV-cache), and Scale-In (BlueField-4 for north-south acceleration)
  • Dual-Role Processing: BlueField-4 functions as both the infrastructure processor for Scale-In and the data/storage processor for CMX, unifying AI factory data movement with pod-level context preservation

Industry Insight

  • AI infrastructure is evolving from compute-centric to a holistic five-pillar model; practitioners should evaluate how north-south networking and DPU offloading will scale alongside their GPU investments to avoid infrastructure bottlenecks
  • The convergence of security, storage, and networking into a unified DPU-accelerated domain signals that multi-tenant AI factories will require purpose-built infrastructure processors rather than general-purpose CPU-based solutions
  • Organizations building agentic AI systems should prioritize DOCA-compatible tooling and Spectrum-X Ethernet adoption to ensure their infrastructure operations can keep pace with growing agent-to-compute ratios and data throughput demands

TL;DR

  • NVIDIA推出Scale-In作为AI网络基础设施第五大支柱,专为Agentic AI工厂的南北向访问提供加速、安全和统一管理
  • BlueField-4 DPU支持800 Gb/s吞吐量,实现主机独立加速,将策略执行、存储访问、安全和遥测等关键服务从宿主CPU卸载
  • DOCA微服务与Spectrum-X以太网结合,提供租户隔离、运行时威胁检测、存储虚拟化和全舰队可观测性等可编程基础设施能力
  • 传统软件定义基础设施已无法满足Agentic AI工厂需求,专用DPU处理成为线速网络、存储和安全的必要条件
  • Scale-In将南北向网络演变为统一的基础设施域,确保AI计算扩展时数据访问、安全和运维不成为瓶颈

为什么值得看

本文揭示了NVIDIA从纯计算加速向全栈基础设施协同设计的战略演进,Scale-In填补了AI工厂在数据访问、多租户安全和运维可观测性方面的关键空白。对AI基础设施架构师和云服务商而言,这提供了Agentic AI时代基础设施设计的完整蓝图。

技术解析

  • 五支柱架构体系:NVIDIA AI网络基础设施形成完整五层——Scale-Up(NVLink连接GPU)、Scale-Out(Spectrum-X/Quantum InfiniBand连接服务器)、Scale-Across(Spectrum-X GS连接分布式工厂)、Context Memory(CMX提供共享KV缓存),以及新增的Scale-In(BlueField-4处理南北向访问)
  • BlueField-4核心规格:800 Gb/s吞吐量,独立于主机CPU执行基础设施服务,包括策略执行、存储访问、安全监控和遥测数据采集,防止这些服务在AI计算扩展时成为瓶颈
  • DOCA软件栈能力:提供统一的基础设施编程模型,支持租户隔离、运行时威胁检测、高效存储虚拟化和全舰队可观测性,实现可编程、高性能的策略驱动基础设施操作
  • 存储架构分工:Scale-In连接AI工厂存储(训练/推理/分析)和企业AI数据系统(检索/多模态索引),CMX提供pod级共享KV缓存存储,BlueField-4同时作为两者的基础设施处理器

行业启示

  • 基础设施协同扩展成为AI时代关键:单纯扩展GPU算力已不足够,数据访问、存储、网络安全和运维必须与计算层同步扩展,否则将成为整体性能瓶颈
  • DPU从可选变为必需:Agentic AI工厂的复杂访问模式(多用户、多代理、多数据源交互)要求专用硬件加速基础设施服务,传统通用CPU方案无法满足线速需求
  • 统一基础设施域是演进方向:安全、网络、存储和运维需要从独立管理层演变为协同工作的统一域,提供一致的控制面和可观测性,以支持AI工作负载的动态扩展

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Chip 芯片 GPU GPU Deployment 部署 Security 安全 Agent Agent