AI News AI资讯 6h ago Updated 1h ago 更新于 1小时前 42

IBM Z and LinuxONE Dual-ISA Processor and AI Acceleration at Hot Chips 2026 IBM Z 和 LinuxONE 双 ISA 处理器及 Hot Chips 2026 上的 AI 加速

IBM unveiled a dual-ISA mainframe processor at Hot Chips 2026 that natively executes both z/Architecture and Arm AArch64 on a single core, implemented in full hardware rather than translation The chip features 11 IBM Z cores at 5.7+ GHz on a 2nm process, with AArch64 v9.3, SVE/SVE2 support, and 2,792 implemented Arm instructions achieving Arm SystemReady compliance A second-generation on-chip AI inference accelerator delivers up to 4x TOPS with FP4/MXFP4 datatypes, 96GB HBM3e at ~4TB/s bandwidth IBM Z处理器采用2nm工艺集成11个核心,每个核心原生支持z/Architecture和AArch64双ISA,无需翻译即可运行Arm软件 AArch64通过完整硬件实现,支持v9.3、SVE/SVE2指令集,实现2792条AArch64指令,达到Arm SystemReady合规 第二代AI加速器配备16个活跃核心加1个冗余核心,支持FP4/MXFP4数据类型,提供高达4x TOPS性能 加速器搭载96GB HBM3e内存,带宽约4TB/s(较上代提升20倍),采用PCIe Gen6低延迟接口 系统目标99.999999%可用性,支持量子安全加密、机密计算和纳秒级线程切换

58
Hot 热度
68
Quality 质量
55
Impact 影响力

Analysis 深度分析

TL;DR

  • IBM unveiled a dual-ISA mainframe processor at Hot Chips 2026 that natively executes both z/Architecture and Arm AArch64 on a single core, implemented in full hardware rather than translation
  • The chip features 11 IBM Z cores at 5.7+ GHz on a 2nm process, with AArch64 v9.3, SVE/SVE2 support, and 2,792 implemented Arm instructions achieving Arm SystemReady compliance
  • A second-generation on-chip AI inference accelerator delivers up to 4x TOPS with FP4/MXFP4 datatypes, 96GB HBM3e at ~4TB/s bandwidth, and includes a redundant core for mission-critical reliability
  • The design targets 99.999999% availability with transparent fault recovery, core sparing, concurrent repair, and RAIM memory protection, while exposing Z accelerators as Linux platform devices to Arm workloads
  • IBM positions the dual-ISA approach as a strategic bridge, letting enterprises retain z/OS and mainframe workloads while tapping the broader Arm software ecosystem through KVM and OpenShift virtualization

Why It Matters

IBM's dual-ISA core represents a bold engineering bet to keep mainframes relevant as Arm-based workloads and AI inference reshape enterprise data centers, rather than treating Z as a closed legacy platform. For AI practitioners and enterprise architects, the integration of a dedicated AI inference accelerator with confidential computing and quantum-safe cryptography directly addresses the growing demand for secure, mission-critical AI deployment in regulated industries. The approach of exposing mainframe accelerators (crypto, compression, AI) as platform devices to Arm Linux also sets a precedent for heterogeneous accelerator sharing across ISAs within a single silicon die.

Technical Details

  • Dual-ISA Core Architecture: 11 IBM Z cores on 2nm at 5.7+ GHz, each natively executing both z/Architecture (big-endian) and AArch64 v9.3 (little-endian) with SVE and SVE2 support; 2,792 AArch64 instructions implemented in full hardware, more than double the Z instruction count
  • Cache and Memory Hierarchy: 36MB private L2 per core, 432MB virtual L3, and 3.5GB virtual L4 cache; SMT=2 design with a dedicated on-chip DPU for I/O acceleration
  • AArch64 Implementation Strategy: IBM reused existing Z core design elements including branch prediction, decode automation from Arm's XML architecture descriptions, and register rename repurposing for GR16-31; new control logic for SVE, new dataflows for FP16, Bfloat16, and crypto, plus CISC instruction reuse for memory operations
  • Second-Generation AI Accelerator: 16 active AI cores plus 1 redundant core, supporting FP4 and MXFP4 datatypes for up to 4x TOPS; 96GB HBM3e with ~4TB/s bandwidth (20x previous generation); PCIe Gen6 peer-to-peer interface for low-latency communication with the host processor
  • Software and Virtualization: Z accelerators exposed as Linux platform devices to Arm workloads with latency comparable to native Z instructions; coexistence of s390x Linux, ARM64 Linux, and IBM z/OS via KVM and OpenShift Virtualization with nanosecond-level thread switching between ISAs
  • Reliability and Security: 99.999999% availability target with error checking across arrays/dataflows/control, transparent transient fault recovery, core sparing, concurrent repair, RAIM memory protection; confidential computing for data at rest/in transit/in use, quantum-safe cryptography, secure boot, and on-chip cryptography tightly integrated with IBM Z operations

Industry Insight

  • IBM's decision to implement AArch64 in full hardware rather than emulation or translation signals a serious commitment to the Arm ecosystem on mainframes, potentially unlocking a new class of Arm-native applications (including AI inference pipelines) to run alongside legacy z/OS workloads on the same silicon—this could accelerate mainframe modernization for enterprises hesitant to abandon Z but needing Arm compatibility
  • The inclusion of a redundant AI core and enterprise-grade confidentiality features positions IBM to compete directly in the mission-critical AI inference market, particularly for regulated industries (finance, healthcare, government) where availability and data protection are non-negotiable; the 4x TOPS with FP4/MXFP4 suggests a focus on efficient inference rather than training, filling a niche between GPU clusters and edge AI chips
  • The dual-ISA coexistence model via KVM/OpenShift, with nanosecond switching between z/Architecture and Arm threads, could become a reference architecture for heterogeneous mainframe design, influencing how other vendors approach ISA convergence; however, the engineering complexity (reusing branch prediction, decode, and register rename across two fundamentally different ISAs) raises questions about long-term maintainability and the pace of future microarchitectural evolution

TL;DR

  • IBM Z处理器采用2nm工艺集成11个核心,每个核心原生支持z/Architecture和AArch64双ISA,无需翻译即可运行Arm软件
  • AArch64通过完整硬件实现,支持v9.3、SVE/SVE2指令集,实现2792条AArch64指令,达到Arm SystemReady合规
  • 第二代AI加速器配备16个活跃核心加1个冗余核心,支持FP4/MXFP4数据类型,提供高达4x TOPS性能
  • 加速器搭载96GB HBM3e内存,带宽约4TB/s(较上代提升20倍),采用PCIe Gen6低延迟接口
  • 系统目标99.999999%可用性,支持量子安全加密、机密计算和纳秒级线程切换

为什么值得看

这篇文章展示了IBM如何将传统大型机的可靠性优势与Arm生态系统和企业级AI推理能力相结合,为Enterprise GenAI提供了全新的硬件基础。这种双ISA原生实现方案代表了大型机厂商在Arm时代保持竞争力的关键战略转型。

技术解析

  • 双ISA核心架构:芯片采用2nm工艺,集成11个IBM Z核心运行在5.7+ GHz,每个核心原生执行z/Architecture(big-endian)和AArch64(little-endian),通过硬件级分支预测复用和自动化XML架构描述实现解码,而非软件翻译。
  • 缓存与内存系统:36MB私有L2缓存组成432MB虚拟L3和3.5GB虚拟L4缓存,配备专用DPU用于I/O加速,支持SMT=2设计,确保双架构工作负载的高效数据访问。
  • AI加速器规格:第二代片上AI加速器包含16个活跃核心加1个冗余核心,支持FP4和MXFP4数据类型,96GB HBM3e内存提供约4TB/s带宽,通过PCIe Gen6实现低延迟对等接口,专为大型企业服务设计。
  • 可靠性工程:实现99.999999%可用性目标,采用数组/数据流/控制的全链路错误检查、瞬态故障透明恢复、持久故障核心备用、并发修复和RAIM内存保护,与NVIDIA在系统级可靠性上的工程思路相似。
  • 软件生态共存:通过Linux KVM和OpenShift虚拟化实现s390x Linux、ARM64 Linux和IBM z/OS在同一处理器上的逻辑分区共存,线程切换仅需纳秒级,Z加速器对Arm Linux暴露为平台设备。

行业启示

  • 大型机演进战略:IBM通过原生双ISA设计而非简单核心堆叠,展示了传统企业计算平台如何通过硬件级创新融合Arm生态,这为其他封闭架构厂商提供了"开放而不失控制"的参考范式。
  • 企业AI推理硬件趋势:冗余核心设计、HBM3e高带宽内存和量子安全加密的组合,表明企业级AI推理正从"性能优先"转向"可靠性+安全+性能"三位一体,适合金融、政务等关键业务场景。
  • 异构计算新路径:纳秒级线程切换和硬件级ISA共存证明了同一核心上运行不同架构工作负载的可行性,为数据中心减少异构硬件管理复杂度提供了技术验证。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Chip 芯片 Inference 推理 LLM 大模型 Deployment 部署