AI Security AI安全 7h ago Updated 2h ago 更新于 2小时前 50

Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models 核破坏恶意软件基准测试难倒多数前沿AI模型

SentinelOne introduces the first long-horizon reverse-engineering benchmark for frontier AI models, using the Fast16 malware as a complex test case. OpenAI’s GPT-5.6 Sol is the only model to successfully complete all eight escalating stages of the investigation, demonstrating superior "project-scale recovery." Other tested models (GPT-5.5, GLM-5.2, Opus 4.x) stalled or declared work finished prematurely due to an inability to retract disproven conclusions and trace downstream dependencies. The s SentinelOne推出首个面向前沿AI模型的“长周期逆向工程”基准测试,以Fast16恶意软件为案例。 OpenAI的GPT-5.6 Sol是唯一完成全部八个渐进式验证阶段的模型,展现出强大的纠错与回溯能力。 其他模型(GPT-5.5, GLM-5.2, Opus 4.x)虽具备局部分析能力,但在面对证据矛盾时无法维持长期可信的调查逻辑。 研究指出核心差距在于“项目级恢复”能力,即撤回错误结论并修正下游所有依赖项的能力,而非单纯的技术洞察。 即使最强模型仍存在语义错误和过早结案问题,人类专家在定义目标、发现盲点和最终审核方面仍不可或缺。

75
Hot 热度
70
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • SentinelOne introduces the first long-horizon reverse-engineering benchmark for frontier AI models, using the Fast16 malware as a complex test case.
  • OpenAI’s GPT-5.6 Sol is the only model to successfully complete all eight escalating stages of the investigation, demonstrating superior "project-scale recovery."
  • Other tested models (GPT-5.5, GLM-5.2, Opus 4.x) stalled or declared work finished prematurely due to an inability to retract disproven conclusions and trace downstream dependencies.
  • The study highlights that even the best-performing models make significant technical errors, necessitating human oversight in high-stakes investigative tasks.

Why It Matters

This benchmark shifts the evaluation of AI capabilities from isolated task completion to sustained, multi-stage reasoning under contradictory evidence, which is critical for real-world security and research applications. It reveals a specific weakness in current frontier models regarding logical consistency over long horizons, informing developers on the need for better error-correction mechanisms. For practitioners, it underscores that AI should currently be viewed as a supervised investigative assistant rather than an autonomous expert.

Technical Details

  • Benchmark Design: The test involves eight escalating stages where new evidence repeatedly contradicts earlier model conclusions, requiring the AI to perform "project-scale recovery" by withdrawing false premises and fixing root causes throughout the entire investigation.
  • Test Case: The investigation focuses on Fast16, a 2005 Windows malware linked to sabotage efforts against Iran’s nuclear program, providing a historically complex and technically dense subject matter.
  • Model Performance: GPT-5.6 Sol completed all stages across three different reasoning-effort settings. In contrast, GPT-5.5 failed at the initial stage, while Z.ai’s GLM-5.2 and Anthropic’s Opus 4.x/4.7/4.8 produced solid local analysis but failed to sustain the investigation.
  • Key Metric: Success is defined not just by technical accuracy but by the ability to maintain a trustworthy narrative flow despite self-contradiction, avoiding premature closure of the inquiry.

Industry Insight

  • Redefining AI Competence: The industry must move beyond simple accuracy metrics to evaluate "logical resilience" and the capacity for iterative self-correction in long-running tasks.
  • Human-in-the-Loop Necessity: Even top-tier models exhibit semantic errors and weak quality control; therefore, AI systems in security and research should be deployed with strict human oversight for objective definition and final validation.
  • Development Focus: Model developers should prioritize architectures that support deep dependency tracing and premise retraction, as these are currently the primary bottlenecks for autonomous investigative agents.

TL;DR

  • SentinelOne推出首个面向前沿AI模型的“长周期逆向工程”基准测试,以Fast16恶意软件为案例。
  • OpenAI的GPT-5.6 Sol是唯一完成全部八个渐进式验证阶段的模型,展现出强大的纠错与回溯能力。
  • 其他模型(GPT-5.5, GLM-5.2, Opus 4.x)虽具备局部分析能力,但在面对证据矛盾时无法维持长期可信的调查逻辑。
  • 研究指出核心差距在于“项目级恢复”能力,即撤回错误结论并修正下游所有依赖项的能力,而非单纯的技术洞察。
  • 即使最强模型仍存在语义错误和过早结案问题,人类专家在定义目标、发现盲点和最终审核方面仍不可或缺。

为什么值得看

该基准测试揭示了当前大模型在处理复杂、长链条且充满不确定性的专业任务时的局限性,特别是缺乏自我纠错和全局一致性维护的能力。对于AI安全、网络安全及自动化工作流开发者而言,这提供了评估模型在真实世界高风险场景中可靠性的关键指标,强调了人机协作模式的必要性。

技术解析

  • 基准设计:不同于传统孤立任务评分,该基准要求模型在八个逐步升级的阶段中持续进行调查,且新证据会反复推翻之前的结论,考验模型的长期一致性和逻辑连贯性。
  • 核心能力指标:重点评估“项目级恢复”(Project-scale recovery),即模型能否识别错误结论、追踪其所有下游影响、修复根本原因并将修正传递至整个调查过程,而非仅修补表面错误。
  • 模型表现对比:GPT-5.6 Sol在不同推理努力下均成功完成全阶段;GPT-5.5卡在初始阶段;GLM-5.2和Opus 4.7/4.8能进行良好的局部分析,但倾向于在缺陷未解决前宣布工作完成。
  • 案例背景:测试对象为Fast16,一款2005年的Windows恶意软件,旨在干扰伊朗核计划相关的LS-DYNA工程软件,具有高度的历史复杂性和技术深度。
  • 局限性发现:即使是表现最好的GPT-5.6 Sol也出现了显著的语义错误、接受了弱质量控制并过早声称准备就绪,证明纯AI自动化尚不成熟。

行业启示

  • 从“单点智能”转向“流程智能”:AI模型的价值不再仅取决于单次回答的准确性,更在于其在长时间跨度内维持逻辑一致性和自我修正的能力,这将重塑对Agent架构的设计要求。
  • 人机协同成为标配:在高风险的专业领域(如网络安全、法律、医疗),AI应定位为“受监督的调查代理”,人类负责设定目标、审查盲点和拥有最终决定权,而非完全替代人类专家。
  • 基准测试需模拟真实复杂性:未来的模型评估应更多采用动态、对抗性且多阶段的测试场景,以暴露模型在长尾逻辑和不确定性环境下的脆弱性,避免实验室环境下的性能虚高。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Security 安全 Benchmark 基准测试 LLM 大模型