AI News AI资讯 7d ago Updated 7d ago 更新于 7天前 49

Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks Z.ai发布GLM-5.3无需重新训练基础模型:在复杂编码和长程任务上表现更佳

GLM-5.3 reuses the same 743B base model as GLM-5.2, with all performance gains derived from scaled post-training rather than architectural changes Terminal-Bench 3.0 scores surged from 4.6 to 28.3, and DeepSWE v1.1 improved from 46.2 to 66.9, marking the largest gains on long-horizon coding benchmarks Cybersecurity capabilities exceeded expectations: CyberGym reached 84.5%, edging past Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), while ExploitBench more than doubled to 54.4% The model demonstrates GLM-5.3复用GLM-5.2的743B基础模型,所有性能提升均来自扩展后训练(更多任务环境、更长训练时间)。 编码能力显著增强,Terminal-Bench 3.0得分从4.6跃升至28.3,DeepSWE v1.1从46.2提升至66.9。 网络安全能力超预期,CyberGym达到84.5%,超越Mythos 5和GPT-5.6 Sol;ExploitBench得分翻倍至54.4%。 模型权重尚未公开,预计发布两周后完成安全评估与加固。 通过Z.ai API、GLM Coding Plan和ZCode部分部署,适合初创公司和中型企业立即采用。

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • GLM-5.3 reuses the same 743B base model as GLM-5.2, with all performance gains derived from scaled post-training rather than architectural changes
  • Terminal-Bench 3.0 scores surged from 4.6 to 28.3, and DeepSWE v1.1 improved from 46.2 to 66.9, marking the largest gains on long-horizon coding benchmarks
  • Cybersecurity capabilities exceeded expectations: CyberGym reached 84.5%, edging past Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), while ExploitBench more than doubled to 54.4%
  • The model demonstrates compounding capability gains as training scales, forming coherent plans across complete exploitation chains rather than just single-bug reasoning
  • Weights will be released approximately two weeks after launch following safety evaluation and hardening; API access is available immediately

Why It Matters

GLM-5.3 demonstrates that significant capability gains can be achieved through post-training scaling alone without retraining the base model, which has major cost and efficiency implications for the industry. The unexpected cybersecurity breakthrough—where vulnerability-discovery training data led to compounding multi-step exploitation planning—signals that agentic AI systems are approaching dangerous levels of autonomous capability, raising urgent questions about safety evaluation and responsible disclosure.

Technical Details

  • Architecture: Retains the 743B base model from GLM-5.2 with no architectural modifications; all improvements come from scaled post-training with more task environments, more environment types, and longer training duration
  • Coding Benchmarks: Terminal-Bench 3.0 improved from 4.6 to 28.3; DeepSWE v1.1 from 46.2 to 66.9; Agents' Last Exam (CLI) from 23.8 to 28.5; GDPval-AA v2 scored 1,769 across 44 occupations; Z.ai Code Bench showed 50% improvement over GLM-5.2 at ~50,000 output tokens per task
  • Cybersecurity Benchmarks: CyberGym reached 84.5% (white-box source vulnerability discovery and validation); ExploitBench jumped from 24.4% to 54.4% (root-cause reasoning with working exploits); ExploitGym completed 105 tasks in 2 hours and 130 in 6 hours
  • Deployment: Available via Z.ai API, GLM Coding Plan, and ZCode; open weights pending safety evaluation and hardening (~2 weeks); internal benchmark used to reduce contamination risk

Industry Insight

  • Organizations should prioritize evaluating GLM-5.3 for long-horizon coding tasks and security-focused applications, as the compounding capability gains in exploitation chains suggest agentic systems are approaching autonomous threat-discovery levels that could reshape both offensive and defensive security operations
  • The post-training-only improvement strategy validates a cost-effective path for model iteration; companies should invest in diverse, high-quality task environments and extended training rather than pursuing expensive base model retraining cycles
  • Enterprises with data-residency or vendor-review requirements should plan for the weight release window, while startups and mid-market teams can begin immediate adoption via API to capture early competitive advantage in developer tooling and application security workflows

TL;DR

  • GLM-5.3复用GLM-5.2的743B基础模型,所有性能提升均来自扩展后训练(更多任务环境、更长训练时间)。
  • 编码能力显著增强,Terminal-Bench 3.0得分从4.6跃升至28.3,DeepSWE v1.1从46.2提升至66.9。
  • 网络安全能力超预期,CyberGym达到84.5%,超越Mythos 5和GPT-5.6 Sol;ExploitBench得分翻倍至54.4%。
  • 模型权重尚未公开,预计发布两周后完成安全评估与加固。
  • 通过Z.ai API、GLM Coding Plan和ZCode部分部署,适合初创公司和中型企业立即采用。

为什么值得看

GLM-5.3展示了仅通过扩展后训练即可大幅提升复杂编码和网络安全任务的能力,为模型优化提供了高效路径。其网络安全能力的突破可能推动安全工具和漏洞发现流程的革新,同时权重延迟发布策略反映了安全评估的重要性。

技术解析

  • 架构与训练:基于743B参数基础模型,未进行重新训练,所有改进来自扩展后训练,包括增加任务环境类型和延长训练时长。
  • 编码基准测试:Terminal-Bench 3.0得分28.3(原4.6),DeepSWE v1.1得分66.9(原46.2),Z.ai Code Bench内部测试提升50%。
  • 网络安全评估:CyberGym得分84.5%,超过Mythos 5(83.8%)和GPT-5.6 Sol(83.6%);ExploitBench得分54.4%,较GLM-5.2的24.4%翻倍。
  • 部署与访问:通过Z.ai API、GLM Coding Plan和ZCode部分可用,权重预计两周后发布,需完成安全评估。
  • 性能对比:在公开基准上落后于GPT-5.6 Sol和Fable 5,但内部测试显示竞争力,强调私有基准减少污染风险。

行业启示

  • 后训练扩展成为关键路径:企业可优先考虑优化训练数据而非重新训练基础模型,以降低成本和时间。
  • 网络安全AI能力快速进步:安全厂商和MSSPs需关注此类模型在漏洞发现和利用链规划上的突破,调整防御策略。
  • 权重延迟发布策略平衡创新与安全:企业应建立安全评估流程,确保模型部署符合合规要求,尤其涉及数据驻留和供应商审查的机构。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Code Generation 代码生成 Agent Agent Training 训练 Benchmark 基准测试