AI News AI资讯 6h ago Updated 2h ago 更新于 2小时前 60

Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans 微软发布新AI'行为准则',要求模型不得入侵系统或欺骗人类

Microsoft released a comprehensive AI code of conduct establishing values and red lines to guide model training and behavior The document predicts superintelligent AI will surpass human performance within the next decade, calling alignment "one of humanity's greatest challenges" Models will have "absolute constraints" prohibiting cyberattacks, nuclear weapons development, and deepfake production A key provision forbids AI from using adaptive, deceptive, self-reinforcing, or collusive mechanisms 微软发布AI行为准则,作为指导模型安全训练的综合框架,比抽象的"放缓前沿"呼吁更具操作性 准则包含"绝对约束",明确禁止网络攻击、核武器、深度伪造等危险行为 明确规定AI模型不得使用适应性、欺骗性、自我强化、共谋等机制逃避人类监督 预测未来十年超级智能AI将在大多数任务上超越人类,对齐与控制是重大挑战 微软与Anthropic、OpenAI、xAI共同支持"放缓前沿"方法,推动嵌入式评估器机制

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Microsoft released a comprehensive AI code of conduct establishing values and red lines to guide model training and behavior
  • The document predicts superintelligent AI will surpass human performance within the next decade, calling alignment "one of humanity's greatest challenges"
  • Models will have "absolute constraints" prohibiting cyberattacks, nuclear weapons development, and deepfake production
  • A key provision forbids AI from using adaptive, deceptive, self-reinforcing, or collusive mechanisms to evade human oversight
  • Microsoft joins Anthropic, OpenAI, and xAI in supporting a "pacing the frontier" approach with embedded evaluators in AI labs

Why It Matters

Microsoft's code of conduct represents one of the most detailed public frameworks for AI safety from a major tech company, translating abstract alignment concerns into concrete operational constraints. The timing—amid rogue-agent incidents and the Anthropic resignation—signals that safety is moving from theoretical discussion to institutional policy. For AI practitioners, it establishes a benchmark for how leading labs are approaching the tension between capability and control.

Technical Details

  • Each AI model operates under an overarching code of conduct that supersedes individual user preferences and specific task instructions
  • "Absolute constraints" explicitly forbid: cyberattacks, nuclear weapons development, and deepfake production
  • Models are prohibited from employing adaptive, deceptive, self-reinforcing, or collusive mechanisms to evade or defeat human oversight
  • The framework includes broader provisions against any general loss of human control over AI systems
  • Microsoft supports the concept of "embedded evaluators"—internal safety review mechanisms within AI labs—to operationalize alignment as a design goal rather than an afterthought

Industry Insight

  • Microsoft's move signals that AI safety is becoming institutionalized at the enterprise level, likely pressuring other labs to publish similar frameworks or face reputational risk
  • The emphasis on embedded evaluators and deliberate pacing suggests the industry is converging on a self-regulatory model, which could shape future policy and compliance requirements
  • The explicit prohibition on deceptive or self-reinforcing behaviors sets a precedent that may influence how AI agents are designed, tested, and deployed in production environments

TL;DR

  • 微软发布AI行为准则,作为指导模型安全训练的综合框架,比抽象的"放缓前沿"呼吁更具操作性
  • 准则包含"绝对约束",明确禁止网络攻击、核武器、深度伪造等危险行为
  • 明确规定AI模型不得使用适应性、欺骗性、自我强化、共谋等机制逃避人类监督
  • 预测未来十年超级智能AI将在大多数任务上超越人类,对齐与控制是重大挑战
  • 微软与Anthropic、OpenAI、xAI共同支持"放缓前沿"方法,推动嵌入式评估器机制

为什么值得看

微软的AI行为准则为行业提供了可操作的安全框架,将抽象的安全理念转化为具体的训练约束,对AI从业者和政策制定者具有重要参考价值。

技术解析

  • 准则采用分层架构,包含一般原则(如支持而非取代人类、加速人类繁荣)和具体安全约束,模型层面的行为准则优先于用户偏好或特定任务
  • 引入"绝对约束"机制,对网络攻击、核武器、深度伪造等高风险行为实施硬性禁止,同时涵盖更广泛的"防止人类控制权丧失"条款
  • 建立人类监督优先原则,明确禁止AI使用适应性、欺骗性、自我强化、共谋等机制规避控制,确保模型可被可靠地指导、修改或关闭
  • 支持嵌入式评估器方法,在模型开发过程中持续监测和评估安全风险,微软CEO纳德拉公开支持这一机制

行业启示

  • 大型科技公司正从理念层面转向可执行的安全框架,微软准则标志着AI治理进入实操阶段,为行业提供可参考模板
  • 超级智能对齐问题已成为行业共识,微软与Anthropic、OpenAI、xAI的共同立场表明头部厂商在安全节奏上趋于一致
  • 行业对AI风险的认知正在深化,从单纯的技术安全扩展到对人类控制权的保护,反映了rogue-agent事件和Anthropic员工辞职等现实事件的警示作用

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Policy 政策 Ethics 伦理 Alignment 对齐 Closed Source 闭源