AI News AI资讯 7h ago Updated 2h ago 更新于 2小时前 48

Stripping safety guardrails from open-weight AI models is now a turnkey commercial service 为开源AI模型剥离安全护栏现已成为一站式商业服务

Abliteration.ai offers a commercial API that strips trained refusal mechanisms from open-weight models like GLM-5.3, suppressing safety guardrails at the model-weight level rather than via prompt jailbreaking The technique, called "abliteration," identifies and suppresses internal activation patterns that trigger refusals while claiming to preserve coding, cybersecurity, and agentic capabilities The service costs $5 per million tokens and requires no GPU infrastructure from customers, lowering b Abliteration.ai推出商业化服务,从GLM-5.3等开源模型中移除训练好的拒绝机制,通过API以$5/百万token定价销售 技术名为"abliteration",通过抑制触发安全拒绝的内部激活模式修改模型权重,声称保留编码、网络安全和智能体能力 目标客户为网络安全团队、红队演练和AI智能体测试,但无身份验证且不留提示/响应日志,降低滥用门槛 原始GLM-5.2在进攻性安全评估中已零拒绝,业界对"abliteration"实际需求存在争议 服务采用API托管模式,不公开权重,企业客户可通过可选策略网关自定义规则

72
Hot 热度
65
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Abliteration.ai offers a commercial API that strips trained refusal mechanisms from open-weight models like GLM-5.3, suppressing safety guardrails at the model-weight level rather than via prompt jailbreaking
  • The technique, called "abliteration," identifies and suppresses internal activation patterns that trigger refusals while claiming to preserve coding, cybersecurity, and agentic capabilities
  • The service costs $5 per million tokens and requires no GPU infrastructure from customers, lowering barriers for both legitimate red-teaming and potential misuse
  • GLM-5.3 was chosen due to its strong cyber capabilities, open weights, and commercially permissive license from Z.AI that allows derivative "Model as a Service" offerings
  • The no-retention policy (no prompt/response logs, no ID verification) creates a significant accountability gap, though an optional enterprise policy gateway is available for controlled use

Why It Matters

This represents the first turnkey commercial service that systemically removes AI safety guardrails from open-weight models, transforming what was previously a technical exercise into an accessible API product. It forces the industry to confront the tension between open-weight model licensing, legitimate security research needs, and the growing market for unrestricted AI capabilities that can be exploited for malicious purposes.

Technical Details

  • Abliteration technique: The process locates internal activation patterns in model weights that trigger refusal behaviors, then modifies those weights to suppress the patterns. This is a structural change to the model, distinct from prompt-level jailbreaks
  • Base model: GLM-5.3 by Z.AI, selected for its strong coding, agentic, and cybersecurity performance combined with an open-weight, commercially usable license permitting modifications and derivative services
  • Performance benchmarks: Abliterated GLM-5.3 scored 84.5% on CyberGym, 41.8% on Terminal-Bench 4.0, and solved 105 ExploitGym tasks in two hours. GPT-5.5 and GPT-5.6 Sol outperformed on some metrics, though cross-harness comparisons are acknowledged as limited
  • Pricing and access: $5 per million input/output tokens via API; no public weight downloads; no identity verification required; no prompt or response retention
  • Safety exceptions retained: Self-harm refusals and child sexual abuse material blocks remain active by default, with optional additional policy rules available through an enterprise gateway

Industry Insight

  • The commoditization of guardrail removal signals a shift from niche technical experiments to accessible commercial services, likely accelerating both defensive security research and malicious exploitation; regulators should consider whether "Model as a Service" providers that remove safety mechanisms fall under emerging AI oversight frameworks
  • The no-retention, no-verification model creates an accountability vacuum that could hinder incident response and forensic investigation; providers in this space may face increasing pressure to implement usage monitoring or tiered access controls as scrutiny grows
  • The skepticism from established red-team providers (who report abliterated models aren't part of routine work) suggests the market may be overestimating the practical necessity of full guardrail removal, pointing to a potential gap between the service's value proposition and actual professional demand

TL;DR

  • Abliteration.ai推出商业化服务,从GLM-5.3等开源模型中移除训练好的拒绝机制,通过API以$5/百万token定价销售
  • 技术名为"abliteration",通过抑制触发安全拒绝的内部激活模式修改模型权重,声称保留编码、网络安全和智能体能力
  • 目标客户为网络安全团队、红队演练和AI智能体测试,但无身份验证且不留提示/响应日志,降低滥用门槛
  • 原始GLM-5.2在进攻性安全评估中已零拒绝,业界对"abliteration"实际需求存在争议
  • 服务采用API托管模式,不公开权重,企业客户可通过可选策略网关自定义规则

为什么值得看

这篇文章揭示了AI安全护栏商业化剥离的新趋势,反映了开源模型在商业应用中的安全边界问题。对AI从业者而言,这提出了关于模型修改权限、安全责任归属以及监管框架的重要讨论。

技术解析

  • Abliteration技术通过识别并抑制模型内部触发拒绝响应的激活模式来修改权重,而非传统的提示注入攻击。测试数据显示该模型在CyberGym达到84.5%、Terminal-Bench 4.0达到41.8%、ExploitGym在2小时内解决105个任务,但基准测试方法存在差异导致直接对比受限。
  • 选择GLM-5.3作为基础模型因其开源权重和允许商业修改的许可证,相比Qwen、DeepSeek和Mistral等其他开源选项更具优势。
  • 服务采用API托管模式而非公开权重,定价为$5/百万token,不存储提示和响应内容,仅保留运营元数据,且无需传统身份验证。
  • 企业客户可通过可选策略网关自定义请求规则,但标准访问权限保持宽松,额外控制需主动启用。

行业启示

  • 开源模型的安全护栏正在成为可剥离的商品化组件,这为红队测试和恶意用途提供了同等便利,反映了AI安全治理中"双刃剑"效应的加剧。
  • 缺乏身份验证和日志留存的服务模式虽然降低了合法安全团队的门槛,但也为滥用行为提供了掩护,行业需要建立更完善的问责机制。
  • 模型修改权限与商业责任的边界正在模糊,开源许可证允许商业修改的同时,服务提供方是否应承担相应的安全责任,这将成为监管和伦理讨论的焦点。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Open Source 开源 LLM 大模型 Security 安全 Alignment 对齐 Ethics 伦理