AI News AI资讯 4h ago Updated 58m ago 更新于 58分钟前 50

AI is becoming AI's biggest customer as agentic token usage jumps 14x on OpenRouter AI正成为AI的最大客户,OpenRouter上智能体Token使用量激增14倍

AI agent token consumption on OpenRouter surged 14x since February 6, 2026, while human usage grew only 2.8x over the same period February 6, 2026, may mark the inflection point when AI agents surpassed humans as the largest consumers of tokens on the platform Agent token usage grew from 0.51 trillion to 7.3 trillion tokens, with agents increasingly operating autonomously over longer stretches and spawning additional AI processes Nearly 70% of agent token usage comes from cached prompts billed a 2026年2月6日可能是人类最后一次在OpenRouter上比AI代理消耗更多token的日期,此后代理token使用量增长14倍,而人类仅增长2.8倍 AI代理越来越多地自主运行,在长时间任务中自行启动额外的AI进程,形成"AI使用AI"的循环 约70%的代理token消耗来自缓存提示(cached prompts),实际成本增速远低于原始数字显示 OpenRouter偏向开源模型,token效率低于OpenAI/Anthropic等闭源模型,但趋势在主流实验室同样存在 推理模型(reasoning models)已开始推动"token通胀",即使不需要深度思考也会消耗更多token

75
Hot 热度
68
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • AI agent token consumption on OpenRouter surged 14x since February 6, 2026, while human usage grew only 2.8x over the same period
  • February 6, 2026, may mark the inflection point when AI agents surpassed humans as the largest consumers of tokens on the platform
  • Agent token usage grew from 0.51 trillion to 7.3 trillion tokens, with agents increasingly operating autonomously over longer stretches and spawning additional AI processes
  • Nearly 70% of agent token usage comes from cached prompts billed at significantly lower rates, meaning actual costs are not rising proportionally to raw token numbers
  • Token inflation is already underway, driven partly by reasoning models that generate longer internal thought processes even when unnecessary

Why It Matters

This trend signals a fundamental shift in how AI infrastructure is consumed, with autonomous agents becoming the dominant workload rather than direct human interaction. For AI practitioners and infrastructure providers, understanding this shift is critical for capacity planning, pricing models, and architectural decisions around caching and efficiency. The data also suggests that the economics of AI deployment are evolving faster than raw token counts indicate, which has implications for cost forecasting and business models.

Technical Details

  • OpenRouter data shows agentic token usage jumped from 0.51 trillion to 7.3 trillion tokens between February and the present, representing a 14x increase, while human usage grew 2.8x
  • The platform skews toward open-weight models, which are generally less token-efficient than proprietary models from OpenAI or Anthropic, though the trend is expected to mirror patterns at major labs
  • Reasoning models contribute to token inflation by generating extended internal thought processes before responding, even in cases where such deliberation is unnecessary
  • Approximately 70% of agent token usage is attributed to cached prompts, which are billed at substantially reduced rates compared to fresh inference, decoupling raw token growth from actual cost growth
  • Agents are increasingly operating in autonomous multi-step workflows, spawning additional AI processes and extending operational horizons beyond single-turn interactions

Industry Insight

  • Infrastructure and API providers should prioritize caching strategies and prompt optimization as the primary lever for managing costs, since the majority of agent token growth is absorbed by cached responses rather than fresh inference
  • Pricing models based purely on token counts may become increasingly misaligned with actual resource consumption; providers should consider tiered or cache-aware billing structures to reflect true costs
  • The 14x agent growth versus 2.8x human growth suggests that enterprise and developer adoption of agentic workflows will continue to outpace consumer usage, making B2B AI infrastructure a increasingly dominant revenue segment

TL;DR

  • 2026年2月6日可能是人类最后一次在OpenRouter上比AI代理消耗更多token的日期,此后代理token使用量增长14倍,而人类仅增长2.8倍
  • AI代理越来越多地自主运行,在长时间任务中自行启动额外的AI进程,形成"AI使用AI"的循环
  • 约70%的代理token消耗来自缓存提示(cached prompts),实际成本增速远低于原始数字显示
  • OpenRouter偏向开源模型,token效率低于OpenAI/Anthropic等闭源模型,但趋势在主流实验室同样存在
  • 推理模型(reasoning models)已开始推动"token通胀",即使不需要深度思考也会消耗更多token

为什么值得看

这篇文章揭示了AI代理经济正在重塑大模型消费格局,对模型厂商、API平台和开发者都有直接影响。理解token使用模式的变化有助于企业评估AI代理部署的真实成本和效率。

技术解析

  • 数据规模:OpenRouter上代理token消耗从2026年2月的0.51万亿增长至7.3万亿,增幅约14倍;同期人类使用量仅增长2.8倍
  • 缓存机制:近70%的代理token来自缓存提示,这类请求以更低费率计费,显著缓解了成本压力
  • 模型效率差异:OpenRouter以开源权重模型为主,其token效率普遍低于OpenAI、Anthropic等闭源模型,但核心趋势具有行业普适性
  • 推理模型通胀:具备深度思考能力的推理模型倾向于在回答问题前进行更长链式推理,即使任务不需要,也导致token消耗增加

行业启示

  • 成本结构重构:缓存机制使代理token消耗的实际成本增速低于表面数字,企业评估AI代理部署时应关注有效成本而非原始token量
  • 模型选择策略:开源模型在代理场景下token效率较低,企业需权衡成本与功能,闭源模型在token经济性上仍具优势
  • 代理架构优化:随着代理自主性增强,设计时应考虑减少不必要的子进程启动和重复推理,以控制token通胀带来的成本风险

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Agent Agent LLM 大模型 Inference 推理 Open Source 开源 Research 科学研究