TAI #218: Enterprise AI Use Is Becoming More Uneven
Two datasets (OpenAI and Ramp) reveal rapidly widening concentration of AI token usage and spending among a small elite of enterprise customers, with the top 10% generating 8.3x more tokens than the middle decile and the top 1% spending $88,806/employee annually—619x the median firm Grok 4.6 (SpaceXAI) achieves frontier-level performance (Artificial Analysis score 61) at the lowest task cost ($0.84) among top-scoring models, released just 27 days after Grok 4.5, signaling accelerating release ca
Analysis
TL;DR
- Two datasets (OpenAI and Ramp) reveal rapidly widening concentration of AI token usage and spending among a small elite of enterprise customers, with the top 10% generating 8.3x more tokens than the middle decile and the top 1% spending $88,806/employee annually—619x the median firm
- Grok 4.6 (SpaceXAI) achieves frontier-level performance (Artificial Analysis score 61) at the lowest task cost ($0.84) among top-scoring models, released just 27 days after Grok 4.5, signaling accelerating release cadences
- OpenAI launches GPT-5.6 Sol Ultrafast at 750 output tokens/sec (14x Standard speed) via Cerebras, with no model downgrade, targeting latency-sensitive use cases like quantitative trading
- Gemini 3.7 Flash delivers substantial coding and workflow gains (DeepSWE: 49%→65.3%) with a 1M-token context window, while Z.AI's GLM-5.3 shows dramatic Terminal-Bench improvements (4.6→28.3) through longer, more realistic engineering training environments
- NVIDIA releases Nemotron 3.5 Lightning (30B MoE, 3B active) for high-volume agent execution layers, and OpenAI expands its Daybreak cybersecurity program with GPT-5.6-Cyber, finding two previously unknown V8 vulnerabilities
Why It Matters
The data reveals a critical divergence: AI adoption is concentrating at an accelerating rate among power users and well-funded firms, but high token spend does not equate to high productivity or return. For AI practitioners, the key challenge is no longer access to frontier models but the ability to design, route, and evaluate agent workflows that convert spend into measurable outcomes. The emergence of specialized tiers—ultrafast for latency-critical work, low-cost frontier for bulk tasks, and cybersecurity-hardened variants—means organizations must develop nuanced model-strategy frameworks rather than treating all AI spend as interchangeable.
Technical Details
- OpenAI Enterprise Concentration Data: Top 10% of enterprise customers by output tokens per active user generated 8.3x more tokens than the middle decile in June (up from 2.6x in January). Codex accounted for 64% of combined ChatGPT+Codex enterprise output tokens. The share of Codex users attempting tasks estimated to take a skilled person 8+ hours rose from 2.1% (Dec) to 25.6% (May).
- Grok 4.6 (SpaceXAI): Artificial Analysis score 61, matching GPT-5.6 Sol Max. Pricing: $0.84/task vs. $1.23 (Sol), $3.14 (Fable 5), $2.34 (Opus 5). API pricing: $2/$6 per million tokens (below 200K context), doubling beyond. 500K context, text+image inputs, new xhigh reasoning level. Training included curated model-generated reasoning, improved optimizer, regenerated SFT from 4.5, and RL environments spanning coding, web dev, CAD, kernel optimization.
- Gemini 3.7 Flash (Google): DeepSWE 65.3% (from 49.0%), FrontierCode 43.6% (from 34.4%), GDP.pdf 34.0% (from 22.0%), AutomationBench 30.4% (from 17.0%). 1M-token context, 64K output. Pricing: $0.75/$3.75 per million input/output through Dec 31, 2027. Supports text, images, audio, video, function calling, search, and computer use.
- GLM-5.3 (Z.AI): Terminal-Bench 3.0: 28.3% (from 4.6%), DeepSWE: 66.9% (from 46.2%), CyberGym: 84.5%, ExploitBench: 54.4% (from 24.4%). 1M context, 128K max output, always-on reasoning. Weights held for 2 weeks for safety evaluation. Discovered 2,436 vulnerabilities across 269 projects.
- NVIDIA Nemotron 3.5 Lightning: 30B-parameter MoE with 3B active parameters, hybrid Mamba-2/MoE/Attention architecture, 1M-token context. Up to 4x faster than similar-sized models, 30% faster PinchBench than Qwen3.6-35B at comparable accuracy. Artificial Analysis score 24. NeMo Switchyard enables routing agent workflow segments to different models. BF16 and NVFP4 weights under OpenMDW-1.1.
- GPT-5.6 Sol Ultrafast: Powered by Cerebras, up to 750 output tokens/sec (14x Standard). Same model as Sol with no downgrade. Access limited, pricing not published.
- GPT-5.6-Cyber (OpenAI Daybreak): Sol-based specialized model for authorized security research. 95% response rate to advanced cyber requests vs. 1.5% for safeguarded Sol. Found two previously unknown V8 vulnerabilities chainable to escape heap sandbox (one fixed as CVE-2026-15903).
Industry Insight
- The productivity gap is the new competitive frontier: Organizations that treat AI spend as a proxy for value will face mounting waste. The data shows a 619x spread between top and median AI spend per employee, but spend ≠ return. Firms should pair every usage metric with outcome measures—quality, review cycles, rework rates, and latency sensitivity—and route tasks to the cheapest model tier that clears their quality and speed targets.
- Model specialization is accelerating across three dimensions: cost (Grok 4.6 at $0.84/task), speed (Sol Ultrafast at 14x), and domain (GPT-5.6-Cyber, Nemotron 3.5 Lightning for agent execution). The winning strategy is not picking one "best" model but building routing infrastructure (like NeMo Switchyard) that directs work to the right tier based on task characteristics.
- The release cadence arms race has operational consequences: With Grok 4.6 arriving 27 days after 4.5 and Gemini 3.7 Flash just 3 weeks after 3.6, firms need agile evaluation pipelines rather than static model procurement. The recommendation to start new workflows with small budgets and clear tests, then scale to proven users, should be institutionalized as a standard operating procedure to avoid locking into underperforming model choices.
Disclaimer: The above content is generated by AI and is for reference only.