GPT-6 Astra
GPT-6 Astra is OpenAI's latest flagship model, rolling out to ChatGPT Plus/Pro/Business/Enterprise users and via API at $10/million input and $50/million output tokens It achieves 99.9% on the ARC-AGI 3 benchmark using OpenAI's custom Provider Adapter harness, though the default harness scored only 62.7% Astra dominates security tasks with 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE-Bench binary reverse engineering It excels at long-context processing, scoring 100% on the eight-n
Analysis
TL;DR
- GPT-6 Astra is OpenAI's latest flagship model, rolling out to ChatGPT Plus/Pro/Business/Enterprise users and via API at $10/million input and $50/million output tokens
- It achieves 99.9% on the ARC-AGI 3 benchmark using OpenAI's custom Provider Adapter harness, though the default harness scored only 62.7%
- Astra dominates security tasks with 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE-Bench binary reverse engineering
- It excels at long-context processing, scoring 100% on the eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens
- Despite these strengths, Astra trails Claude Fable 5.1 and Meta's Muse Spark 1.3 on the Artificial Analysis Intelligence Index, though it leads on cost efficiency for coding agent tasks
Why It Matters
GPT-6 Astra represents a significant competitive move by OpenAI to counter Claude Fable and Meta's Muse in the frontier model race, particularly in specialized domains like security and long-context reasoning. Its pricing strategy—matching Claude Fable 5/5.1—signals an aggressive push for market share in the API space, while its near-perfect ARC-AGI 3 score (albeit with a custom harness) suggests meaningful progress on general intelligence benchmarking.
Technical Details
- ARC-AGI 3 Performance: Achieved 99.9% using OpenAI's custom "Provider Adapter harness," which preserves opaque reasoning state between requests and uses compaction for longer conversations; the default harness yielded only 62.7%, highlighting the impact of evaluation methodology
- Security Capabilities: 100% on ExploitBench (vs. GPT-5.6 Sol's 78.5%), 42.4% on ExploitGym (vs. Sol's 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering (vs. Sol's 68.7%)
- Long Context: 100% accuracy on OpenAI's eight-needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, addressing a persistent challenge in the field
- Pricing: API priced at $10/million input tokens and $50/million output tokens, matching Claude Fable 5 and 5.1
- Intelligence Index Positioning: Scores 61 on Artificial Analysis Intelligence Index, matching GPT-5.6 Sol but trailing Claude Fable 5.1 (66) and Meta's Muse Spark 1.3
Industry Insight
- The dramatic gap between Astra's custom harness (99.9%) and default harness (62.7%) on ARC-AGI 3 underscores the importance of standardized evaluation methodologies; practitioners should be cautious about benchmark claims that depend on proprietary evaluation infrastructure
- Astra's cost efficiency in coding agent tasks—less than half the cost of Claude Fable 5 for the same score—makes it a compelling option for production workloads that prioritize economics over raw intelligence, particularly for high-volume coding automation
- OpenAI's focus on security and long-context capabilities appears strategically targeted at enterprise adoption, suggesting the company is positioning Astra as a drop-in replacement for organizations with compliance and large-document processing needs
Disclaimer: The above content is generated by AI and is for reference only.