Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price
Meta released Muse Spark 1.3, its fourth model in five months, with xhigh tier publicly available and max tier in limited preview pending safety testing The model scores 61 (xhigh) and 62 (max) on the Intelligence Index, with gains concentrated in agentic task benchmarks like τ³-Banking where max leads at 52% At $0.55 per index task, it is the cheapest model in its performance tier, undercutting rivals by $0.39 to $0.68 per task despite a price increase from version 1.2 Meta buys the max variant
Analysis
TL;DR
- Meta released Muse Spark 1.3, its fourth model in five months, with xhigh tier publicly available and max tier in limited preview pending safety testing
- The model scores 61 (xhigh) and 62 (max) on the Intelligence Index, with gains concentrated in agentic task benchmarks like τ³-Banking where max leads at 52%
- At $0.55 per index task, it is the cheapest model in its performance tier, undercutting rivals by $0.39 to $0.68 per task despite a price increase from version 1.2
- Meta buys the max variant's edge with significantly more compute, burning 62% more reasoning tokens than xhigh, while open-weights and larger models are announced for the future
- The model trails top performers like Claude Fable 5.1 across most benchmarks and shows regressions in AA-LCR (83→79%) and factual accuracy due to increased refusal behavior
Why It Matters
Meta's aggressive pricing strategy with Muse Spark 1.3 signals a growing commoditization of frontier AI capabilities, making high-scoring models accessible to a broader range of developers and enterprises. The concentration of gains in agentic benchmarks highlights the industry's shift toward practical, tool-using AI systems rather than pure reasoning or knowledge recall. For practitioners, this model represents a cost-effective option for agent-heavy workflows, though those requiring top-tier scientific reasoning or factual accuracy may need to look elsewhere.
Technical Details
- Model tiers: xhigh (publicly available) scores 61 on the Intelligence Index; max (limited preview) scores 62, requiring 62% more reasoning tokens than xhigh
- Benchmark performance: τ³-Banking max leads at 52% (only outright win); Terminal-Bench 2.1 reaches 85-86%; GDPval-AA v2 scores 1,709 (xhigh) and 1,754 (max) vs. Claude Fable 5.1's 1,853; GPQA Diamond at 94%; CritPt at 26%
- Pricing: $1.25 per million input tokens and $4.25 per million output tokens, translating to $0.55 per index task—cheapest in its class
- Regression areas: AA-LCR dropped from 83% to 79%; factual accuracy in AA-Omniscience slipped up to 3 points due to increased refusal-to-answer behavior
- Availability: Released through Muse Code and Meta Model API; open-weights version and larger models announced as upcoming
Industry Insight
- Meta's pricing pressure at the 59+ index tier forces competitors to reconsider their cost-performance ratios, potentially accelerating margin compression across the industry as agentic capabilities become table stakes
- The trade-off between factual accuracy and refusal behavior suggests ongoing tension in alignment strategies; practitioners deploying this model in production should implement verification layers for critical applications
- The rapid iteration cadence (four models in five months) indicates Meta is prioritizing speed-to-market and ecosystem lock-in through Muse Code, suggesting that open-weights availability could significantly shift the competitive landscape for self-hosted deployments
Disclaimer: The above content is generated by AI and is for reference only.