Google says its new Gemini 3.8 Flash model 'works harder' but might cost more
Google launched Gemini 3.8 Flash, claiming more reasoning steps and iterative tool calling compared to its predecessor, with unchanged per-token pricing but potential for higher overall costs due to increased token usage Early benchmarks show Gemini 3.8 Flash outperforming competitors including Anthropic's Fable 5 on DeepSWE v1.1, Vals Finance Agent V2, and Harvey's Legal Agent benchmark Artificial Analysis estimates effective pricing is up ~40% from 3.7 Flash due to a 30% increase in output tok
Analysis
TL;DR
- Google launched Gemini 3.8 Flash, claiming more reasoning steps and iterative tool calling compared to its predecessor, with unchanged per-token pricing but potential for higher overall costs due to increased token usage
- Early benchmarks show Gemini 3.8 Flash outperforming competitors including Anthropic's Fable 5 on DeepSWE v1.1, Vals Finance Agent V2, and Harvey's Legal Agent benchmark
- Artificial Analysis estimates effective pricing is up ~40% from 3.7 Flash due to a 30% increase in output tokens per task and more agentic evaluation turns
- Google released a specialized "Cyber" variant alongside the Fairwind Program, a limited-access initiative for governments and trusted security partners like CrowdStrike and CISA
- The model includes built-in safeguards against misuse in CBRN and cyber offense domains, while remaining available to consumers, developers, and enterprise users
Why It Matters
Gemini 3.8 Flash represents Google's continued push into the agentic AI space, directly competing with Anthropic and other frontier models on software engineering and autonomous task performance. The pricing dynamics—unchanged per-token rates but significantly higher effective costs—signal a shift in how AI economics may evolve as models consume more tokens to achieve better results. For practitioners, this highlights the importance of monitoring actual usage patterns rather than relying solely on published pricing.
Technical Details
- Gemini 3.8 Flash performs more reasoning steps on complex tasks and calls tools iteratively, with configurable effort levels that can increase token consumption
- Benchmarked on DeepSWE v1.1 (software engineering), Vals Finance Agent V2 (finance), and Harvey's Legal Agent benchmark (legal), outperforming both its predecessor and Anthropic's Fable 5
- Pricing remains at $0.75 per million input tokens and $3.75 per million output tokens, but effective cost per task rose ~40% due to increased output tokens and agentic turns
- Gemini 3.8 Flash Cyber is a specialized variant with CBRN and cyber offense misuse safeguards, distributed through the Fairwind Program to 650 government and trusted partner organizations
- Google's CodeMender agent, bundled with the Cyber variant, autonomously finds and fixes vulnerabilities in critical infrastructure and national security systems
Industry Insight
- The "cheaper per token but more expensive per task" dynamic may become a recurring theme as agentic models grow more capable, forcing buyers to evaluate total cost of ownership rather than unit pricing
- Google's differentiation strategy increasingly hinges on vertical-specific benchmarks (software engineering, finance, legal) and security-focused variants, suggesting the market is fragmenting along use-case lines
- The Fairwind Program's restricted access model for cybersecurity applications mirrors a growing trend of AI vendors creating government-grade tiers with enhanced safeguards, potentially creating a two-tier ecosystem for sensitive domains
Disclaimer: The above content is generated by AI and is for reference only.