Model Update2026-09-03The Verge

Google Launches Gemini 3.8 Flash, Promises Harder Work

Google has unveiled Gemini 3.8 Flash, a new addition to its AI model lineup that arrives just weeks after the release of its predecessor, 3.7 Flash. The tech giant is positioning this iteration as a 'harder worker,' emphasizing its ability to perform deeper reasoning chains and iteratively call external tools to solve complex, multi-step problems. According to Google's announcement, the model is designed to think longer before responding, which should translate to higher accuracy on tasks involving math, coding, and logical deduction. Perhaps the most striking aspect of this release is the pricing strategy. Google has maintained the introductory rate of $0.75 per million input tokens, matching the 3.7 Flash tier. However, the company warns that the cost per successful query could rise significantly. Because the model now engages in more extensive internal deliberation and tool usage, a single complex request might consume multiple times the token volume of a simpler prompt. This 'thinking tax' means developers will need to carefully monitor their usage patterns, as a query that previously cost a few cents could now cost several dollars depending on the complexity involved. The launch signals a competitive shift in the industry, where the focus is moving from raw speed to 'cognitive effort.' Google is betting that users will accept higher variable costs in exchange for fewer errors and less need for prompt engineering. Early benchmarks suggest that 3.8 Flash outperforms its predecessor on standard reasoning tests by a significant margin. For businesses integrating AI into their workflows, this model offers a trade-off: budget for higher operational costs, but potentially save on human review time. As the AI race intensifies, Google is clearly prioritizing capability over cost-per-token as the primary marketing metric.

Related news