Three weeks. That’s the gap between 3.6 Flash and 3.7 Flash, and it’s the part that actually changes how I work. I’ve been treating model choice like picking a chartplotter — research it once, mount it, forget it for two seasons. That instinct is now wrong. If the workhorse tier turns over monthly, the thing worth building isn’t a pipeline tuned to one model, it’s a pipeline where the model name is a config value and I have an eval I trust.
The pricing is the other tell. Introductory rate of $0.75 per million input and $3.75 output, half what 3.6 cost, and it expires December 31, 2026 — then $1.50 and $7.50. So the cheap window is a window, not a floor. I’d use it to run the batch jobs I keep deferring: chewing through PDFs (they claim 34.0% vs 22.0% on a document-processing eval), backfilling summaries, the unglamorous stuff where volume matters more than polish.
What I’d actually test first is the claim about fewer retries — better instruction-following and multi-step tool calls. For anything agentic and self-hosted, retries are the real cost, not tokens.
The story — Google introduced Gemini 3.7 Flash on Aug 13, 2026, three weeks after 3.6 Flash, pitching it as its most intelligent workhorse model for coding and agents, with gains on FrontierCode 1.1 Main (43.6% vs 34.4%), DeepSWE v1.1 (65.3% vs 49.0%), and WebDev Arena Elo (1588 vs 1538). It powers Gemini Spark for AI Pro and Ultra subscribers and is available via the Gemini API, Google AI Studio, and Antigravity. (Source)