Gemini 3.7 Flash
Google's fast workhorse model for coding, agents, and everyday knowledge work at a budget price.
Quality
Modality
multimodal
Context
1M tokens
Access
closed
Fabian's Take
"Flash sits in a curious spot: clearly behind the frontier models on hard problems, yet so fast and cheap that I stopped caring for a whole class of work. I use it most for experimenting with websites, where it iterates faster than I can review. Know where its ceiling is and it's a very good deal."
Gemini 3.7 Flash is Google’s self-described “most intelligent workhorse model,” released in August 2026 just weeks after 3.6 Flash. The pitch is speed and cost rather than peak intelligence: it keeps the instant response times the Flash line is known for while making large jumps on coding and agent benchmarks over its predecessor.
What changed
The gains concentrate exactly where workhorse models earn their keep: software engineering, agentic tasks, and web development. It ships with tunable thinking levels (low, medium, high), so you can trade a little speed for more careful answers, and a computer-use capability in preview. It is available in the Gemini app, Google AI Studio, the Gemini API, and as the engine inside Google Antigravity.
How to use it
Treat Flash as the model for volume, and know where the volume ends. Rapid UI iterations, multi-language copy, analyzing screenshots and mockups, and small coding tasks all land in its sweet spot, and the introductory API pricing makes it one of the cheapest capable models you can call. For genuinely hard logic or high-stakes reasoning, step up to a frontier model and let Flash handle everything around it.
The Verdict
Best for: Fast, cheap execution of simple-to-medium tasks: coding, agents, copy, and UI/UX analysis.
Pros
- Exceptionally fast response times
- Large benchmark gains over 3.5/3.6 Flash on software engineering and agent tasks
- Tunable thinking levels and computer-use support (preview)
- Very aggressive introductory pricing
Cons
- Clearly behind the frontier models on genuinely hard reasoning
- Output capped at 64K tokens
Specs
- Pricing $0.75/M input, $3.75/M output (introductory until end of 2026, then $1.50/$7.50)
- Cost Tier budget
- ⚡Speed Tier instant
- License Proprietary