Gemini 3.8 Flash
Google's fast workhorse model, with gains on coding and agents over 3.7 Flash and the same introductory price.
Quality
Modality
multimodal
Context
1M tokens
Access
closed
Fabian's Take
"On raw model capability Google has been left behind by OpenAI and Anthropic, and that matters less than it sounds. A good harness and good connections to everything else you use are often worth more than the last few points of capability. I use a lot of Google services, so Flash is an inexpensive, fast, capable coworker, especially inside the Antigravity harness. It's also excellent at watching and reviewing videos for me now."
Gemini 3.8 Flash landed on 2 September 2026, three weeks after 3.7 Flash. Google kept the pitch the same: speed and price rather than peak intelligence, with the gains concentrated in software engineering, agentic tasks, and multi-step reasoning.
The price to watch
$0.75 per million input tokens and $3.75 per million output, which is what 3.7 Flash cost. The number to put in your calendar is 1 January 2027, when the introductory rate ends and both figures double to $1.50 and $7.50. If you’re sizing a workload on Flash economics, size it on the 2027 price.
The Cyber variant
Google shipped a sibling, 3.8 Flash Cyber, on the same day. It runs with deliberately looser mitigations for security work and is restricted to trusted defenders through the new Fairwind Program, which covers government authorities, critical infrastructure operators, and software maintainers who apply and get approved. It isn’t something a reader can pick in the Gemini app, which is why it doesn’t get its own entry here.
Where it fits
On raw capability, Google has been left behind by the frontier models from OpenAI and Anthropic. That matters less than the benchmark tables imply. How clever the model is turns out to be only one part of a working setup. The other part is the software wrapped around it: what it can see, what it can reach, which of your other tools it can act on. In AI circles that wrapper gets called the harness, and for day-to-day work it’s often worth more than the last few points of raw capability.
That’s why Flash works as a coworker for me specifically. I use a lot of Google services, so the model already sits next to the things I work on, and inside Antigravity it’s inexpensive, fast, and capable enough. It’s also become excellent at watching and reviewing video, which is a job most models are still bad at.
The rest of the advice for Flash hasn’t changed: use it for volume, and know where the volume ends. For hard logic or high-stakes reasoning, step up to a frontier model and let Flash handle everything around it.
The Verdict
Best for: High-volume execution work: coding, agents, copy, and anything where speed and price beat peak intelligence.
Pros
- Gains over 3.7 Flash on software engineering, agentic tasks, and multi-step reasoning
- Introductory pricing holds until the end of 2026
- Instant response times at 1M tokens of context
- Strong at watching and reviewing video
- Already sits inside the Google services and tools you may be working in anyway
Cons
- Prices double on 1 January 2027, to $1.50/$7.50
- Output still capped at 64K tokens
- Behind the frontier models on genuinely hard reasoning
Specs
- Pricing $0.75/M input, $3.75/M output (introductory until 2026-12-31, then $1.50/$7.50 from 2027-01-01)
- Cost Tier budget
- ⚡Speed Tier instant
- License Proprietary