Gemini 3.6 Flash's headline scores barely moved. Its coding and agent benchmarks moved a lot.
Artificial Analysis's Intelligence Index puts Gemini 3.6 Flash level with the model it replaces, 50.1 against 50.2. Underneath that, agentic coding is up 15.1 points on GBENCH and 11.2 on DeepSWE, social reasoning is down 19.1, and the model runs 35.5% faster for 16.7% less per output token.
Event date Published
Gemini 3.6 Flash arrives faster and cheaper than the model it replaces
Google released Gemini 3.6 Flash and Gemini 3.5 Flash Lite on 2026-07-21. Compared with Gemini 3.5 Flash, released 2026-05-19, AA Intelligence Index is effectively flat at 50.2 versus 50.1. The throughput and price changes are unambiguous, and the task-level benchmarks below are where the model actually changed.
- Median output speed
- 248 tokens/sec Artificial Analysis median across served endpoints, against 183 tokens/sec for Gemini 3.5 Flash — 35.5% faster.
- Output price
- $7.50 per 1M tokens Down 16.7% from $9.00 for Gemini 3.5 Flash. Input pricing is unchanged at $1.50 per 1M tokens.
- Context window
- 1,000,000 tokens Artificial Analysis reported, unchanged from Gemini 3.5 Flash.
- LMArena WebDev Elo
- 1,537 Tenth of 72 models with a reported LMArena WebDev Elo, behind Kimi K3 at 1678 and Claude Fable 5 at 1634. This is the WebDev arena specifically, not coding in general. Gemini 3.5 Flash has no reported WebDev Elo, so this is not a generational comparison.
Task-level benchmarks, Gemini 3.5 Flash to 3.6 Flash
GBENCH Agentic Coding
DeepSWE
Vals Vibe Code
AutomationBench-AA
LiveBench Agentic Coding
GBENCH Social Intelligence
| Panel | Item | Score |
|---|---|---|
| GBENCH Agentic Coding | Gemini 3.5 Flash | 60.4 |
| GBENCH Agentic Coding | Gemini 3.6 Flash | 75.5 |
| DeepSWE | Gemini 3.5 Flash | 37.4 |
| DeepSWE | Gemini 3.6 Flash | 48.6 |
| Vals Vibe Code | Gemini 3.5 Flash | 48.7 |
| Vals Vibe Code | Gemini 3.6 Flash | 57.3 |
| AutomationBench-AA | Gemini 3.5 Flash | 42.6 |
| AutomationBench-AA | Gemini 3.6 Flash | 51.1 |
| LiveBench Agentic Coding | Gemini 3.5 Flash | 49 |
| LiveBench Agentic Coding | Gemini 3.6 Flash | 43.4 |
| GBENCH Social Intelligence | Gemini 3.5 Flash | 81.6 |
| GBENCH Social Intelligence | Gemini 3.6 Flash | 62.5 |
The four largest gains and the two largest losses among the benchmarks both models report, each on a 0 to 100 axis and each reported by a single reporter for both models. Note that the two agentic-coding measures disagree in direction: GBENCH's rises 15.1 points while LiveBench's falls 5.6.
What the aggregate index hides
Each pair compares Gemini 3.5 Flash to Gemini 3.6 Flash on the same benchmark and the same reporter, so the differences are like-for-like. Each difference is the change between the two rounded values shown. The Intelligence Index row at the bottom is the aggregate most people will see first.
Gemini 3.5 Flash Lite is the cheaper, faster tier
The second model in the same 2026-07-21 release is a genuinely different price point rather than a variant of Flash. It runs 450 tokens per second at $0.30 in and $2.50 out per million tokens. It sits well below both Flash models on AA Intelligence Index, though it is not uniformly worse: it scores higher than 3.6 Flash on hallucination resistance, at 66.5 against 46.5 on AA's non-hallucination measure.
- AA Intelligence Index
- 36.5 Against 50.1 for Gemini 3.6 Flash and 50.2 for Gemini 3.5 Flash. Artificial Analysis reported.
- LiveBench
- 63.9 Against 73.6 for Gemini 3.6 Flash. Thirty-third of 34 models with a reported LiveBench score, 2026-06-25 release.
- Median output speed
- 450 tokens/sec Artificial Analysis median across served endpoints — roughly 1.8x Gemini 3.6 Flash's 248.
What this means if you are on Flash today
The Intelligence Index barely changed; the task benchmarks say the coding and agent workloads did. If that is what you run, Gemini 3.6 Flash is worth testing on your own harness — four independent coding and agent benchmarks gained between 8.5 and 15.1 points, and it is 35.5% faster at 16.7% less per output token. Two cautions: LiveBench's agentic-coding measure moved the other way, so the agent gains are not unanimous, and GBENCH's social-intelligence measure dropped 19.1 points, the largest move in this comparison. None of these results speak to reliability or long-context behavior.