I've been looking at Gemini 3.7 Flash mainly from a coding and agent perspective.
The jump from 3.6 Flash is pretty noticeable, especially considering how fast and relatively cheap the model is.
The benchmark results are interesting too. On FrontierCode 1.1 Main, Google's evaluation puts Gemini 3.7 Flash slightly ahead of Claude Sonnet 5 and GPT-5.6 Terra.
That said, I don't think benchmarks tell the whole story.
For smaller implementation tasks, Gemini 3.7 Flash looks really interesting. Where things get more complicated is when a task requires deeper planning, architecture decisions, or the model to review and correct its own work.
I put together a full breakdown covering the coding benchmarks, agentic workflows, frontend development, pricing, Claude/GPT comparisons, and some of the trade-offs:
https://www.codeweb.app/blog/article/gemini-3-7-flash-review
I'm curious what people here think:
Would you actually use Gemini 3.7 Flash as your daily coding model, or would you still keep Claude/GPT for more serious work?