AZ Labs

AI Research

Gemini 3.7 Flash Is Now Generally Available

13 August 20264 min read

Google has released Gemini 3.7 Flash as a stable model for coding, agent workflows and multimodal reasoning. It has a 1,048,576-token input limit, 65,536-token output, and direct API pricing from $0.75 per million input tokens through 2026.

Listen
0:00 / 0:00

AI-narrated summary, not a word-for-word reading.

Official Google DeepMind artwork for Gemini 3.7 Flash

Key takeaways

  • check_circleGoogle lists Gemini 3.7 Flash as generally available and stable, not a preview model.
  • check_circleThe model accepts text, images, video, audio and PDFs, with a 1,048,576-token input limit and 65,536-token output limit.
  • check_circleAIMI first saw the Google API route at 18:42:47 SAST. OpenRouter followed 37 minutes and 4 seconds later, then OpenCode Zen listed the model 2 hours, 31 minutes and 17 seconds after Google.

A stable Flash release for coding and agents

Google added Gemini 3.7 Flash to the Gemini API changelog on 13 August 2026 and marked it generally available. The model page describes it as the next Gemini 3 reasoning model, while Google DeepMind calls it its most intelligent workhorse model yet for coding and agents.

The stable API model code is gemini-3.7-flash. Google supports low, medium and high thinking effort. Minimal effort is not supported and returns an error, which matters when applications reuse reasoning settings across model families.

Context, inputs and tools

Gemini 3.7 Flash accepts text, images, video, audio and PDF input, then returns text. The official limit is 1,048,576 input tokens and 65,536 output tokens.

Google lists caching, code execution, file search, function calling, Google Maps grounding, Search grounding, structured outputs and URL context as supported. Computer use is available in preview. Audio generation, image generation and the Live API are not supported by this model.

Pricing and access

Google lists a free tier for standard Gemini API use. Paid standard pricing is $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. Those rates double on 1 January 2027. Google notes that free-tier content may be used to improve its products, while paid-tier content is not.

Batch and Flex inference cost $0.375 per million input tokens and $1.875 per million output tokens through 31 December. OpenRouter currently lists both google/gemini-3.7-flash and google/gemini-3.7-flash:batch at those lower rates. Provider pricing can change independently, so check the active route before sending a large workload.

OpenCode Zen now lists gemini-3.7-flash in its official model registry. That confirms route availability, but the registry does not state route-specific pricing, limits or capability flags. Those details remain unverified for OpenCode Zen.

How the release surfaced today

AIMI first observed models/gemini-3.7-flash on Google's official models endpoint at 16:42:47 UTC, or 18:42:47 SAST. At that point, the public catalogue and changelog had not yet confirmed a maker release, so AZ Labs kept it as a candidate rather than publishing it.

OpenRouter created its route metadata 20 minutes and 14 seconds after the first Google endpoint sighting. AIMI first observed the OpenRouter standard and batch aliases at 17:19:51 UTC, exactly 37 minutes and 4 seconds after the Google route. Google's public model catalogue, changelog and DeepMind model page now confirm the release. The changelog gives a date but no announcement time, so the exact GA timestamp remains unknown.

AIMI first saw gemini-3.7-flash on OpenCode Zen at 19:14:04 UTC, or 21:14:04 SAST. That was 2 hours, 31 minutes and 17 seconds after the Google route, and 1 hour, 54 minutes and 13 seconds after AIMI saw the OpenRouter aliases. This is another provider route for the same model, not a second maker release.

Frequently asked questions

Is Gemini 3.7 Flash a preview model?

No. Google marks gemini-3.7-flash as stable and generally available. Computer use remains a preview capability inside the stable model.

What is the context window?

The official input limit is 1,048,576 tokens and the output limit is 65,536 tokens.

What does Gemini 3.7 Flash cost?

Google offers a free tier. Paid standard API pricing is $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. Batch and Flex cost half those rates.

Where is it available?

Directly through the Gemini API as gemini-3.7-flash, through OpenRouter as google/gemini-3.7-flash or google/gemini-3.7-flash:batch, and through OpenCode Zen as gemini-3.7-flash.

Sources

Copyright & image credits

The article image is credited to © Google DeepMind, used for news reporting. Source: Google DeepMind: Gemini 3.7 Flash.

Share this guidePost on X