Gemini 3.7 Flash: Benchmarks, Pricing, API Availability and What Changed
Google released Gemini 3.7 Flash on August 13, 2026, only three weeks after Gemini 3.6 Flash. The update is aimed squarely at coding, agentic workflows, web development and document-heavy knowledge work, with Google publishing substantial gains over 3.6 Flash on several software-engineering and automation benchmarks.
The unusual part is pricing: Gemini 3.7 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens on the standard paid API through December 31, 2026. Those rates are temporary and double on January 1, 2027. This guide explains the benchmark results, API availability, model limits, supported tools, pricing tiers and the practical differences from Gemini 3.6 Flash.
Key takeaways
- Release: Gemini 3.7 Flash became generally available on August 13, 2026.
- API model ID: The stable Gemini API identifier is
gemini-3.7-flash. - Introductory pricing: Standard paid usage costs $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026. From January 1, 2027, those rates become $1.50 and $7.50.
- Benchmark gains: Google reports 43.6% on FrontierCode 1.1 Main, 65.3% on DeepSWE v1.1, 1588 Elo on Code Arena and 30.4% on AutomationBench.
- Context: The model accepts up to 1,048,576 input tokens and can return up to 65,536 output tokens.
- Availability: Google distributes it through the Gemini API, Google AI Studio, Google Antigravity, Gemini Enterprise products and Gemini Spark for eligible Google AI Pro and Ultra subscribers.
Gemini 3.7 Flash at a glance
| Specification | Gemini 3.7 Flash |
|---|---|
| Release date | August 13, 2026 |
| Launch stage | Generally available (GA) |
| Stable API model ID | gemini-3.7-flash |
| Input modalities | Text, image, video, audio and PDF |
| Output modality | Text |
| Input token limit | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Thinking levels | Low, medium and high; minimal is not supported |
| Standard paid API price through Dec. 31, 2026 | $0.75 / 1M input tokens; $3.75 / 1M output tokens |
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is the next iteration of Google's Gemini 3 Flash family. Google DeepMind says it is based on Gemini 3.6 Flash and adds algorithmic improvements to the model's core reasoning foundation. The positioning has also become more explicit: Google calls 3.7 Flash its primary "workhorse" for coding and agents rather than a model designed simply to maximize raw throughput.
That distinction matters. Flash models have traditionally been selected because they are faster and cheaper than the largest reasoning models. With 3.7, Google is trying to move more multi-step work into that cost-efficient tier: repository-level coding, browser or computer interaction, function calling, complex document processing and business workflows that require several tool calls before an answer is complete.
If you are coming from the predecessor, Zerlo's Gemini 3.6 Flash overview provides the launch context for the earlier model. The important update now is that 3.7 Flash keeps the same million-token context window and broad tool set while raising performance on many of Google's published agent and coding evaluations.
What changed from Gemini 3.6 Flash?
| Area | Gemini 3.7 Flash change | Practical effect |
|---|---|---|
| Coding | Higher FrontierCode, DeepSWE and Terminal-bench scores | Stronger case for repository work, debugging and long-running coding agents |
| Web development | Code Arena rises from 1538 to 1588 Elo in Google's model card | Better published performance on functional layouts and feature-complete web tasks |
| Enterprise automation | AutomationBench rises from 17.0% to 30.4% | More promising for workflows that combine reasoning, tools and business systems |
| Document reasoning | GDP.pdf rises from 22.0% to 34.0% | Improved published result for difficult PDF-based knowledge work |
| Long context | GDM-MRCR v2 at 128k rises from 91.8% to 97.0% | Better retrieval and reasoning across long inputs in that evaluation |
| Developer behavior | Google says the model adapts better to roadblocks, clarifies intent and follows instructions more faithfully | Potentially fewer retries and less manual oversight in multi-step workflows |
| Pricing | 3.7 launches at a temporary $0.75 input / $3.75 output rate per 1M tokens | Lower short-term cost than the original 3.6 launch rate, although 3.6 is now on the same temporary rate |
The final row is easy to misunderstand. Google describes Gemini 3.7 Flash as launching at half the original Gemini 3.6 Flash cost. However, Google's current model card lists both 3.6 Flash and 3.7 Flash at $0.75 input and $3.75 output per million tokens under the same temporary promotion. So 3.7 is not currently cheaper per standard token than 3.6; the comparison is against 3.6's original launch pricing.
Gemini 3.7 Flash benchmarks
Google's August 2026 model card publishes a broad benchmark set covering software engineering, agentic terminal use, enterprise automation, knowledge work, PDFs, long context, computer use and scientific reasoning. The numbers below compare Gemini 3.7 Flash with Gemini 3.6 Flash using the values in that model card.
| Benchmark | What it measures | Gemini 3.6 Flash | Gemini 3.7 Flash |
|---|---|---|---|
| Artificial Analysis Intelligence Index | Composite model intelligence | 52 | 56 |
| FrontierCode 1.1 Main | Production code quality | 34.4% | 43.6% |
| DeepSWE v1.1 | Long-horizon software engineering | 48.6% | 65.3% |
| Code Arena | Web development | 1538 Elo | 1588 Elo |
| Terminal-bench 2.1 | Agentic terminal coding | 78.0% | 85.8% |
| AutomationBench | Enterprise workflow automation | 17.0% | 30.4% |
| GDP.pdf | Expert PDF comprehension | 22.0% | 34.0% |
| GDM-MRCR v2 at 128k | Long-context retrieval and reasoning | 91.8% | 97.0% |
| OSWorld-2.0 | Agentic computer use | 33.8% | 47.9% |

Source: blog.google
Google's August 2026 evaluation table shows especially large 3.7 Flash gains in long-horizon software engineering, enterprise workflow automation, document comprehension and computer use. The benchmarks measure different workloads and should not be combined into one universal score.
Where the gains are strongest
The clearest movement is in workloads that require the model to keep working rather than answer once. DeepSWE rises from 48.6% to 65.3%, while AutomationBench moves from 17.0% to 30.4%. That supports Google's claim that the model is more useful for agents that need to plan, call tools, inspect results and recover from roadblocks.
Document work also improves substantially in Google's published data. GDP.pdf moves from 22.0% to 34.0%, and the company specifically highlights finance, law and biosciences as knowledge-dense domains where it expects better reasoning. That does not mean the model can replace expert review in regulated or high-stakes work; it means the benchmark result is materially higher than 3.6 Flash on the tested tasks.
3.7 Flash does not win every benchmark
The release is not a clean sweep. On CharXiv without tools, Gemini 3.7 Flash scores 84.5% versus 85.2% for 3.6 Flash; with tools it scores 88.7% versus 89.4%. In Google's cross-model table, GPT-5.6 Terra also remains ahead on some coding and terminal evaluations. This is why a benchmark table should be used to identify promising workloads, not to declare a universal winner.
There is also a small reporting difference worth noting: Google's launch article rounds the Gemini 3.6 Flash DeepSWE result to 49.0%, while the detailed DeepMind model card lists 48.6%. This article uses the more precise model-card value in the comparison table.

Source: blog.google
Google also frames the release around cost per completed software-engineering task, not just price per token. That is the right production metric for agents: retries, reasoning tokens and repeated tool calls can matter more than the headline token rate.
Gemini 3.7 Flash pricing
Gemini 3.7 Flash has a time-limited pricing schedule. The current introductory rates run through December 31, 2026. Google's pricing documentation then doubles the token rates on January 1, 2027.
| Consumption option | Input / 1M tokens through Dec. 31, 2026 | Output / 1M tokens through Dec. 31, 2026 | Input / output from Jan. 1, 2027 |
|---|---|---|---|
| Standard paid | $0.75 | $3.75 | $1.50 / $7.50 |
| Batch | $0.375 | $1.875 | $0.75 / $3.75 |
| Flex | $0.375 | $1.875 | $0.75 / $3.75 |
| Priority | $1.35 | $6.75 | $2.70 / $13.50 |
Output pricing includes thinking tokens. Standard context caching costs $0.075 per million cached tokens during the introductory period, plus $0.50 per million tokens per hour for cache storage; those rates also double in 2027. Google Search grounding and Google Maps grounding can add separate charges after the shared monthly allowance described in Google's pricing documentation.
The Standard API also has a free tier listed as free of charge. That should not be treated as guaranteed production capacity: rate limits depend on the model, project, usage tier and account status, and Google explicitly says published limits are not guaranteed. For paid production workloads, model costs are only one part of the budget; grounding, external APIs, storage and failed or repeated agent steps can change the total.
Example standard API costs
| Monthly token usage | Estimated cost through Dec. 31, 2026 | Estimated cost from Jan. 1, 2027 |
|---|---|---|
| 10M input + 2M output | $15.00 | $30.00 |
| 50M input + 5M output | $56.25 | $112.50 |
| 100M input + 20M output | $150.00 | $300.00 |
These examples include model input and output tokens only. They do not include grounding, caching storage, external services or application infrastructure.
Gemini 3.7 Flash API availability
Gemini 3.7 Flash is not a preview-only announcement. Google Cloud lists gemini-3.7-flash as GA with a release date of August 13, 2026, and the Gemini API documentation identifies the same string as the stable model version.
| User group | Where Gemini 3.7 Flash is available | Important detail |
|---|---|---|
| API developers | Gemini API and Google AI Studio | Use the stable model ID gemini-3.7-flash |
| Android developers | Android Studio | Google lists Android Studio as a launch surface for developers |
| Agent developers | Google Antigravity | Designed for agent-first coding and multi-step workflows |
| Enterprises | Gemini Enterprise Agent Platform and Gemini Enterprise app | Enterprise governance, billing and regional controls apply |
| Individuals | Gemini Spark in the Gemini app | Google says Spark uses 3.7 Flash for Google AI Pro and Ultra subscribers in more than 160 supported countries |
For developers who have not set up access yet, Zerlo's Gemini API key guide covers the Google AI Studio setup. Once access is ready, the critical deployment change is the model identifier; production migration should still include regression tests for tools, schemas, latency and total token use.

Source: pexels.com
Moving to a new model ID is the easy part. The safer migration is to test representative coding and agent tasks in staging, record task success, latency, token usage, retries and tool calls, and only then shift production traffic.
Rate limits are project-specific
There is no single Gemini 3.7 Flash RPM or TPM number that applies to every account. Google measures limits using requests per minute, tokens per minute and requests per day, applies them per project rather than per API key, and adjusts them by usage tier and account status. Google recommends checking the active limits for the project directly in AI Studio. Batch traffic has separate limits; for example, the documented Tier 1 batch queue allows 3,000,000 enqueued tokens for Gemini 3.7 Flash.
Context window, modalities and supported tools
The core API envelope is familiar to anyone already using 3.6 Flash: a 1,048,576-token input context and a maximum 65,536-token text output. The model can ingest text, images, video, audio and PDF documents, but it does not generate images or audio itself.
| Capability | Status in Gemini 3.7 Flash |
|---|---|
| Caching | Supported |
| Code execution | Supported |
| Computer use | Supported in Preview |
| File Search | Supported |
| Function calling | Supported |
| Google Search grounding | Supported |
| Google Maps grounding | Supported |
| Structured outputs | Supported |
| URL context | Supported |
| Thinking levels | Low, medium and high |
| Live API | Not supported |
| Image generation | Not supported |
| Audio generation | Not supported |
One migration trap is the thinking configuration. Gemini 3.7 Flash supports low, medium and high. Google Cloud documents medium as the default in its Enterprise Agent Platform, while explicitly setting minimal returns a validation error. Applications that previously used a minimal setting should map that behavior to a supported level and re-measure latency and output quality.

Source: pexels.com
Gemini 3.7 Flash improves Google's published PDF and knowledge-work results, which is relevant to document-heavy workflows. A larger benchmark score does not remove hallucination risk, so important extracted facts and consequential decisions still need verification.
What the changes mean in practice
For coding agents
Gemini 3.7 Flash has the strongest upgrade story when an application does long-horizon coding rather than isolated snippets. The DeepSWE, FrontierCode and Terminal-bench gains all point in the same direction: the model is more capable of continuing through multi-step engineering tasks. Google also says it is more disciplined about adapting to roadblocks and clarifying intent, which can matter when an agent has permission to edit files or call tools repeatedly.
For document and knowledge workflows
The million-token context window is unchanged, but the quality of working across dense material appears to improve in the release evaluations. That makes 3.7 Flash a stronger candidate for annual reports, legal-document pipelines, scientific literature, internal knowledge bases and mixed PDF workflows. A huge context window is not the same as perfect retrieval, though; retrieval systems, File Search and prompt segmentation can still be more reliable and cheaper than sending everything on every request.
For web development
Google highlights stronger first-pass code accuracy, more functional layouts and fewer prompts to reach feature-complete apps. Code Arena's 1588 Elo score is the clearest published indicator. Teams should still evaluate design adherence, browser behavior, accessibility and framework-specific conventions on their own codebase because an aggregate web benchmark cannot represent every frontend stack.
For high-volume simple tasks
Do not assume 3.7 Flash should replace every cheaper model. If the workload is classification, extraction or templated transformation at very high volume, a Flash-Lite model may still win on cost and throughput. The purpose of 3.7 Flash is to push stronger reasoning and agentic performance into the Flash tier, not to become the lowest-cost option for every request.
Should you migrate from Gemini 3.6 Flash?
| Workload | Recommended direction | Why |
|---|---|---|
| Repository-level coding agent | Strongly test 3.7 Flash | Large published gains on DeepSWE and FrontierCode |
| Business process automation | Test 3.7 Flash | AutomationBench improves from 17.0% to 30.4% |
| PDF and document analysis | Test 3.7 Flash | Higher GDP.pdf and long-context results |
| Existing stable 3.6 Flash pipeline | Benchmark before switching | Current introductory token rates are the same, so quality, latency and task cost should decide |
| High-volume simple extraction | Compare with Flash-Lite | A lighter model may still have a better throughput/cost profile |
| Real-time voice application | Use a Live API-compatible model | Gemini 3.7 Flash does not support Live API |
| Image or audio generation | Use another model | 3.7 Flash produces text output, not generated images or audio |
Migration checklist for Gemini 3.7 Flash
- Change the target model to
gemini-3.7-flashin a staging environment first. - Review the thinking configuration. Do not send
minimal; choose low, medium or high. - Re-run function-calling tests, structured-output validation and tool permission checks.
- Test long-context and PDF workflows with the same representative documents used in production.
- Measure successful task completion, latency, input tokens, output and thinking tokens, retries, turns and tool calls.
- Include grounding and caching charges when comparing total cost per task.
- Roll out gradually and keep a tested rollback path to the existing production model.
Limitations to keep in mind
Google DeepMind's model card lists the standard limitations of foundation models, including hallucinations, and notes that occasional slowness or timeout issues can occur. It gives Gemini 3.7 Flash a March 2026 knowledge cutoff, while warning that some domains may effectively be limited to January 2025. Current information therefore still requires grounding or retrieval from up-to-date sources when freshness matters.
Computer use is marked as Preview, so applications should restrict permissions, validate actions and keep human confirmation for consequential steps. The model also lacks Live API, image generation and audio generation. Finally, the attractive launch pricing is explicitly temporary; applications with meaningful 2027 usage should budget against the post-promotion rates rather than assuming the August 2026 prices will persist.
For a broader model-selection view, Zerlo's Gemini vs. Claude technical comparison provides additional context. Because model generations and prices move quickly, use current vendor documentation for procurement decisions and benchmark the exact models you plan to deploy.
FAQ
When was Gemini 3.7 Flash released?
Google announced and released Gemini 3.7 Flash on August 13, 2026. Google Cloud lists the stable model as generally available rather than preview-only.
What is the Gemini 3.7 Flash API model name?
The stable model identifier is gemini-3.7-flash. Google lists this exact ID in both the Gemini API model documentation and its enterprise model documentation.
How much does Gemini 3.7 Flash cost?
Through December 31, 2026, Standard paid API usage is $0.75 per million input tokens and $3.75 per million output tokens, including thinking tokens. Starting January 1, 2027, those rates rise to $1.50 and $7.50. Batch, Flex and Priority have separate rates.
Is Gemini 3.7 Flash cheaper than Gemini 3.6 Flash?
Not at the current introductory token rates. Google's August 2026 model card lists both Gemini 3.6 Flash and Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Google's "half the cost" launch wording compares 3.7 with the original 3.6 Flash launch price.
Is Gemini 3.7 Flash available in the API?
Yes. It is available through the Gemini API and Google AI Studio with the stable model ID gemini-3.7-flash. Google also distributes it through Antigravity and enterprise Gemini products.
What is the Gemini 3.7 Flash context window?
The input token limit is 1,048,576 tokens and the maximum output is 65,536 tokens. Inputs can include text, images, video, audio and PDF documents; output is text.
Does Gemini 3.7 Flash support the Live API?
No. Google's model documentation marks Live API as not supported. Applications that require native real-time voice interaction should choose a Live API-compatible model.
Is Gemini 3.7 Flash better than Gemini 3.6 Flash?
Google's published evaluations show substantial improvements on several coding, automation, PDF, long-context and computer-use benchmarks, but not every benchmark improves. The practical answer depends on the workload, latency target, tool behavior and cost per successful task, so production teams should run their own representative evaluations.
Bottom line
Gemini 3.7 Flash is a meaningful upgrade for the part of the Flash family that matters most to developers: multi-step coding, agents, automation and document-heavy workflows. Google's strongest published improvements are not small leaderboard nudges; DeepSWE rises from 48.6% to 65.3%, AutomationBench from 17.0% to 30.4% and GDP.pdf from 22.0% to 34.0%, while the model keeps the same million-token context window and broad Gemini tool support.
The pricing story is equally important. Standard API usage is temporarily $0.75 per million input tokens and $3.75 per million output tokens, but those rates double on January 1, 2027. For teams already on Gemini 3.6 Flash, the best reason to migrate is therefore better task performance at the same current introductory token rate, not a permanent price cut. Test 3.7 Flash on real workloads, include tool and retry costs, and make the migration decision based on successful-task cost and reliability rather than the model name alone.