Gemini 3.7 Flash: Benchmarks, Pricing, API Availability and What Changed

Avatar
Lisa Ernst · 17.08.2026 · Artificial Intelligence · 15 min read

Google released Gemini 3.7 Flash on August 13, 2026, only three weeks after Gemini 3.6 Flash. The update is aimed squarely at coding, agentic workflows, web development and document-heavy knowledge work, with Google publishing substantial gains over 3.6 Flash on several software-engineering and automation benchmarks.

The unusual part is pricing: Gemini 3.7 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens on the standard paid API through December 31, 2026. Those rates are temporary and double on January 1, 2027. This guide explains the benchmark results, API availability, model limits, supported tools, pricing tiers and the practical differences from Gemini 3.6 Flash.

Key takeaways

Gemini 3.7 Flash at a glance

Specification Gemini 3.7 Flash
Release date August 13, 2026
Launch stage Generally available (GA)
Stable API model ID gemini-3.7-flash
Input modalities Text, image, video, audio and PDF
Output modality Text
Input token limit 1,048,576 tokens
Maximum output 65,536 tokens
Thinking levels Low, medium and high; minimal is not supported
Standard paid API price through Dec. 31, 2026 $0.75 / 1M input tokens; $3.75 / 1M output tokens

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is the next iteration of Google's Gemini 3 Flash family. Google DeepMind says it is based on Gemini 3.6 Flash and adds algorithmic improvements to the model's core reasoning foundation. The positioning has also become more explicit: Google calls 3.7 Flash its primary "workhorse" for coding and agents rather than a model designed simply to maximize raw throughput.

That distinction matters. Flash models have traditionally been selected because they are faster and cheaper than the largest reasoning models. With 3.7, Google is trying to move more multi-step work into that cost-efficient tier: repository-level coding, browser or computer interaction, function calling, complex document processing and business workflows that require several tool calls before an answer is complete.

If you are coming from the predecessor, Zerlo's Gemini 3.6 Flash overview provides the launch context for the earlier model. The important update now is that 3.7 Flash keeps the same million-token context window and broad tool set while raising performance on many of Google's published agent and coding evaluations.

What changed from Gemini 3.6 Flash?

Area Gemini 3.7 Flash change Practical effect
Coding Higher FrontierCode, DeepSWE and Terminal-bench scores Stronger case for repository work, debugging and long-running coding agents
Web development Code Arena rises from 1538 to 1588 Elo in Google's model card Better published performance on functional layouts and feature-complete web tasks
Enterprise automation AutomationBench rises from 17.0% to 30.4% More promising for workflows that combine reasoning, tools and business systems
Document reasoning GDP.pdf rises from 22.0% to 34.0% Improved published result for difficult PDF-based knowledge work
Long context GDM-MRCR v2 at 128k rises from 91.8% to 97.0% Better retrieval and reasoning across long inputs in that evaluation
Developer behavior Google says the model adapts better to roadblocks, clarifies intent and follows instructions more faithfully Potentially fewer retries and less manual oversight in multi-step workflows
Pricing 3.7 launches at a temporary $0.75 input / $3.75 output rate per 1M tokens Lower short-term cost than the original 3.6 launch rate, although 3.6 is now on the same temporary rate

The final row is easy to misunderstand. Google describes Gemini 3.7 Flash as launching at half the original Gemini 3.6 Flash cost. However, Google's current model card lists both 3.6 Flash and 3.7 Flash at $0.75 input and $3.75 output per million tokens under the same temporary promotion. So 3.7 is not currently cheaper per standard token than 3.6; the comparison is against 3.6's original launch pricing.

Gemini 3.7 Flash benchmarks

Google's August 2026 model card publishes a broad benchmark set covering software engineering, agentic terminal use, enterprise automation, knowledge work, PDFs, long context, computer use and scientific reasoning. The numbers below compare Gemini 3.7 Flash with Gemini 3.6 Flash using the values in that model card.

Benchmark What it measures Gemini 3.6 Flash Gemini 3.7 Flash
Artificial Analysis Intelligence Index Composite model intelligence 52 56
FrontierCode 1.1 Main Production code quality 34.4% 43.6%
DeepSWE v1.1 Long-horizon software engineering 48.6% 65.3%
Code Arena Web development 1538 Elo 1588 Elo
Terminal-bench 2.1 Agentic terminal coding 78.0% 85.8%
AutomationBench Enterprise workflow automation 17.0% 30.4%
GDP.pdf Expert PDF comprehension 22.0% 34.0%
GDM-MRCR v2 at 128k Long-context retrieval and reasoning 91.8% 97.0%
OSWorld-2.0 Agentic computer use 33.8% 47.9%
Google benchmark table comparing Gemini 3.7 Flash with Gemini 3.6 Flash and other AI models

Source: blog.google

Google's August 2026 evaluation table shows especially large 3.7 Flash gains in long-horizon software engineering, enterprise workflow automation, document comprehension and computer use. The benchmarks measure different workloads and should not be combined into one universal score.

Where the gains are strongest

The clearest movement is in workloads that require the model to keep working rather than answer once. DeepSWE rises from 48.6% to 65.3%, while AutomationBench moves from 17.0% to 30.4%. That supports Google's claim that the model is more useful for agents that need to plan, call tools, inspect results and recover from roadblocks.

Document work also improves substantially in Google's published data. GDP.pdf moves from 22.0% to 34.0%, and the company specifically highlights finance, law and biosciences as knowledge-dense domains where it expects better reasoning. That does not mean the model can replace expert review in regulated or high-stakes work; it means the benchmark result is materially higher than 3.6 Flash on the tested tasks.

3.7 Flash does not win every benchmark

The release is not a clean sweep. On CharXiv without tools, Gemini 3.7 Flash scores 84.5% versus 85.2% for 3.6 Flash; with tools it scores 88.7% versus 89.4%. In Google's cross-model table, GPT-5.6 Terra also remains ahead on some coding and terminal evaluations. This is why a benchmark table should be used to identify promising workloads, not to declare a universal winner.

There is also a small reporting difference worth noting: Google's launch article rounds the Gemini 3.6 Flash DeepSWE result to 49.0%, while the detailed DeepMind model card lists 48.6%. This article uses the more precise model-card value in the comparison table.

Google cost-versus-performance chart for Gemini 3.7 Flash on DeepSWE v1.1

Source: blog.google

Google also frames the release around cost per completed software-engineering task, not just price per token. That is the right production metric for agents: retries, reasoning tokens and repeated tool calls can matter more than the headline token rate.

Gemini 3.7 Flash pricing

Gemini 3.7 Flash has a time-limited pricing schedule. The current introductory rates run through December 31, 2026. Google's pricing documentation then doubles the token rates on January 1, 2027.

Consumption option Input / 1M tokens through Dec. 31, 2026 Output / 1M tokens through Dec. 31, 2026 Input / output from Jan. 1, 2027
Standard paid $0.75 $3.75 $1.50 / $7.50
Batch $0.375 $1.875 $0.75 / $3.75
Flex $0.375 $1.875 $0.75 / $3.75
Priority $1.35 $6.75 $2.70 / $13.50

Output pricing includes thinking tokens. Standard context caching costs $0.075 per million cached tokens during the introductory period, plus $0.50 per million tokens per hour for cache storage; those rates also double in 2027. Google Search grounding and Google Maps grounding can add separate charges after the shared monthly allowance described in Google's pricing documentation.

The Standard API also has a free tier listed as free of charge. That should not be treated as guaranteed production capacity: rate limits depend on the model, project, usage tier and account status, and Google explicitly says published limits are not guaranteed. For paid production workloads, model costs are only one part of the budget; grounding, external APIs, storage and failed or repeated agent steps can change the total.

Example standard API costs

Monthly token usage Estimated cost through Dec. 31, 2026 Estimated cost from Jan. 1, 2027
10M input + 2M output $15.00 $30.00
50M input + 5M output $56.25 $112.50
100M input + 20M output $150.00 $300.00

These examples include model input and output tokens only. They do not include grounding, caching storage, external services or application infrastructure.

Gemini 3.7 Flash API availability

Gemini 3.7 Flash is not a preview-only announcement. Google Cloud lists gemini-3.7-flash as GA with a release date of August 13, 2026, and the Gemini API documentation identifies the same string as the stable model version.

User group Where Gemini 3.7 Flash is available Important detail
API developers Gemini API and Google AI Studio Use the stable model ID gemini-3.7-flash
Android developers Android Studio Google lists Android Studio as a launch surface for developers
Agent developers Google Antigravity Designed for agent-first coding and multi-step workflows
Enterprises Gemini Enterprise Agent Platform and Gemini Enterprise app Enterprise governance, billing and regional controls apply
Individuals Gemini Spark in the Gemini app Google says Spark uses 3.7 Flash for Google AI Pro and Ultra subscribers in more than 160 supported countries

For developers who have not set up access yet, Zerlo's Gemini API key guide covers the Google AI Studio setup. Once access is ready, the critical deployment change is the model identifier; production migration should still include regression tests for tools, schemas, latency and total token use.

Developer working with source code on a laptop

Source: pexels.com

Moving to a new model ID is the easy part. The safer migration is to test representative coding and agent tasks in staging, record task success, latency, token usage, retries and tool calls, and only then shift production traffic.

Rate limits are project-specific

There is no single Gemini 3.7 Flash RPM or TPM number that applies to every account. Google measures limits using requests per minute, tokens per minute and requests per day, applies them per project rather than per API key, and adjusts them by usage tier and account status. Google recommends checking the active limits for the project directly in AI Studio. Batch traffic has separate limits; for example, the documented Tier 1 batch queue allows 3,000,000 enqueued tokens for Gemini 3.7 Flash.

Context window, modalities and supported tools

The core API envelope is familiar to anyone already using 3.6 Flash: a 1,048,576-token input context and a maximum 65,536-token text output. The model can ingest text, images, video, audio and PDF documents, but it does not generate images or audio itself.

Capability Status in Gemini 3.7 Flash
Caching Supported
Code execution Supported
Computer use Supported in Preview
File Search Supported
Function calling Supported
Google Search grounding Supported
Google Maps grounding Supported
Structured outputs Supported
URL context Supported
Thinking levels Low, medium and high
Live API Not supported
Image generation Not supported
Audio generation Not supported

One migration trap is the thinking configuration. Gemini 3.7 Flash supports low, medium and high. Google Cloud documents medium as the default in its Enterprise Agent Platform, while explicitly setting minimal returns a validation error. Applications that previously used a minimal setting should map that behavior to a supported level and re-measure latency and output quality.

Office workers reviewing documents beside a laptop

Source: pexels.com

Gemini 3.7 Flash improves Google's published PDF and knowledge-work results, which is relevant to document-heavy workflows. A larger benchmark score does not remove hallucination risk, so important extracted facts and consequential decisions still need verification.

What the changes mean in practice

For coding agents

Gemini 3.7 Flash has the strongest upgrade story when an application does long-horizon coding rather than isolated snippets. The DeepSWE, FrontierCode and Terminal-bench gains all point in the same direction: the model is more capable of continuing through multi-step engineering tasks. Google also says it is more disciplined about adapting to roadblocks and clarifying intent, which can matter when an agent has permission to edit files or call tools repeatedly.

For document and knowledge workflows

The million-token context window is unchanged, but the quality of working across dense material appears to improve in the release evaluations. That makes 3.7 Flash a stronger candidate for annual reports, legal-document pipelines, scientific literature, internal knowledge bases and mixed PDF workflows. A huge context window is not the same as perfect retrieval, though; retrieval systems, File Search and prompt segmentation can still be more reliable and cheaper than sending everything on every request.

For web development

Google highlights stronger first-pass code accuracy, more functional layouts and fewer prompts to reach feature-complete apps. Code Arena's 1588 Elo score is the clearest published indicator. Teams should still evaluate design adherence, browser behavior, accessibility and framework-specific conventions on their own codebase because an aggregate web benchmark cannot represent every frontend stack.

For high-volume simple tasks

Do not assume 3.7 Flash should replace every cheaper model. If the workload is classification, extraction or templated transformation at very high volume, a Flash-Lite model may still win on cost and throughput. The purpose of 3.7 Flash is to push stronger reasoning and agentic performance into the Flash tier, not to become the lowest-cost option for every request.

Should you migrate from Gemini 3.6 Flash?

Workload Recommended direction Why
Repository-level coding agent Strongly test 3.7 Flash Large published gains on DeepSWE and FrontierCode
Business process automation Test 3.7 Flash AutomationBench improves from 17.0% to 30.4%
PDF and document analysis Test 3.7 Flash Higher GDP.pdf and long-context results
Existing stable 3.6 Flash pipeline Benchmark before switching Current introductory token rates are the same, so quality, latency and task cost should decide
High-volume simple extraction Compare with Flash-Lite A lighter model may still have a better throughput/cost profile
Real-time voice application Use a Live API-compatible model Gemini 3.7 Flash does not support Live API
Image or audio generation Use another model 3.7 Flash produces text output, not generated images or audio

Migration checklist for Gemini 3.7 Flash

  1. Change the target model to gemini-3.7-flash in a staging environment first.
  2. Review the thinking configuration. Do not send minimal; choose low, medium or high.
  3. Re-run function-calling tests, structured-output validation and tool permission checks.
  4. Test long-context and PDF workflows with the same representative documents used in production.
  5. Measure successful task completion, latency, input tokens, output and thinking tokens, retries, turns and tool calls.
  6. Include grounding and caching charges when comparing total cost per task.
  7. Roll out gradually and keep a tested rollback path to the existing production model.

Limitations to keep in mind

Google DeepMind's model card lists the standard limitations of foundation models, including hallucinations, and notes that occasional slowness or timeout issues can occur. It gives Gemini 3.7 Flash a March 2026 knowledge cutoff, while warning that some domains may effectively be limited to January 2025. Current information therefore still requires grounding or retrieval from up-to-date sources when freshness matters.

Computer use is marked as Preview, so applications should restrict permissions, validate actions and keep human confirmation for consequential steps. The model also lacks Live API, image generation and audio generation. Finally, the attractive launch pricing is explicitly temporary; applications with meaningful 2027 usage should budget against the post-promotion rates rather than assuming the August 2026 prices will persist.

For a broader model-selection view, Zerlo's Gemini vs. Claude technical comparison provides additional context. Because model generations and prices move quickly, use current vendor documentation for procurement decisions and benchmark the exact models you plan to deploy.

FAQ

When was Gemini 3.7 Flash released?

Google announced and released Gemini 3.7 Flash on August 13, 2026. Google Cloud lists the stable model as generally available rather than preview-only.

What is the Gemini 3.7 Flash API model name?

The stable model identifier is gemini-3.7-flash. Google lists this exact ID in both the Gemini API model documentation and its enterprise model documentation.

How much does Gemini 3.7 Flash cost?

Through December 31, 2026, Standard paid API usage is $0.75 per million input tokens and $3.75 per million output tokens, including thinking tokens. Starting January 1, 2027, those rates rise to $1.50 and $7.50. Batch, Flex and Priority have separate rates.

Is Gemini 3.7 Flash cheaper than Gemini 3.6 Flash?

Not at the current introductory token rates. Google's August 2026 model card lists both Gemini 3.6 Flash and Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Google's "half the cost" launch wording compares 3.7 with the original 3.6 Flash launch price.

Is Gemini 3.7 Flash available in the API?

Yes. It is available through the Gemini API and Google AI Studio with the stable model ID gemini-3.7-flash. Google also distributes it through Antigravity and enterprise Gemini products.

What is the Gemini 3.7 Flash context window?

The input token limit is 1,048,576 tokens and the maximum output is 65,536 tokens. Inputs can include text, images, video, audio and PDF documents; output is text.

Does Gemini 3.7 Flash support the Live API?

No. Google's model documentation marks Live API as not supported. Applications that require native real-time voice interaction should choose a Live API-compatible model.

Is Gemini 3.7 Flash better than Gemini 3.6 Flash?

Google's published evaluations show substantial improvements on several coding, automation, PDF, long-context and computer-use benchmarks, but not every benchmark improves. The practical answer depends on the workload, latency target, tool behavior and cost per successful task, so production teams should run their own representative evaluations.

Bottom line

Gemini 3.7 Flash is a meaningful upgrade for the part of the Flash family that matters most to developers: multi-step coding, agents, automation and document-heavy workflows. Google's strongest published improvements are not small leaderboard nudges; DeepSWE rises from 48.6% to 65.3%, AutomationBench from 17.0% to 30.4% and GDP.pdf from 22.0% to 34.0%, while the model keeps the same million-token context window and broad Gemini tool support.

The pricing story is equally important. Standard API usage is temporarily $0.75 per million input tokens and $3.75 per million output tokens, but those rates double on January 1, 2027. For teams already on Gemini 3.6 Flash, the best reason to migrate is therefore better task performance at the same current introductory token rate, not a permanent price cut. Test 3.7 Flash on real workloads, include tool and retry costs, and make the migration decision based on successful-task cost and reliability rather than the model name alone.

Share our post!
Sources