Gemini 3.7 Flash Released: Specs, Benchmarks, Pricing, and Model Comparison
A fact-checked Gemini 3.7 Flash analysis covering its 1M context, 64K output, capabilities, official pricing, full comparison with 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2, plus API migration guidance.
Google released Gemini 3.7 Flash on August 13, 2026, only three weeks after Gemini 3.6 Flash. It is generally available under the stable API model ID gemini-3.7-flash and is positioned as Google's most capable Flash model for coding, agents, multimodal reasoning, and reliable multi-step execution.
The headline is not a larger context window or a disclosed parameter count. Gemini 3.7 Flash keeps the 1M-input and 64K-output class of 3.6 Flash, while Google reports substantial gains in coding, tool use, knowledge work, document understanding, and computer use. Its introductory token price is also unchanged from 3.6 Flash.
This analysis separates official specifications, Google-published benchmark results, temporary launch pricing, and current Modelflare availability. Volatile facts were checked on August 15, 2026.
Gemini 3.7 Flash at a glance
| Property | Official value |
|---|---|
| Release date | August 13, 2026 |
| API status | Generally available; stable model ID |
| Model ID | gemini-3.7-flash |
| Foundation | Based on Gemini 3.6 Flash with algorithmic improvements |
| Input modalities | Text, image, video, audio, and PDF |
| Output modality | Text |
| Maximum input | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Thinking levels | low, medium, high; medium is the documented default |
| Knowledge cutoff | March 2026, with some domains potentially limited to January 2025 |
| Parameter count | Not disclosed |
| Latest documented update | August 2026 |
Supported capabilities include context caching, code execution, Computer Use in preview, file search, function calling, Google Maps grounding, Google Search grounding, structured outputs, thinking, and URL context. Batch, Flex, and Priority consumption modes are available. Audio generation, image generation, and the Live API are not supported by this model.
What actually changed from Gemini 3.6 Flash
Gemini 3.7 Flash is an iteration on 3.6 Flash rather than a new context or modality class. Google attributes the release to algorithmic improvements in the reasoning foundation and feedback from production developers.
The practical changes are concentrated in:
- higher first-pass accuracy for code generation and software engineering;
- stronger design adherence when turning screenshots, images, or design systems into web interfaces;
- better multi-step planning, tool calls, recovery from roadblocks, and intent clarification;
- better reasoning over complex documents in finance, law, and biosciences;
- improved long-context retrieval and agentic computer use;
- the same introductory Standard price as Gemini 3.6 Flash through December 31, 2026.
The last point matters. At launch, 3.7 is not a higher-priced replacement for 3.6. The migration decision is mainly about behavior, latency, and compatibility, not a token-price premium.
Official API pricing
The following paid-tier prices are in US dollars. Token prices are per 1 million tokens; cache storage is per 1 million tokens per hour.
| Consumption mode | Meter | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|---|
| Standard | Input | $0.75 | $1.50 |
| Standard | Output, including thinking tokens | $3.75 | $7.50 |
| Standard | Cached input | $0.075 | $0.15 |
| Standard | Cache storage per hour | $0.50 | $1.00 |
| Batch | Input | $0.375 | $0.75 |
| Batch | Output, including thinking tokens | $1.875 | $3.75 |
| Batch | Cached input | $0.0375 | $0.075 |
| Flex | Input | $0.375 | $0.75 |
| Flex | Output, including thinking tokens | $1.875 | $3.75 |
| Flex | Cached input | $0.0375 | $0.075 |
| Priority | Input | $1.35 | $2.70 |
| Priority | Output, including thinking tokens | $6.75 | $13.50 |
| Priority | Cached input | $0.135 | $0.27 |
Google Search grounding includes 5,000 free search requests per month shared across Gemini 3.x models, then costs $14 per 1,000 requests. Google Maps grounding uses a shared allowance of 5,000 prompts and then the same $14 per 1,000 search queries.
At the introductory Standard rate, 100K input tokens plus 10K output tokens cost about $0.1125 before caching, tools, storage, or retries. One million cached input tokens plus 100K output tokens cost about $0.45, excluding cache storage. Both examples double after the introductory period.
Full horizontal benchmark comparison
Google's August 2026 model card compares Gemini 3.7 Flash with Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Higher is better for every row below. A dash means the model card did not publish a result.
| Benchmark | Gemini 3.7 Flash | Gemini 3.6 Flash | Claude Sonnet 5 | GPT-5.6 Terra | Muse Spark 1.2 |
|---|---|---|---|---|---|
| Artificial Analysis Intelligence Index | 56 | 52 | 55 | 57 | 57 |
| FrontierCode 1.1 Main | 43.6% | 34.4% | 42.7% | 41.3% | — |
| DeepSWE v1.1 | 65.3% | 48.6% | 53.8% | 69.6% | 54.9% |
| Code Arena, WebDev Elo | 1588 | 1538 | 1541 | 1523 | 1535 |
| Terminal-bench 2.1 | 85.8% | 78.0% | 80.4% | 87.4% | 82.9% |
| Terminal-bench 3.0 | 14.9% | 5.4% | 14.6% | 20.8% | — |
| AutomationBench, private set | 30.4% | 17.0% | 10.7% | 23.6% | — |
| GDPVal-AA v2, Elo | 1525 | 1422 | 1598 | 1578 | 1628 |
| Harvey LAB-AA | 90.7% | 85.1% | 90.1% | 85.2% | — |
| GDP.pdf | 34.0% | 22.0% | 28.0% | 24.7% | 16.0% |
| CharXiv Reasoning, no tools | 84.5% | 85.2% | 77.0% | 85.9% | — |
| CharXiv Reasoning, with tools | 88.7% | 89.4% | 88.3% | — | — |
| LVBench | 85.4% | 84.2% | 68.5% | 78.9% | — |
| GDM-MRCR v2, 128K average | 97.0% | 91.8% | 81.5% | 93.5% | — |
| OSWorld 2.0 | 47.9% | 33.8% | — | 50.2% | — |
| Agent's Last Exam | 26.3% | 24.2% | 33.3% | 28.0% | — |
| HLE-Verified | 53.6% | 51.2% | 31.0% | 51.1% | — |
| BioMysteryBench, human-solvable | 87.1% | 80.6% | 87.5% | 83.8% | — |
| BioMysteryBench, human-difficult | 43.5% | 41.2% | 34.1% | 49.4% | — |
| LABBench2 | 82.1% | 76.1% | 80.1% | 81.2% | — |
These numbers are useful, but they are not one clean independent tournament. Google says Gemini scores are generally pass@1 with default sampling. Non-Gemini numbers are provider-reported unless the methodology says otherwise, and maximum available reasoning settings are preferred. Several evaluations were self-computed; harnesses, tool access, run aggregation, and even video frame limits vary by model.
Capability and performance analysis
Against Gemini 3.6 Flash, the upgrade is broad rather than cosmetic. FrontierCode rises 9.2 percentage points, DeepSWE 16.7 points, AutomationBench 13.4 points, GDP.pdf 12 points, OSWorld 14.1 points, and Code Arena gains 50 Elo. The long-context GDM-MRCR score rises from 91.8% to 97.0%.
The comparison with other model families is more nuanced:
- Gemini 3.7 Flash leads the published table on FrontierCode, Code Arena, AutomationBench, Harvey LAB-AA, GDP.pdf, LVBench, GDM-MRCR, HLE-Verified, and LABBench2.
- GPT-5.6 Terra remains ahead on DeepSWE, both Terminal-bench rows, no-tool CharXiv, OSWorld, and the human-difficult BioMysteryBench split.
- Claude Sonnet 5 leads Agent's Last Exam and narrowly leads the human-solvable BioMysteryBench split.
- Muse Spark 1.2 leads GDPVal-AA and ties GPT-5.6 Terra on the Artificial Analysis Intelligence Index, but many rows are unreported.
- Gemini 3.7 is not uniformly better than 3.6: CharXiv is 0.7 points lower both without and with tools.
The practical reading is that 3.7 Flash is unusually balanced for its launch price. It is a strong default candidate for coding agents, browser or desktop workflows, long documents, multimodal analysis, and high-volume knowledge work. GPT-5.6 Terra still deserves a direct test for long-horizon software engineering and terminal agents; Claude Sonnet 5 remains competitive for desktop-agent tasks; Muse Spark's knowledge-work result is strong but its comparison surface is incomplete.
Parameters, context, and multimodality
Google has not disclosed the total parameter count, active parameter count, expert count, architecture details, training corpus size, or compute budget. The official model card only states that Gemini 3.7 Flash is based on Gemini 3.6 Flash with algorithmic improvements. Any exact parameter estimate presented as fact would be speculation.
The parameters developers can actually depend on are the API contract:
- 1,048,576 input tokens and 65,536 output tokens;
- text, image, video, audio, and PDF input with text output;
- low, medium, and high thinking levels; minimal returns an error;
- function calling, structured output, search, Maps, code execution, file search, URL context, caching, and preview Computer Use;
- no Live API, audio generation, or image generation on this model.
A 1M context window is capacity, not a target. Long prompts still increase latency and cost, and retrieval quality does not eliminate the need to preserve authoritative state, compact repeated history, and verify tool side effects.
Call Gemini 3.7 Flash
Google's launch guide uses the Interactions API and documents medium as the default thinking level:
export GEMINI_API_KEY="<YOUR_GEMINI_API_KEY>"
curl "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-X POST \
-d '{
"model": "gemini-3.7-flash",
"input": "Review this retrying payment workflow for race conditions and propose a safe transaction boundary.",
"generation_config": {
"thinking_level": "medium"
}
}'
Modelflare currently exposes the exact model through its OpenAI-compatible Chat Completions route:
export MODELFLARE_API_KEY="<YOUR_MODELFLARE_API_KEY>"
curl "https://origin.modelflare.dev/v1/chat/completions" \
-H "Authorization: Bearer $MODELFLARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.7-flash",
"messages": [
{
"role": "user",
"content": "Review this retrying payment workflow for race conditions and propose a safe transaction boundary."
}
]
}'
The two examples use different wire formats. Do not assume every Google-native built-in tool or Interactions field is available through an OpenAI-compatible adapter.
Migration and production notes
Google's 3.7 migration guide asks developers to remove temperature, top_p, and top_k, replace thinking_budget with the string enum thinking_level, and remove unsupported candidate_count. Multi-turn Interactions workflows should use server-side previous_interaction_id.
Before replacing an existing model:
- pin the exact gemini-3.7-flash ID and reject silent alias substitution;
- replay a fixed, privacy-safe task set against the current model and 3.7;
- compare low, medium, and high thinking by successful-task cost and wall time;
- test streaming, structured output, function calling, multimodal payloads, cancellation, and refusal paths separately;
- preserve thought signatures and tool-call correlation fields required by the chosen API;
- verify cache usage, retries, timeout behavior, and duplicate-side-effect protection;
- record the final model, route, input, cached input, output, tools, latency, and charge;
- keep a rollback route until real traffic meets the acceptance threshold.
An HTTP 200 proves one request completed. It does not prove protocol parity, stable latency, correct tool behavior, or lower cost per successful task.
Gemini 3.7 Flash on Modelflare
At the production check on August 15, 2026, Modelflare's public Models & Pricing catalog listed the exact gemini-3.7-flash model ID in the gemini-stable group. The catalog reports support for the Gemini-native generateContent route and OpenAI-compatible Chat Completions. It does not list a Responses route for this model.
The displayed group multiplier was 0.15x at that check. Google's introductory price, Modelflare's configured reference price, group access, and the final recorded request charge are separate facts and can change independently. Use the live pricing page and a completed non-sensitive request as the current source of truth.
For production, create or update a key with an eligible group, use the exact model ID, and validate the protocol features your application actually uses before shifting traffic.
Bottom line
Gemini 3.7 Flash is a meaningful upgrade over 3.6 Flash at the same temporary launch price. Its strongest case is breadth: production coding, web development, tool-driven agents, complex documents, multimodal inputs, long context, and computer use all improve without moving into a premium price tier.
It is not the winner on every benchmark, its parameter count is unknown, and the launch discount expires at the end of 2026. The sensible default is to benchmark 3.7 Flash against the model already handling your real tasks, then choose by completion rate, latency, retries, and total cost—not by one composite score.
Official sources and update log
Primary sources:
- Google: Introducing Gemini 3.7 Flash
- Google AI: Gemini 3.7 Flash model page
- Google AI: latest model and migration guide
- Google AI: Gemini API pricing
- Google DeepMind: Gemini 3.7 Flash model card
- Google DeepMind: evaluation methodology
Update log:
- August 15, 2026: Initial publication. Specifications, capabilities, benchmark methodology, official prices, migration guidance, and Modelflare catalog availability checked.