Gemini 3.7 Flash Released: Specs, Benchmarks, Pricing, and Model Comparison

A fact-checked Gemini 3.7 Flash analysis covering its 1M context, 64K output, capabilities, official pricing, full comparison with 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2, plus API migration guidance.

Google released Gemini 3.7 Flash on August 13, 2026, only three weeks after Gemini 3.6 Flash. It is generally available under the stable API model ID gemini-3.7-flash and is positioned as Google's most capable Flash model for coding, agents, multimodal reasoning, and reliable multi-step execution.

The headline is not a larger context window or a disclosed parameter count. Gemini 3.7 Flash keeps the 1M-input and 64K-output class of 3.6 Flash, while Google reports substantial gains in coding, tool use, knowledge work, document understanding, and computer use. Its introductory token price is also unchanged from 3.6 Flash.

This analysis separates official specifications, Google-published benchmark results, temporary launch pricing, and current Modelflare availability. Volatile facts were checked on August 15, 2026.

Gemini 3.7 Flash at a glance

Property Official value
Release date August 13, 2026
API status Generally available; stable model ID
Model ID gemini-3.7-flash
Foundation Based on Gemini 3.6 Flash with algorithmic improvements
Input modalities Text, image, video, audio, and PDF
Output modality Text
Maximum input 1,048,576 tokens
Maximum output 65,536 tokens
Thinking levels low, medium, high; medium is the documented default
Knowledge cutoff March 2026, with some domains potentially limited to January 2025
Parameter count Not disclosed
Latest documented update August 2026

Supported capabilities include context caching, code execution, Computer Use in preview, file search, function calling, Google Maps grounding, Google Search grounding, structured outputs, thinking, and URL context. Batch, Flex, and Priority consumption modes are available. Audio generation, image generation, and the Live API are not supported by this model.

What actually changed from Gemini 3.6 Flash

Gemini 3.7 Flash is an iteration on 3.6 Flash rather than a new context or modality class. Google attributes the release to algorithmic improvements in the reasoning foundation and feedback from production developers.

The practical changes are concentrated in:

  • higher first-pass accuracy for code generation and software engineering;
  • stronger design adherence when turning screenshots, images, or design systems into web interfaces;
  • better multi-step planning, tool calls, recovery from roadblocks, and intent clarification;
  • better reasoning over complex documents in finance, law, and biosciences;
  • improved long-context retrieval and agentic computer use;
  • the same introductory Standard price as Gemini 3.6 Flash through December 31, 2026.

The last point matters. At launch, 3.7 is not a higher-priced replacement for 3.6. The migration decision is mainly about behavior, latency, and compatibility, not a token-price premium.

Official API pricing

The following paid-tier prices are in US dollars. Token prices are per 1 million tokens; cache storage is per 1 million tokens per hour.

Consumption mode Meter Through Dec 31, 2026 From Jan 1, 2027
Standard Input $0.75 $1.50
Standard Output, including thinking tokens $3.75 $7.50
Standard Cached input $0.075 $0.15
Standard Cache storage per hour $0.50 $1.00
Batch Input $0.375 $0.75
Batch Output, including thinking tokens $1.875 $3.75
Batch Cached input $0.0375 $0.075
Flex Input $0.375 $0.75
Flex Output, including thinking tokens $1.875 $3.75
Flex Cached input $0.0375 $0.075
Priority Input $1.35 $2.70
Priority Output, including thinking tokens $6.75 $13.50
Priority Cached input $0.135 $0.27

Google Search grounding includes 5,000 free search requests per month shared across Gemini 3.x models, then costs $14 per 1,000 requests. Google Maps grounding uses a shared allowance of 5,000 prompts and then the same $14 per 1,000 search queries.

At the introductory Standard rate, 100K input tokens plus 10K output tokens cost about $0.1125 before caching, tools, storage, or retries. One million cached input tokens plus 100K output tokens cost about $0.45, excluding cache storage. Both examples double after the introductory period.

Full horizontal benchmark comparison

Google's August 2026 model card compares Gemini 3.7 Flash with Gemini 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2. Higher is better for every row below. A dash means the model card did not publish a result.

Benchmark Gemini 3.7 Flash Gemini 3.6 Flash Claude Sonnet 5 GPT-5.6 Terra Muse Spark 1.2
Artificial Analysis Intelligence Index 56 52 55 57 57
FrontierCode 1.1 Main 43.6% 34.4% 42.7% 41.3%
DeepSWE v1.1 65.3% 48.6% 53.8% 69.6% 54.9%
Code Arena, WebDev Elo 1588 1538 1541 1523 1535
Terminal-bench 2.1 85.8% 78.0% 80.4% 87.4% 82.9%
Terminal-bench 3.0 14.9% 5.4% 14.6% 20.8%
AutomationBench, private set 30.4% 17.0% 10.7% 23.6%
GDPVal-AA v2, Elo 1525 1422 1598 1578 1628
Harvey LAB-AA 90.7% 85.1% 90.1% 85.2%
GDP.pdf 34.0% 22.0% 28.0% 24.7% 16.0%
CharXiv Reasoning, no tools 84.5% 85.2% 77.0% 85.9%
CharXiv Reasoning, with tools 88.7% 89.4% 88.3%
LVBench 85.4% 84.2% 68.5% 78.9%
GDM-MRCR v2, 128K average 97.0% 91.8% 81.5% 93.5%
OSWorld 2.0 47.9% 33.8% 50.2%
Agent's Last Exam 26.3% 24.2% 33.3% 28.0%
HLE-Verified 53.6% 51.2% 31.0% 51.1%
BioMysteryBench, human-solvable 87.1% 80.6% 87.5% 83.8%
BioMysteryBench, human-difficult 43.5% 41.2% 34.1% 49.4%
LABBench2 82.1% 76.1% 80.1% 81.2%

These numbers are useful, but they are not one clean independent tournament. Google says Gemini scores are generally pass@1 with default sampling. Non-Gemini numbers are provider-reported unless the methodology says otherwise, and maximum available reasoning settings are preferred. Several evaluations were self-computed; harnesses, tool access, run aggregation, and even video frame limits vary by model.

Capability and performance analysis

Against Gemini 3.6 Flash, the upgrade is broad rather than cosmetic. FrontierCode rises 9.2 percentage points, DeepSWE 16.7 points, AutomationBench 13.4 points, GDP.pdf 12 points, OSWorld 14.1 points, and Code Arena gains 50 Elo. The long-context GDM-MRCR score rises from 91.8% to 97.0%.

The comparison with other model families is more nuanced:

  • Gemini 3.7 Flash leads the published table on FrontierCode, Code Arena, AutomationBench, Harvey LAB-AA, GDP.pdf, LVBench, GDM-MRCR, HLE-Verified, and LABBench2.
  • GPT-5.6 Terra remains ahead on DeepSWE, both Terminal-bench rows, no-tool CharXiv, OSWorld, and the human-difficult BioMysteryBench split.
  • Claude Sonnet 5 leads Agent's Last Exam and narrowly leads the human-solvable BioMysteryBench split.
  • Muse Spark 1.2 leads GDPVal-AA and ties GPT-5.6 Terra on the Artificial Analysis Intelligence Index, but many rows are unreported.
  • Gemini 3.7 is not uniformly better than 3.6: CharXiv is 0.7 points lower both without and with tools.

The practical reading is that 3.7 Flash is unusually balanced for its launch price. It is a strong default candidate for coding agents, browser or desktop workflows, long documents, multimodal analysis, and high-volume knowledge work. GPT-5.6 Terra still deserves a direct test for long-horizon software engineering and terminal agents; Claude Sonnet 5 remains competitive for desktop-agent tasks; Muse Spark's knowledge-work result is strong but its comparison surface is incomplete.

Parameters, context, and multimodality

Google has not disclosed the total parameter count, active parameter count, expert count, architecture details, training corpus size, or compute budget. The official model card only states that Gemini 3.7 Flash is based on Gemini 3.6 Flash with algorithmic improvements. Any exact parameter estimate presented as fact would be speculation.

The parameters developers can actually depend on are the API contract:

  • 1,048,576 input tokens and 65,536 output tokens;
  • text, image, video, audio, and PDF input with text output;
  • low, medium, and high thinking levels; minimal returns an error;
  • function calling, structured output, search, Maps, code execution, file search, URL context, caching, and preview Computer Use;
  • no Live API, audio generation, or image generation on this model.

A 1M context window is capacity, not a target. Long prompts still increase latency and cost, and retrieval quality does not eliminate the need to preserve authoritative state, compact repeated history, and verify tool side effects.

Call Gemini 3.7 Flash

Google's launch guide uses the Interactions API and documents medium as the default thinking level:

export GEMINI_API_KEY="<YOUR_GEMINI_API_KEY>"

curl "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -X POST \
  -d '{
    "model": "gemini-3.7-flash",
    "input": "Review this retrying payment workflow for race conditions and propose a safe transaction boundary.",
    "generation_config": {
      "thinking_level": "medium"
    }
  }'

Modelflare currently exposes the exact model through its OpenAI-compatible Chat Completions route:

export MODELFLARE_API_KEY="<YOUR_MODELFLARE_API_KEY>"

curl "https://origin.modelflare.dev/v1/chat/completions" \
  -H "Authorization: Bearer $MODELFLARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.7-flash",
    "messages": [
      {
        "role": "user",
        "content": "Review this retrying payment workflow for race conditions and propose a safe transaction boundary."
      }
    ]
  }'

The two examples use different wire formats. Do not assume every Google-native built-in tool or Interactions field is available through an OpenAI-compatible adapter.

Migration and production notes

Google's 3.7 migration guide asks developers to remove temperature, top_p, and top_k, replace thinking_budget with the string enum thinking_level, and remove unsupported candidate_count. Multi-turn Interactions workflows should use server-side previous_interaction_id.

Before replacing an existing model:

  1. pin the exact gemini-3.7-flash ID and reject silent alias substitution;
  2. replay a fixed, privacy-safe task set against the current model and 3.7;
  3. compare low, medium, and high thinking by successful-task cost and wall time;
  4. test streaming, structured output, function calling, multimodal payloads, cancellation, and refusal paths separately;
  5. preserve thought signatures and tool-call correlation fields required by the chosen API;
  6. verify cache usage, retries, timeout behavior, and duplicate-side-effect protection;
  7. record the final model, route, input, cached input, output, tools, latency, and charge;
  8. keep a rollback route until real traffic meets the acceptance threshold.

An HTTP 200 proves one request completed. It does not prove protocol parity, stable latency, correct tool behavior, or lower cost per successful task.

Gemini 3.7 Flash on Modelflare

At the production check on August 15, 2026, Modelflare's public Models & Pricing catalog listed the exact gemini-3.7-flash model ID in the gemini-stable group. The catalog reports support for the Gemini-native generateContent route and OpenAI-compatible Chat Completions. It does not list a Responses route for this model.

The displayed group multiplier was 0.15x at that check. Google's introductory price, Modelflare's configured reference price, group access, and the final recorded request charge are separate facts and can change independently. Use the live pricing page and a completed non-sensitive request as the current source of truth.

For production, create or update a key with an eligible group, use the exact model ID, and validate the protocol features your application actually uses before shifting traffic.

Bottom line

Gemini 3.7 Flash is a meaningful upgrade over 3.6 Flash at the same temporary launch price. Its strongest case is breadth: production coding, web development, tool-driven agents, complex documents, multimodal inputs, long context, and computer use all improve without moving into a premium price tier.

It is not the winner on every benchmark, its parameter count is unknown, and the launch discount expires at the end of 2026. The sensible default is to benchmark 3.7 Flash against the model already handling your real tasks, then choose by completion rate, latency, retries, and total cost—not by one composite score.

Official sources and update log

Primary sources:

Update log:

  • August 15, 2026: Initial publication. Specifications, capabilities, benchmark methodology, official prices, migration guidance, and Modelflare catalog availability checked.