Skip to content

Google: Gemini 3.6 Flash provider comparison

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Endpoints:
6
Context:
1M
Input:
text, image, video, file, audio
Output:
text
Compare this model in a set →

At a glance

Comparison winners

Based on 1,000 input and 500 output tokens using the latest published median speed data.

Best cost/speed trade-off

Google AI Studiogoogle-ai-studio

$3.5 / 1M tokens

5.69 s estimated response

Cheapest

Google Vertexgoogle-vertex/global/flex

$1.75 / 1M tokens

$0.002625 for the sample request

Fastest

Google AI Studiogoogle-ai-studio

5.69 s estimated response

$3.5 / 1M tokens

Visual comparison

Price and performance charts

Compare published endpoint pricing and estimated response time visually.

Effective price by endpoint

USD per 1 million tokens for input and output. Lower is better.

Google Vertex has the lowest estimated cost for 1,000 input and 500 output tokens at $0.002625.

Estimated cost vs. response time

Based on 1,000 input and 500 output tokens. Cost is shown in USD per 1 million tokens. Lower and further left is better; the single best cost/speed trade-off appears at full opacity.

2 providers have complete pricing and speed data. Google AI Studio is closest to the ideal combination of lowest cost and fastest response.

Server-rendered comparison

Provider endpoints

Prompt and completion prices are effective USD per token as published by OpenRouter. Missing speed data never removes an endpoint.

Endpoint comparison for Google: Gemini 3.6 Flash
Provider / routePricing ↗BenchmarksAPI support
Provider / route
google-vertex/global
1M contextunknown65.5K max output
Pricing
Input
$1.5 / 1M
Output
$7.5 / 1M
Cache
$0.15 / 1M
Blended
$3.5 / 1M
Benchmarks
Speed— Worst comparable value
51.31 s
TTFT— Best comparable value
1.31 s
TPS— Worst comparable value
10.0 tok/s
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
10 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools

Provider / route
google-vertex/global/flex
1M contextunknown65.5K max output
Cheapest
Pricing
Input
$0.75 / 1M
Output
$3.75 / 1M
Cache
$0.075 / 1M
Blended— Best comparable value
$1.75 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
10 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools

Provider / route
google-vertex/global/priority
1M contextunknown65.5K max output
Pricing
Input
$2.7 / 1M
Output
$13.5 / 1M
Cache
$0.27 / 1M
Blended— Worst comparable value
$6.3 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
10 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, tool_choice, tools

Provider / route
google-ai-studio
1M contextunknown65.5K max output
FastestBest cost/speed trade-off
Pricing
Input
$1.5 / 1M
Output
$7.5 / 1M
Cache
$0.15 / 1M
Blended
$3.5 / 1M
Benchmarks
Speed— Best comparable value
5.69 s
TTFT— Worst comparable value
1.59 s
TPS— Best comparable value
122.0 tok/s
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
11 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p

Provider / route
google-ai-studio/flex
1M contextunknown65.5K max output
Pricing
Input
$0.75 / 1M
Output
$3.75 / 1M
Cache
$0.075 / 1M
Blended— Best comparable value
$1.75 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
11 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p

Provider / route
google-ai-studio/priority
1M contextunknown65.5K max output
Pricing
Input
$2.7 / 1M
Output
$13.5 / 1M
Cache
$0.27 / 1M
Blended— Worst comparable value
$6.3 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
11 parameters

include_reasoning, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, temperature, tool_choice, tools, top_p

Green value Best comparable resultRed value Worst comparable result

Endpoint data fetched .

Gemini 3.6 Flash deployment guide

How to choose a Gemini 3.6 Flash inference provider

Gemini 3.6 Flash is indexed here as a text model by Google, with 6 provider endpoints from Google Vertex, Google AI Studio. The comparison preserves each exact OpenRouter routing tag so pricing and performance observations can be connected to the route an application would actually request.

Match the endpoint to the workload

The model publishes a 1M-token context window. It accepts text, image, video, file, audio input and returns text output. 12 distinct supported parameters appear across the listed routes. Confirm limits on the specific endpoint rather than assuming every host exposes the same configuration.

Compare the real request economics

Google Vertex currently has the lowest estimated cost for the standard 1,000-input/500-output-token sample at $0.002625. Input-heavy and output-heavy applications can produce a different result, so review both per-million-token prices in the endpoint table.

Balance response start and generation speed

Google AI Studio currently has the shortest estimated 500-token response at 5.69 seconds. Google AI Studio is the single endpoint closest to the current ideal cost/speed combination. Recent observations can change, so validate finalists with your own prompts.