Skip to content

Cheapest Gemini 3.6 Flash Inference Providers

Ranks providers by estimated cost using a 1,000-input/500-output token ratio, shown per 1 million total tokens.

Median input price:
$1.5 / 1M
Median output price:
$7.5 / 1M
Median cache price:
$0.15 / 1M

Cheapest endpoint ranking

Cheapest Gemini 3.6 Flash Inference Providers
RankProvider / routePricingBenchmarksAPI support
#1Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.75 / 1M
Output
$3.75 / 1M
Cache
$0.075 / 1M
Blended— Best comparable value
$1.75 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#2Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.75 / 1M
Output
$3.75 / 1M
Cache
$0.075 / 1M
Blended— Best comparable value
$1.75 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Median blended price · $3.5 / 1M
Endpoints below this line are priced at or above the median blended price.
#3Provider / route
1M contextunknown65.5K max output
Pricing
Input
$1.5 / 1M
Output
$7.5 / 1M
Cache
$0.15 / 1M
Blended
$3.5 / 1M
Benchmarks
Speed— Worst comparable value
51.31 s
TTFT— Best comparable value
1.31 s
TPS— Worst comparable value
10.0 tok/s
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#4Provider / route
1M contextunknown65.5K max output
Pricing
Input
$1.5 / 1M
Output
$7.5 / 1M
Cache
$0.15 / 1M
Blended
$3.5 / 1M
Benchmarks
Speed— Best comparable value
5.69 s
TTFT— Worst comparable value
1.59 s
TPS— Best comparable value
122.0 tok/s
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#5Provider / route
1M contextunknown65.5K max output
Pricing
Input
$2.7 / 1M
Output
$13.5 / 1M
Cache
$0.27 / 1M
Blended— Worst comparable value
$6.3 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.87%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#6Provider / route
1M contextunknown65.5K max output
Pricing
Input
$2.7 / 1M
Output
$13.5 / 1M
Cache
$0.27 / 1M
Blended— Worst comparable value
$6.3 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.07%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Green value Best comparable resultRed value Worst comparable result

Endpoint data fetched .

Gemini 3.6 Flash endpoint guide

How to interpret the Cheapest Gemini 3.6 Flash Inference Providers

6 of 6 Gemini 3.6 Flash endpoints currently have the published data required for this ranking. Google Vertex leads at $1.75 / 1M via google-vertex/global/flex. The table keeps unranked routes visible so missing measurements do not look like missing provider availability.

How estimated token cost is ranked

The cheapest ranking uses a 1,000-input/500-output-token ratio and reports the blended result per 1 million total tokens. Because providers can price input and output differently, a workload with a different token ratio may produce a different order.

Compare Google Vertex with the next option

Google AI Studio currently ranks second at $1.75 / 1M. Compare that gap with input and output price, response-start time, output speed, 30-minute uptime, context length, and the exact route before deciding whether first place is meaningful for your workload.

Use the route, not only the provider name

Gemini 3.6 Flash has 6 published provider options, and performance data is matched to each exact OpenRouter routing tag. Different quantization, context, regional deployment, or provider configuration can change price and behavior even when the underlying model name is identical.