Skip to content

Fastest First-Token Gemini 3.5 Flash-Lite Inference Providers

Ranks providers by how quickly the first generated token appears. Providers without enough recent data remain visible after ranked rows.

Median input price:
$0.3 / 1M
Median output price:
$2.5 / 1M
Median cache price:
$0.03 / 1M

Fastest first token endpoint ranking

Fastest First-Token Gemini 3.5 Flash-Lite Inference Providers
RankProvider / routePricingBenchmarksAPI support
#1Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.3 / 1M
Output
$2.5 / 1M
Cache
$0.03 / 1M
Blended
$1.0333 / 1M
Benchmarks
Speed— Worst comparable value
3.99 s
TTFT— Best comparable value
0.49 s
TPS— Worst comparable value
143.0 tok/s
Uptime— Worst comparable value
99.73%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#2Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.3 / 1M
Output
$2.5 / 1M
Cache
$0.03 / 1M
Blended
$1.0333 / 1M
Benchmarks
Speed— Best comparable value
3.29 s
TTFT— Worst comparable value
0.56 s
TPS— Best comparable value
183.0 tok/s
Uptime— Best comparable value
99.97%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.15 / 1M
Output
$1.25 / 1M
Cache
$0.015 / 1M
Blended— Best comparable value
$0.5167 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.73%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.54 / 1M
Output
$4.5 / 1M
Cache
$0.054 / 1M
Blended— Worst comparable value
$1.86 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Worst comparable value
99.73%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.15 / 1M
Output
$1.25 / 1M
Cache
$0.015 / 1M
Blended— Best comparable value
$0.5167 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.97%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Provider / route
1M contextunknown65.5K max output
Pricing
Input
$0.54 / 1M
Output
$4.5 / 1M
Cache
$0.054 / 1M
Blended— Worst comparable value
$1.86 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime— Best comparable value
99.97%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Green value Best comparable resultRed value Worst comparable result

Endpoint data fetched .

Gemini 3.5 Flash-Lite endpoint guide

How to interpret the Fastest First-Token Gemini 3.5 Flash-Lite Inference Providers

2 of 6 Gemini 3.5 Flash-Lite endpoints currently have the published data required for this ranking. Google AI Studio leads at 0.49 s via google-ai-studio. The table keeps unranked routes visible so missing measurements do not look like missing provider availability.

How response-start time is ranked

The first-token ranking orders endpoints by the recent median time before output begins. This is useful for interactive applications, but it should be read alongside throughput because the first provider to start is not always the first to finish a long answer.

Compare Google AI Studio with the next option

Google Vertex currently ranks second at 0.56 s. Compare that gap with input and output price, response-start time, output speed, 30-minute uptime, context length, and the exact route before deciding whether first place is meaningful for your workload.

Use the route, not only the provider name

Gemini 3.5 Flash-Lite has 6 published provider options, and performance data is matched to each exact OpenRouter routing tag. Different quantization, context, regional deployment, or provider configuration can change price and behavior even when the underlying model name is identical.