Skip to content

Fastest Claude Opus 5 Inference Providers

Ranks providers by estimated time to return 500 tokens, combining the wait for the first token with output speed.

Median input price:
$5 / 1M
Median output price:
$25 / 1M
Median cache price:
$0.5 / 1M

Fastest response endpoint ranking

Fastest Claude Opus 5 Inference Providers
RankProvider / routePricingBenchmarksAPI support
#1Provider / route
1M contextunknown128K max output
Pricing
Input
$5 / 1M
Output
$25 / 1M
Cache
$0.5 / 1M
Blended— Best comparable value
$11.6667 / 1M
Benchmarks
Speed— Best comparable value
9.55 s
TTFT— Best comparable value
2.15 s
TPS— Best comparable value
67.5 tok/s
Uptime
N/a
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#2Provider / route
1M contextunknown128K max output
Pricing
Input
$5 / 1M
Output
$25 / 1M
Cache
$0.5 / 1M
Blended— Best comparable value
$11.6667 / 1M
Benchmarks
Speed
10.97 s
TTFT
2.35 s
TPS
58.0 tok/s
Uptime— Worst comparable value
99.97%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#3Provider / route
1M contextunknown128K max output
Pricing
Input
$5 / 1M
Output
$25 / 1M
Cache
$0.5 / 1M
Blended— Best comparable value
$11.6667 / 1M
Benchmarks
Speed
11.35 s
TTFT— Worst comparable value
3.77 s
TPS
66.0 tok/s
Uptime— Best comparable value
100.00%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#4Provider / route
1M contextunknown128K max output
Pricing
Input
$5 / 1M
Output
$25 / 1M
Cache
$0.5 / 1M
Blended— Best comparable value
$11.6667 / 1M
Benchmarks
Speed
11.39 s
TTFT
2.30 s
TPS
55.0 tok/s
Uptime— Best comparable value
100.00%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
#5Provider / route
1M contextunknown128K max output
Pricing
Input
$5 / 1M
Output
$25 / 1M
Cache
$0.5 / 1M
Blended— Best comparable value
$11.6667 / 1M
Benchmarks
Speed— Worst comparable value
13.41 s
TTFT
3.60 s
TPS— Worst comparable value
51.0 tok/s
Uptime
99.99%
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Provider / route
1M contextunknown128K max output
Pricing
Input
$5.5 / 1M
Output
$27.5 / 1M
Cache
$0.55 / 1M
Blended— Worst comparable value
$12.8333 / 1M
Benchmarks
Speed
N/a
TTFT
N/a
TPS
N/a
Uptime
N/a
API support
  • Tool calling
  • Tool choice
  • Structured output
  • Parallel calls
Green value Best comparable resultRed value Worst comparable result

Endpoint data fetched .

Claude Opus 5 endpoint guide

How to interpret the Fastest Claude Opus 5 Inference Providers

5 of 6 Claude Opus 5 endpoints currently have the published data required for this ranking. Azure leads at 9.55 s via azure/us-east-2. The table keeps unranked routes visible so missing measurements do not look like missing provider availability.

How complete response time is ranked

The fastest ranking estimates a 500-token answer by adding the recent median first-token wait to 500 divided by median output throughput. It represents an example complete response, not a guarantee for every prompt or request size.

Compare Azure with the next option

Amazon Bedrock currently ranks second at 10.97 s. Compare that gap with input and output price, response-start time, output speed, 30-minute uptime, context length, and the exact route before deciding whether first place is meaningful for your workload.

Use the route, not only the provider name

Claude Opus 5 has 6 published provider options, and performance data is matched to each exact OpenRouter routing tag. Different quantization, context, regional deployment, or provider configuration can change price and behavior even when the underlying model name is identical.