Skip to content

CoreWeave

(0)
HQ: United StatesDatacenters: United StatesPrivacyTermsCompare CoreWeave

Average response time

13.39 s

Estimated 500-token answer. Across 20 models with recent complete speed data.

Average response start

0.53 s

Wait until the first token. Across 20 models with recent response start data.

Average output speed

71.3 tok/s

Generated tokens per second. Across 20 models with recent output speed data.

AI Models available from CoreWeave

Prices and recent speed measurements below are for CoreWeave endpoints. A 500-token response combines the wait for the first token with the time to generate 500 output tokens. Winner and loser results compare this exact model across every provider with published comparable data.

Filter options

Filter options

Input price range

Input price range

USD per 1 million input tokens.

Models

Showing 20 models

  • Model
    Z.ai: GLM 5.2262.1K context
    Pricing
    Input
    $1.39 / 1M
    Output
    $4.4 / 1M
    Cache
    $0.26 / 1M
    Blended
    $2.3933 / 1M
    Benchmarks
    Speed
    20.46 s
    TTFT
    1.94 s
    TPS
    27.0 tok/s
    Uptime
    99.96%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $0.94 / 1M
    Output
    $4 / 1M
    Cache
    $0.19 / 1M
    Blended
    $1.96 / 1M
    Benchmarks
    Speed— Best comparable value
    7.35 s
    TTFT
    0.68 s
    TPS— Best comparable value
    75.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    IBM: Granite 4.1 8B131.1K context
    Pricing
    Input
    $0.05 / 1M
    Output
    $0.1 / 1M
    Cache
    $0.05 / 1M
    Blended
    $0.0667 / 1M
    Benchmarks
    Speed
    14.48 s
    TTFT
    0.19 s
    TPS
    35.0 tok/s
    Uptime
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Qwen: Qwen3.6 35B A3B262.1K context
    Pricing
    Input
    $0.25 / 1M
    Output
    $1.25 / 1M
    Cache
    $0.25 / 1M
    Blended
    $0.5833 / 1M
    Benchmarks
    Speed
    3.72 s
    TTFT— Best comparable value
    0.26 s
    TPS
    144.5 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Qwen: Qwen3.6 27B262.1K context
    Pricing
    Input
    $0.6 / 1M
    Output
    $3.6 / 1M
    Cache
    $0.12 / 1M
    Blended— Worst comparable value
    $1.6 / 1M
    Benchmarks
    Speed
    18.18 s
    TTFT
    0.33 s
    TPS
    28.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Pricing
    Input
    $1.74 / 1M
    Output
    $3.48 / 1M
    Cache
    $0.14 / 1M
    Blended— Worst comparable value
    $2.32 / 1M
    Benchmarks
    Speed— Worst comparable value
    73.01 s
    TTFT
    1.50 s
    TPS— Worst comparable value
    7.0 tok/s
    Uptime
    98.59%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Pricing
    Input
    $0.14 / 1M
    Output
    $0.28 / 1M
    Cache
    $0.07 / 1M
    Blended
    $0.1867 / 1M
    Benchmarks
    Speed
    17.27 s
    TTFT— Best comparable value
    0.59 s
    TPS
    30.0 tok/s
    Uptime
    99.85%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    MoonshotAI: Kimi K2.6262.1K context
    Pricing
    Input
    $0.95 / 1M
    Output
    $4 / 1M
    Cache
    $0.16 / 1M
    Blended
    $1.9667 / 1M
    Benchmarks
    Speed— Best comparable value
    3.30 s
    TTFT
    0.34 s
    TPS— Best comparable value
    169.0 tok/s
    Uptime
    99.89%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Z.ai: GLM 5.1202.8K context
    Pricing
    Input
    $1.4 / 1M
    Output
    $4.4 / 1M
    Cache
    $0.26 / 1M
    Blended
    $2.4 / 1M
    Benchmarks
    Speed— Best comparable value
    4.38 s
    TTFT
    0.50 s
    TPS— Best comparable value
    129.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Google: Gemma 4 31B262.1K context
    Pricing
    Input
    $0.12 / 1M
    Output
    $0.35 / 1M
    Cache
    $0.09 / 1M
    Blended
    $0.1967 / 1M
    Benchmarks
    Speed
    30.82 s
    TTFT
    1.41 s
    TPS
    17.0 tok/s
    Uptime
    77.67%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Qwen: Qwen3.5-35B-A3B262.1K context
    Pricing
    Input
    $0.25 / 1M
    Output
    $1.25 / 1M
    Cache
    $0.25 / 1M
    Blended
    $0.5833 / 1M
    Benchmarks
    Speed
    4.86 s
    TTFT
    0.35 s
    TPS
    111.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    MiniMax: MiniMax M2.5196.6K context
    Pricing
    Input
    $0.3 / 1M
    Output
    $1.2 / 1M
    Cache
    $0.3 / 1M
    Blended— Worst comparable value
    $0.6 / 1M
    Benchmarks
    Speed
    6.55 s
    TTFT— Best comparable value
    0.38 s
    TPS
    81.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $0.55 / 1M
    Output
    $1.65 / 1M
    Cache
    $0.55 / 1M
    Blended
    $0.9167 / 1M
    Benchmarks
    Speed
    9.97 s
    TTFT— Best comparable value
    0.35 s
    TPS
    52.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    OpenAI: gpt-oss-120b131.1K context
    Pricing
    Input
    $0.04 / 1M
    Output
    $0.14 / 1M
    Cache
    $0.04 / 1M
    Blended— Best comparable value
    $0.0733 / 1M
    Benchmarks
    Speed
    5.92 s
    TTFT
    0.36 s
    TPS
    90.0 tok/s
    Uptime
    99.68%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    OpenAI: gpt-oss-20b131.1K context
    Pricing
    Input
    $0.03 / 1M
    Output
    $0.13 / 1M
    Cache
    $0.03 / 1M
    Blended— Best comparable value
    $0.0633 / 1M
    Benchmarks
    Speed
    4.18 s
    TTFT
    0.28 s
    TPS
    128.0 tok/s
    Uptime
    87.61%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Pricing
    Input
    $0.1 / 1M
    Output
    $0.3 / 1M
    Cache
    $0.1 / 1M
    Blended
    $0.1667 / 1M
    Benchmarks
    Speed
    8.22 s
    TTFT
    0.28 s
    TPS
    63.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $1 / 1M
    Output
    $1.5 / 1M
    Cache
    $1 / 1M
    Blended
    $1.1667 / 1M
    Benchmarks
    Speed
    9.06 s
    TTFT— Best comparable value
    0.29 s
    TPS
    57.0 tok/s
    Uptime
    N/a
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $0.71 / 1M
    Output
    $0.71 / 1M
    Cache
    $0.71 / 1M
    Blended
    $0.71 / 1M
    Benchmarks
    Speed
    18.01 s
    TTFT— Best comparable value
    0.15 s
    TPS
    28.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $0.8 / 1M
    Output
    $0.8 / 1M
    Cache
    $0.8 / 1M
    Blended— Worst comparable value
    $0.8 / 1M
    Benchmarks
    Speed— Best comparable value
    12.50 s
    TTFT
    0.31 s
    TPS— Best comparable value
    41.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
  • Model
    Pricing
    Input
    $0.22 / 1M
    Output
    $0.22 / 1M
    Cache
    $0.22 / 1M
    Blended— Worst comparable value
    $0.22 / 1M
    Benchmarks
    Speed— Best comparable value
    4.63 s
    TTFT— Best comparable value
    0.17 s
    TPS— Best comparable value
    112.0 tok/s
    Uptime— Best comparable value
    100.00%
    API support
    • Tool calling
    • Tool choice
    • Structured output
    • Parallel calls
Green valueBest among providers for the exact modelRed valueWorst among providers for the exact model

Model data fetched .

Same-model comparisons calculated . How results are calculated.

Community experience

CoreWeave reviews and ratings

Ratings come from email-verified reviewers and are separate from ProviderBench performance measurements.

N/a

(0)

0 verified reviews

5 star
0
4 star
0
3 star
0
2 star
0
1 star
0
No reviews yet. Be the first to share first-hand experience with CoreWeave.

View all reviews for CoreWeave

Rate CoreWeave

Share first-hand experience with this provider’s inference service. Your email is used only to verify and manage this review.

Overall rating

30–2,000 characters. Links and HTML are not allowed.

Human verification is required before publishing.

Review guidelines

Practical comparison guide

How to evaluate CoreWeave as an AI inference provider

CoreWeave currently has 20 indexed models on this site, including 20 text. This page combines the provider-specific catalog with published pricing and recent performance data so you can identify suitable models and then compare CoreWeave with alternative hosts on each model page.

Review model coverage

CoreWeave offers model-specific routes rather than one interchangeable service. Check modality, context window, supported inputs, and parameter support before comparing price. The largest published context window in this catalog is 1M tokens.

Interpret price and speed together

The lowest published input price shown here is $0.03 / 1M. Recent 500-token response measurements are available for 20 models; missing measurements remain visible and are never estimated. CoreWeave currently leads or shares the lead on 16 exact models, including 2 with the best blended price and 5 with the highest output speed.

Confirm deployment requirements

CoreWeave lists its headquarters as US. Published datacenter coverage: US. Headquarters and hosting regions describe different things, so verify data residency, service terms, and current availability directly with the provider.