Skip to content

Same-model provider benchmark

Nebius vs Together: LLM provider comparison

Compare Nebius and Together on 3 exact shared text models. ProviderBench keeps speed, price, and catalog coverage separate so naturally faster model catalogs cannot distort the result.

Nebius

14 indexed models

Headquarters
Netherlands
Server regions
N/a
Model types
13 text, 1 embeddings

Together

23 indexed models

Headquarters
United States
Server regions
N/a
Model types
19 text, 2 speech, 2 transcription

At a glance

Metric winners

There is no overall score. Each winner answers one specific question using only directly comparable data.

Fastest on shared models

1.43× typical advantage

3 exact models with complete recent speed data

Lowest token cost

$2.5067 / 1M

3 exact models using a 1K-input/500-output mix

Most models available

23 models

Complete catalog coverage across all indexed modalities

Shared-model benchmark summary

MetricNebius(0)Together(0)
SpeedNebius11.76 sTogether16.85 s
TTFTNebius0.83 sTogether1.75 s
TPSNebius46.0 tok/sTogether37.0 tok/s
UptimeNebius94.36%Together94.97%
BlendedNebius$2.5067 / 1MTogether$2.78 / 1M
Green value Better comparable resultRed value Worse comparable result

Visual comparison

Price and performance charts

NebiusTogether
500-token response by shared model

Estimated seconds using recent median response-start and output-speed data. Lower is better.

Nebius has the lower typical same-model response ratio across 3 measured models.

Blended token price by shared model

USD per 1 million tokens using a 1,000-input/500-output mix. Lower is better.

Nebius has the lower typical price ratio across 3 priced shared models.

Shared text models

3 exact models · newest first

ModelNebius(0)Together(0)
Kimi K3Winner · Nebius
Nebius
Speed— Best comparable value
11.76 s
TTFT— Best comparable value
0.89 s
TPS— Best comparable value
46.0 tok/s
Uptime— Worst comparable value
94.74%
Context
8K
Route
nebius/fp4
Blended
$7 / 1M
Together
Speed— Worst comparable value
16.85 s
TTFT— Worst comparable value
3.34 s
TPS— Worst comparable value
37.0 tok/s
Uptime— Best comparable value
94.97%
Context
1M
Route
together
Blended
$7 / 1M
gpt-oss-120bWinner · Nebius
Nebius
Speed— Best comparable value
5.14 s
TTFT— Best comparable value
0.42 s
TPS— Best comparable value
106.0 tok/s
Uptime— Best comparable value
94.36%
Context
131.1K
Route
nebius/fp4
Blended
$0.3 / 1M
Together
Speed— Worst comparable value
15.27 s
TTFT— Worst comparable value
1.75 s
TPS— Worst comparable value
37.0 tok/s
Uptime— Worst comparable value
63.29%
Context
131.1K
Route
together
Blended
$0.3 / 1M
Llama 3.3 70B InstructWinner · Nebius
Nebius
Speed— Worst comparable value
39.29 s
TTFT— Worst comparable value
0.83 s
TPS— Worst comparable value
13.0 tok/s
Uptime— Worst comparable value
57.28%
Context
131.1K
Route
nebius/fp8
Blended— Best comparable value
$0.22 / 1M
Together
Speed— Best comparable value
27.04 s
TTFT— Best comparable value
0.72 s
TPS— Best comparable value
19.0 tok/s
Uptime— Best comparable value
96.04%
Context
131.1K
Route
together/fp8
Blended— Worst comparable value
$1.04 / 1M

A per-model winner combines blended price and estimated 500-token response time with equal proportional weight. Ties and rows missing either measurement receive no badge. Comparison data calculated . Values use one deterministic route per provider and model; missing measurements remain visible as N/a.

Nebius vs Together analysis

How Nebius and Together compare for AI inference

Nebius and Together share 3 indexed text models, including Kimi K3, gpt-oss-120b, Llama 3.3 70B Instruct. Nebius has the stronger typical response-time result on the directly measured set.

Same-model speed evidence

3 shared models currently have complete response-start and output-speed measurements on both providers. The speed comparison uses per-model ratios before taking the median, so naturally faster model catalogs do not improve the result.

Token pricing on one workload

3 shared models have complete input and output prices on both providers. Prices use the same 1,000-input/500-output-token mix and are normalized to one million tokens for readability.

Catalog and deployment differences

Nebius has 14 indexed models and lists no published server regions; Together has 23 models and lists no regions. Verify data residency, privacy terms, limits, and production latency directly before choosing.

Related same-model benchmarks

Compare with other providers

Explore qualified alternatives with the most shared measured models. Recommendations include comparisons for both Nebius and Together.