PLATFORM · PROVIDER BENCHMARKS

Fastest providers for
GLM 5.3 and Kimi K3

Independent benchmarking of every API provider serving GLM 5.3 and Moonshot AI Kimi K3. BLACKBOX AI ranks first on both output speed tests.

OUTPUT SPEED · TOKENS PER SECONDRANKED FASTEST FIRST
01
FASTEST

GLM 5.3

#1BLACKBOX AI461.0t/s
#2Inco (FAST)388.2t/s
#3Nebius (FP4)373.2t/s
#4Makora295.3t/s
#5Together AI240.6t/s
+19% AHEAD OF Inco (FAST)VIEW MODEL →
02
FASTEST

Kimi K3MAX

#1BLACKBOX AI292.6t/s
#2Inco (FAST)255.0t/s
#3Databricks143.0t/s
#4Baseten94.0t/s
#5Together AI42.0t/s
+15% AHEAD OF Inco (FAST)VIEW MODEL →
THE MARGIN

First on both boards

Two models, two boards, one provider at the top of both. The gap over the runner-up is the part that shows up in your own streaming latency.

461.0 t/s
GLM 5.3 · RANKED #1
+19%
AHEAD OF THE #2 PROVIDER
292.6 t/s
KIMI K3 MAX · RANKED #1
+15%
AHEAD OF THE #2 PROVIDER
READING THE BOARDS

What the
numbers mean

Speed claims are only useful with their units, their scope, and their shelf life attached.

OUTPUT SPEED
Tokens per second on streaming completions, measured per provider endpoint. Higher is better. Every value on a board uses the same unit and the same scale.
ENDPOINT TIERS
FAST, ULTRA, and MAX are the providers' own names for their speed tiers. A row is one endpoint, not a company average, so a provider can appear more than once.
SAME WEIGHTS
Every provider on a board serves the same open weights. The spread is the serving stack: kernels, quantization, hardware, and scheduling.
A SNAPSHOT
Provider speeds move as endpoints are re-tuned and capacity changes. Read a board as the state of the run, not a standing guarantee.

Measure it on your own prompts.

Both models run on the same key and the same endpoint. Point your SDK at it, stream a completion, and count the tokens yourself.