Independent measurements, not claims
Artificial Analysis verified us as the #1 Nemotron 3 Ultra provider at 454 tokens/sec, July 2026. Same methodology and same token counting for every provider, published in the open.
Same model. Same prompt.
Watch the gap open.
Identical weights and an identical prompt, streamed side by side against Nebius — the #2 provider on the leaderboard. Blackbox answers sooner and streams 47% faster, so the same completion simply lands first. The readouts are the published Artificial Analysis figures.
Every provider.
One clear #1.
Artificial Analysis continuously benchmarks every API provider that serves Nemotron 3 Ultra. Blackbox holds the top output speed, with one of the lowest prices on the board. Verify the numbers on Artificial Analysis.
Output speed · ranked
t/s · TTFA · $ / 1MSame weights as everyone else.
A very different engine.
Every provider on the leaderboard serves the same open weights. The gap comes from the serving stack.
Custom CUDA kernels
We generate tokens, we do not resell them. Direct access to the open weights lets us run the model on our own infrastructure, with kernels tuned for exactly this architecture.
FP4 on NVIDIA B300
Advanced quantization on latest-generation hardware. On the same bench, FP4 leads BF16 on every axis — with no measurable quality loss on evals.
Multi-token prediction
Nemotron ships a built-in MTP head. Our engine uses it for speculative decoding, and emits multiple tokens per forward pass instead of one.
End-to-end encrypted, zero retention
Speed without a trade-off. We provide end-to-end encryption and retain zero data — we process prompts and completions in memory and discard them after the response.
Point your SDK at it.
Nothing else changes.
The Blackbox API is OpenAI-compatible. Swap the base URL and the key, set the model to blackboxai/nvidia/nemotron-3-ultra, and every token arrives 47% sooner. Streaming, function calling, and JSON mode work unchanged.
From the blog
Artificial Analysis: BLACKBOX AI Is the #1 Fastest Nemotron 3 Ultra Provider
Artificial Analysis independently benchmarks every API provider serving NVIDIA Nemotron 3 Ultra. In the latest snapshot, BLACKBOX AI holds #1 output speed — 454.4 tokens per second, 47% ahead of the runner-up, at roughly a third of its price.
READ ARTICLE8 MIN READNemotron on BLACKBOX: Open Weights, Encrypted Inference, and 420.2 tok/s
Blackbox is the orchestration layer for coding agents — unifying the best open- and closed-source models behind one secure, cost-efficient interface. Here is why Nemotron has become a cornerstone of the platform: frontier American open weights, 20–30× cheaper than closed-source, behind end-to-end encrypted inference, and 420.2 tok/s of concurrency-1 output.
READ ARTICLEEvery token, 47% sooner.
Get a key, point your SDK at the endpoint, and measure it on your own prompts. The leaderboard is public — so is the gap.