01
FASTEST
GLM 5.3
#1BLACKBOX AI461.0t/s
#2Inco (FAST)388.2t/s
#3Nebius (FP4)373.2t/s
#4Makora295.3t/s
#5Together AI240.6t/s
+19% AHEAD OF Inco (FAST)VIEW MODEL →
Two models, two boards, one provider at the top of both. The gap over the runner-up is the part that shows up in your own streaming latency.
Speed claims are only useful with their units, their scope, and their shelf life attached.
Both models run on the same key and the same endpoint. Point your SDK at it, stream a completion, and count the tokens yourself.