This week's chart comes from Artificial Analysis, whose new Endpoint Accuracy Index measures how much of an open-weight model's quality each API provider actually delivers. The reference is a self-hosted SGLang deployment of GLM 5.2 running the lab's own serving recipe. Seven providers match it. Then the scores fall. CoreWeave lands at 90%, Scaleway at 75%, DeepInfra at 73%, and Blackbox AI at 52%.
The competition in open models has been about cost and speed. Those only matter if the quality holds, and now we can see where it does not. The cause is mundane: restrictive output-token limits cut the model off before it finishes reasoning, and on the hardest science questions the worst endpoints score half the reference. Quantisation is not the culprit. Nebius serves FP4 at 100% and Novita serves FP8 at 98%, while DeepInfra's FP4 reaches 73%, so the fault likely lies in configuration, not compression.
