Z.ai says its week-long anonymous GLM-5.3-Flash preview ran entirely on domestic Chinese accelerators, at per-token cost it calls comparable to Nvidia GPUs. It named no chip vendor and released no throughput or power figures. Serving is also the easier half of the problem.
https://huggingnews.com/ai/ox-alpha-stealth-model-launches-with-100t-token-capacity-and-glm-53-fing-4a9cff7e
Assuming they are related; A capacity @ 100 Trillion tokens a day, is something of a statement ! According to this site that is 1/4 of all global current token generation, which supposedly are at 390T tokens a day.
Wait can tokens just be directly compared like that? My impression is that token cost can vary by 2 orders of magnitude, depending on model, because the actual work of computation varies by that much between models.