Z.ai says its week-long anonymous GLM-5.3-Flash preview ran entirely on domestic Chinese accelerators, at per-token cost it calls comparable to Nvidia GPUs. It named no chip vendor and released no throughput or power figures. Serving is also the easier half of the problem.
That China is embargoed and is supposed to have no access to that types of chips.
Recent months showed a huge deal of ingenuity achieve almost competitive chips. Parts of their domestic use seems to he covered, already.