Z.ai: GLM 5.2 Performance
TrustedRouter performance signals and provider route posture for Z.ai: GLM 5.2.
z-ai/glm-5.2
Measured performance
Continuously sampled p50/p95 time-to-first-token (TTFT), time-to-first-byte (TTFB), effective throughput, and success rate for Z.ai: GLM 5.2. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.
| Provider | p50 TTFT | p95 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|---|
| friendli | 2546 ms | 7418 ms | 2546 ms | — | 100.00% | — | 33 |
| venice | 2578 ms | 3395 ms | 2578 ms | — | 100.00% | — | 18 |
| deepinfra | 2827 ms | 4557 ms | 2827 ms | — | 100.00% | — | 2 |
| parasail | 2972 ms | 3234 ms | 2972 ms | — | 100.00% | — | 9 |
| siliconflow | 3045 ms | 3911 ms | 3045 ms | — | 100.00% | — | 12 |
| together | 3057 ms | 14456 ms | 3057 ms | — | 100.00% | — | 12 |
| fireworks | 3074 ms | 3990 ms | 3074 ms | — | 100.00% | — | 14 |
| baseten | 3509 ms | 4487 ms | 3509 ms | — | 100.00% | — | 20 |
| gmi | 4074 ms | 5168 ms | 4074 ms | — | 100.00% | — | 44 |
| zai | 4400 ms | 5469 ms | 4400 ms | — | 100.00% | — | 22 |
| novita | 4681 ms | 7302 ms | 4681 ms | — | 100.00% | — | 3 |
| phala | 3738 ms | 14010 ms | 3738 ms | — | 95.69% | — | 116 |
| wafer | 3674 ms | 5123 ms | 3674 ms | — | 94.74% | — | 38 |
| crusoe | 3288 ms | 4267 ms | 3288 ms | — | 93.75% | — | 16 |
| alibaba | — | — | — | — | 0.00% | — | 9 |
| atlas-cloud | — | — | — | — | 0.00% | — | 2 |
| chutes | — | — | — | — | 0.00% | — | 23 |
| cloudflare-workers-ai | — | — | — | — | 0.00% | — | 7 |
| digitalocean | — | — | — | — | 0.00% | — | 14 |
| inceptron | — | — | — | — | 0.00% | — | 57 |
| makora | — | — | — | — | 0.00% | — | 38 |
| morph | — | — | — | — | 0.00% | — | 27 |
| telnyx | — | — | — | — | 0.00% | — | 21 |
| tinfoil | — | — | — | — | 0.00% | — | 64 |
| zero-g | — | — | — | — | 0.00% | — | 65 |
Full provider & model leaderboard.
44 routes.
More routes give the auto router more room to fail over around provider 429 and 5xx responses.
Gateway overhead is measured separately.
Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.
Metadata rollups.
Status samples store latency, outcome, provider, model, route, cost, and region metadata only.
View public status or inspect provider routes.