OpenAI compatible API · Attested · Public status
Nebius Token Factory performance
Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Nebius Token Factory.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
nebius
194 samples
Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 2543 ms |
|---|---|
| p95 TTFT | 6616 ms |
| p50 TTFB | 2641 ms |
| Effective throughput | — |
| Uptime | 98.97% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 | 2220 ms | 2219 ms | — | 100.00% | — | 8 |
| google/gemma-3-27b-it | 2226 ms | 2226 ms | — | 100.00% | — | 9 |
| openbmb/MiniCPM-V-4_5 | 2235 ms | 2235 ms | — | 100.00% | — | 14 |
| nvidia/Cosmos3-Super-Reasoner | 2245 ms | 2245 ms | — | 100.00% | — | 8 |
| Qwen/Qwen3-32B | 2281 ms | 2281 ms | — | 100.00% | — | 6 |
| Qwen/Qwen2.5-VL-72B-Instruct | 2286 ms | 2286 ms | — | 100.00% | — | 8 |
| NousResearch/Hermes-4-70B | 2301 ms | 2301 ms | — | 100.00% | — | 7 |
| Qwen/Qwen3-30B-A3B-Instruct-2507 | 2324 ms | 2324 ms | — | 100.00% | — | 6 |
| nvidia/Nemotron-3-Nano-Omni | 2328 ms | 2328 ms | — | 100.00% | — | 11 |
| Qwen/Qwen3-235B-A22B-Instruct-2507 | 2425 ms | 2425 ms | — | 100.00% | — | 11 |
| NousResearch/Hermes-4-405B | 2428 ms | 2428 ms | — | 100.00% | — | 7 |
| openai/gpt-oss-120b | 2543 ms | 2543 ms | — | 100.00% | — | 8 |
| nvidia/nemotron-3-ultra-550b-a55b | 2706 ms | 2706 ms | — | 100.00% | — | 5 |
| deepseek-ai/DeepSeek-V4-Pro | 2736 ms | 2736 ms | — | 100.00% | — | 13 |
| zai-org/GLM-5.2 | 2783 ms | 2783 ms | — | 100.00% | — | 5 |
| nvidia/nemotron-3-super-120b-a12b | 2895 ms | 2895 ms | — | 100.00% | — | 8 |
| moonshotai/Kimi-K2.6 | 2900 ms | 2899 ms | — | 100.00% | — | 10 |
| moonshotai/Kimi-K2.7-Code | 2914 ms | 2914 ms | — | 100.00% | — | 6 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B | 2927 ms | 2927 ms | — | 100.00% | — | 2 |
| Qwen/Qwen3-Next-80B-A3B-Thinking | 3999 ms | 3999 ms | — | 100.00% | — | 9 |
| zai-org/GLM-5.1 | 5242 ms | 5241 ms | — | 100.00% | — | 11 |
| meta-llama/Llama-3.3-70B-Instruct | 5629 ms | 5629 ms | — | 100.00% | — | 8 |
| MiniMaxAI/MiniMax-M3 | 2781 ms | 2781 ms | — | 85.71% | — | 7 |
| moonshotai/kimi-k3 | 2827 ms | 2827 ms | — | 85.71% | — | 7 |