Every model the API serves. Each page has the model's documentation: limits, measured throughput and time-to-first-token, and how it compares with other providers of the same model.
served as Qwen/Qwen3.8-27B · fp8
text + image + video input · full-rate video to 780 s · 262,144-token context · 32k max output
$0.40 / $0.05 cached / $2.75 per M tokens (ZDR)
documentation, measured throughput and latency →
served as nvidia/Qwen3.8-Flash-Next-NVFP4 · nvfp4
text + image + video input · full-rate video to 780 s · 262,144-token context · 32k max output
$0.80 / $0.10 cached / $5.50 per M tokens (ZDR)
documentation, measured throughput and latency →
served as nvidia/Kimi-K3-NVFP4 · fp4
text + image + video input · native video at 8 fps to 780 s · 262,144-token context · 64k max output
$0.88 / $0.33 cached / $10.53 per M tokens (ZDR)
documentation, measured throughput and latency →
Availability is read live from the API's model document when this page loads.