Models

Every model the API serves. Each page has the model's documentation: limits, measured throughput and time-to-first-token, and how it compares with other providers of the same model.

Qwen3.8-27B

listed

served as Qwen/Qwen3.8-27B · fp8

text + image + video input · full-rate video to 780 s · 262,144-token context · 32k max output

$0.40 / $0.05 cached / $2.75 per M tokens (ZDR)

documentation, measured throughput and latency →

Qwen3.8-Flash-Next-NVFP4

listed

served as nvidia/Qwen3.8-Flash-Next-NVFP4 · nvfp4

text + image + video input · full-rate video to 780 s · 262,144-token context · 32k max output

$0.80 / $0.10 cached / $5.50 per M tokens (ZDR)

documentation, measured throughput and latency →

Kimi-K3-NVFP4

listed

served as nvidia/Kimi-K3-NVFP4 · fp4

text + image + video input · native video at 8 fps to 780 s · 262,144-token context · 64k max output

$0.88 / $0.33 cached / $10.53 per M tokens (ZDR)

documentation, measured throughput and latency →

Availability is read live from the API's model document when this page loads.

© 2026 CostPlusIQ, a product of Ensoul Inc · API docs · Terms · Privacy · Refunds