We are an inference endpoint for open-weight models, served at cost-plus.
API Console → Docs → OpenAI-compatible; video, streaming, limits, and what we do when we are full
| model | context | native video support † | ZDR | $ per M — in / cached-in / out | status |
|---|---|---|---|---|---|
| Qwen3.6-27B bf16 | 262k | ✓ full-rate, 685 s | ✓ | $0.28 / $0.07 / $2.00 | serving |
| Qwen3.6-35B-A3B | 262k | ✓ full-rate | ✓ | — | next |
| Qwen3.5-9B | 262k | ✓ full-rate | ✓ | — | next |
† Video is decoded and sampled at the model's native rate of 2 frames per second across the entire clip — up to 1,370 frames, i.e. clips as long as 685 s — with every sampled frame tokenized by the vision encoder (~64 tokens per second of 360p video). Prompt tokens therefore grow linearly with clip duration: a 10-minute clip is ~38,000 tokens of visual signal. Many endpoints instead cap videos at a fixed frame budget (~30 frames total, one frame every ~20 s on a long clip) regardless of duration. Context is 262,144 tokens, so even the longest clip fits comfortably. ZDR: zero data retention — prompts and media are processed in memory and never persisted or trained on; billing keeps token counts only. Each model page shows measured throughput and TTFT by clip length.
api.costplusiq.com coming online with the OpenRouter listing