We are an inference endpoint for open-weight models, served at cost-plus.

API Console →  Docs → OpenAI-compatible; video, streaming, limits, and what we do when we are full

Models

modelcontextnative video support †ZDR$ per M — in / cached-in / outstatus
Qwen3.6-27B bf16 262k✓ full-rate, 685 s$0.28 / $0.07 / $2.00 serving
Qwen3.6-35B-A3B262k✓ full-rate next
Qwen3.5-9B262k✓ full-rate next

† Video is decoded and sampled at the model's native rate of 2 frames per second across the entire clip — up to 1,370 frames, i.e. clips as long as 685 s — with every sampled frame tokenized by the vision encoder (~64 tokens per second of 360p video). Prompt tokens therefore grow linearly with clip duration: a 10-minute clip is ~38,000 tokens of visual signal. Many endpoints instead cap videos at a fixed frame budget (~30 frames total, one frame every ~20 s on a long clip) regardless of duration. Context is 262,144 tokens, so even the longest clip fits comfortably. ZDR: zero data retention — prompts and media are processed in memory and never persisted or trained on; billing keeps token counts only. Each model page shows measured throughput and TTFT by clip length.

© 2026 CostPlusIQ · API documentation · served from Nebius me-west1 · api.costplusiq.com coming online with the OpenRouter listing