Dedicated B200 serving · 262,144 context · 32k max output · video up to 685 s / 1,370 frames at native 2 fps · OpenAI-compatible API
Their p50s are text-dominated traffic (OpenRouter published stats); ours is measured on full-rate video with thinking enabled — the hardest case we serve.
| input | video tokens | TTFT |
|---|---|---|
| 4 s clip | 309 tokens | 0.58 s |
| 60 s clip | 4,391 tokens | 1.03 s |
| 600 s clip | 44,311 tokens | 6.64 s |
A 10-minute clip is ~38,000 tokens here (native 2 fps sampling). Every other ZDR endpoint of this model caps it at ~1,200 tokens (~30 frames) — measurably unusable for temporal tasks. We are the only endpoint combining full-rate video with zero data retention.
Metrics regenerate from committed benchmarks; market stats snapshotted 2026-08-11T06:19:04Z.