| model | context | native video † | ZDR § $ per M — in / cached-in / out |
standard § $ per M — in / cached-in / out |
|---|---|---|---|---|
| Qwen3.8-27B fp8 | 262,144 | ✓ full-rate, 780 s | $0.40 / $0.05 / $2.75 | $0.30 / $0.05 / $2.06 |
| Qwen3.8-Flash-Next-NVFP4 nvfp4 | 262,144 | ✓ full-rate, 780 s | $0.80 / $0.10 / $5.50 | $0.60 / $0.10 / $4.12 |
| Kimi-K3-NVFP4 fp4 | 262,144 | ✓ native 8 fps, 780 s | $0.88 / $0.33 / $10.53 | $0.66 / $0.33 / $7.90 |
† Most other endpoints of this model cap how much of a video they ingest and do not tell you: measured, a ten-minute clip reaches them as 3–10% of its frames, returned as a normal 200. We tokenize up to 1,560 frames (780 s @ 2hz), so prompt tokens grow with clip duration — see the measured plot, where their curves flatten and ours does not.
§ Zero-Data-Retention (ZDR): prompts and media are processed in memory and never stored or trained on.