← Models

Qwen3.8-Flash-Next-NVFP4

262,144-token context · 32k max output · full-rate 2 fps native video · zero data retention · OpenAI-compatible API · pricing · served as nvidia/Qwen3.8-Flash-Next-NVFP4

Reasoning effort

Set chat_template_kwargs.reasoning_effort to low, medium, or xhigh (the default). high is accepted as an alias for xhigh and uses the same reasoning budget. The top-level reasoning_effort field accepts the same alias. See the API parameters.

Throughput vs other providers of this model

Their p50s are text-dominated traffic (OpenRouter published stats); ours is measured on full-rate video with thinking enabled — the hardest case we serve. Both sides refresh nightly.

Latency

Time to first token. Theirs is OpenRouter's published p50 on text traffic; ours is measured nightly on a short clip, so it carries a video's prefill and theirs does not. Read it with the scaling plot: most of these endpoints get faster on long clips by reading less of them.

Time to first token vs clip length (measured)

inputvideo tokensTTFT
4 s clip931 tokens0.47 s
10 s clip2,299 tokens0.66 s
60 s clip13,749 tokens2.17 s
600 s clip137,909 tokens19.51 s

Video length scaling — how much of your clip is actually read

The same clip ladder sent to us and to every other OpenRouter provider of this model, re-measured weekly. The top plot is the one that matters: video tokens ingested. Ours climbs with the clip, because we tokenize every frame at 2 fps. Most others flatten — they do not refuse a long video, they accept it, silently sample a fraction of the frames and answer confidently about footage they never saw. A flattened line is drawn dashed from the length where it stops keeping up, with a dotted marker showing where full rate would have put it.

Which is why time-to-first-token below must be read together with the plot above, never on its own: a provider that drops nine frames in ten answers faster. Being slower on a long clip is the cost of having read it.

The video token budget

Video becomes prompt tokens. At 2 fps the processor merges frame pairs into one temporal frame per second, and each contributes (H∕32)×(W∕32) tokens — ~220 at the 640×360 presentation, ≈229 tok/s with the per-second timestamp. A whole-clip pixel budget, size.longest_edge = 351,462,400 (T·H·W, sized as 1,560 frames × 640×352), keeps presentation clips at native resolution across the whole contract; oversized inputs are downscaled together, so tokens plateau at ~160k rather than overflow — served at reduced resolution, never refused.

knobrecommendedwhy
resolution≤ 640×360~220 tokens/temporal-frame; larger input is downscaled past the budget
clip length≤ 13 min (1,560 frames)full-rate 2 fps to 780 s; longer clips sample sparser
longest_edge351,462,400whole-video T·H·W budget; video plateaus at ~160k tokens
context262,144 tokens~80k left for prompt + full 32k output at the video plateau

Measured, not theoretical: a maximum-size 13-minute clip serves at 179,309 prompt tokens in ~328 s cold / ~26 s warm.

No other OpenRouter provider of this model could be measured on the clip ladder when this page was built; the plot above shows this endpoint alone.

Charts are drawn from a snapshot the nightly canary measures against this endpoint's public API, paired with what OpenRouter publishes for the other providers of this model. Each carries the date it was taken. The clip ladder is re-measured weekly.