← Models

Kimi-K3-NVFP4

262,144-token context · 64k max output · text + image + video input · zero data retention · OpenAI-compatible API · pricing · served as nvidia/Kimi-K3-NVFP4

Video

Video is read by the model's own vision stack, not converted to images: frames are sampled at 8 fps over the whole clip, first frame to last, and every 4 frames become one temporal chunk the vision tower encodes together, each stamped with its time in the video. Nothing is dropped from the timeline. The whole clip shares one budget of 40,960 vision tokens, so a short clip is seen at native resolution and a long one at lower resolution, still at every sampled moment. Supported up to 13 minutes. The only hard limit is the context: a clip whose tokens would not leave room to answer is refused with a 400, never cut short.

Throughput vs other providers of this model

Their p50s are OpenRouter published stats on text-dominated traffic; ours is the median decode rate after the first token over streamed production requests to this endpoint in the last 7 days (see Production traffic below), whatever prompt, thinking setting and concurrency they came with; it is drawn once at least 30 of them carry a first-token time.

Latency

Time to first token: OpenRouter's published p50 for the other providers; ours the median over the same streamed production requests, at their real prompt lengths.

Quality: the same benchmark, every provider of this model

Serving stacks differ — quantization, kernels, speculative decoding, context handling — so the same model can score differently at different providers. OpenRouter runs its own benchmarks against every endpoint of this model; ours is the same benchmark run against this endpoint's public API, as a customer, at the model's defaults.

GPQA Diamond

Ours: 93.6% (95% interval 90.3–96.5%), mean of 4 runs of 198 questions, measured 2026-09-25 — tied for 1st of 21. 17 of 20 providers score inside that interval, so the order among them is within noise. Theirs: OpenRouter's per-provider runs, fetched 2026-09-25.

providerscoreruns
Parasail93.6%4
CostPlusIQ93.6%4
Morph Fast93.5%3
Fireworks Fast93.3%4
Relace93.1%3
Fireworks93.0%4
auto-routing92.9%3
Makora92.8%3
Morph92.6%4
Wafer92.3%6
Alibaba Cloud Int.92.2%4
Baseten92.1%4
Modal92.1%4
Chutes92.0%4
Phala91.9%4
Sail Research91.9%5
Moonshot AI91.6%4
Together90.3%4
inference.net90.0%4
DeepInfra89.6%4
DigitalOcean88.4%5

The model itself (Artificial Analysis)

Independent model-level scores as OpenRouter publishes them, at the named effort; measured on Artificial Analysis's own harness, so they are the model's reference, not a measurement of any provider.

benchmarkscore
Kimi K3 (max) Intelligence Index43.6
Kimi K3 (max) Coding Index76.2
Kimi K3 (max) Agentic Index50.0
Kimi K3 (max) GPQA Diamond93.5%
Kimi K3 (max) HLE46.9%
Kimi K3 (max) AA-LCR88.7%
Kimi K3 (max) GDPval-AA51.2%
Kimi K3 (max) CritPt23.4%
Kimi K3 (max) SciCode59.5%
Kimi K3 (max) AA-Omniscience Accuracy47.6%
Kimi K3 (max) AA-Omniscience Non-Hallucination Rate46.8%
Kimi K3 (low) Coding Index72.0
Kimi K3 (low) GPQA Diamond84.2%
Kimi K3 (low) HLE25.0%
Kimi K3 (low) AA-LCR79.3%
Kimi K3 (low) CritPt3.1%
Kimi K3 (low) SciCode52.7%
Kimi K3 (low) AA-Omniscience Accuracy45.8%
Kimi K3 (low) AA-Omniscience Non-Hallucination Rate22.9%

How ours was measured: GPQA Diamond (simple-evals CSV, 198 questions, options reshuffled per run), simple-evals prompt ending in 'Answer: $LETTER', 4 independent runs against https://api.costplusiq.com at the model's defaults (thinking on, default effort, no sampling overrides), max_tokens 65,536; unanswered or failed calls count as wrong. projects/inference/experiments/2026-09/2026-09-25_gpqa-kimi-k3-costplusiq

Charts are drawn from a nightly snapshot — ours from this endpoint's production requests over a rolling 7-day window, refreshed nightly, the others from what OpenRouter published for the other providers of this model. Each carries the date it was taken.