262,144-token context · 64k max output · text + image + video input ·
zero data retention · OpenAI-compatible API ·
pricing · served as nvidia/Kimi-K3-NVFP4
Video is read by the model's own vision stack, not converted to images: frames are sampled at 8 fps over the whole clip, first frame to last, and every 4 frames become one temporal chunk the vision tower encodes together, each stamped with its time in the video. Nothing is dropped from the timeline. The whole clip shares one budget of 40,960 vision tokens, so a short clip is seen at native resolution and a long one at lower resolution, still at every sampled moment. Supported up to 13 minutes. The only hard limit is the context: a clip whose tokens would not leave room to answer is refused with a 400, never cut short.
Their p50s are OpenRouter published stats on text-dominated traffic; ours is the median decode rate after the first token over streamed production requests to this endpoint in the last 7 days (see Production traffic below), whatever prompt, thinking setting and concurrency they came with; it is drawn once at least 30 of them carry a first-token time.
Time to first token: OpenRouter's published p50 for the other providers; ours the median over the same streamed production requests, at their real prompt lengths.
Serving stacks differ — quantization, kernels, speculative decoding, context handling — so the same model can score differently at different providers. OpenRouter runs its own benchmarks against every endpoint of this model; ours is the same benchmark run against this endpoint's public API, as a customer, at the model's defaults.
Ours: 93.6% (95% interval 90.3–96.5%), mean of 4 runs of 198 questions, measured 2026-09-25 — tied for 1st of 21. 17 of 20 providers score inside that interval, so the order among them is within noise. Theirs: OpenRouter's per-provider runs, fetched 2026-09-25.
| provider | score | runs |
|---|---|---|
| Parasail | 93.6% | 4 |
| CostPlusIQ | 93.6% | 4 |
| Morph Fast | 93.5% | 3 |
| Fireworks Fast | 93.3% | 4 |
| Relace | 93.1% | 3 |
| Fireworks | 93.0% | 4 |
| auto-routing | 92.9% | 3 |
| Makora | 92.8% | 3 |
| Morph | 92.6% | 4 |
| Wafer | 92.3% | 6 |
| Alibaba Cloud Int. | 92.2% | 4 |
| Baseten | 92.1% | 4 |
| Modal | 92.1% | 4 |
| Chutes | 92.0% | 4 |
| Phala | 91.9% | 4 |
| Sail Research | 91.9% | 5 |
| Moonshot AI | 91.6% | 4 |
| Together | 90.3% | 4 |
| inference.net | 90.0% | 4 |
| DeepInfra | 89.6% | 4 |
| DigitalOcean | 88.4% | 5 |
Independent model-level scores as OpenRouter publishes them, at the named effort; measured on Artificial Analysis's own harness, so they are the model's reference, not a measurement of any provider.
| benchmark | score |
|---|---|
| Kimi K3 (max) Intelligence Index | 43.6 |
| Kimi K3 (max) Coding Index | 76.2 |
| Kimi K3 (max) Agentic Index | 50.0 |
| Kimi K3 (max) GPQA Diamond | 93.5% |
| Kimi K3 (max) HLE | 46.9% |
| Kimi K3 (max) AA-LCR | 88.7% |
| Kimi K3 (max) GDPval-AA | 51.2% |
| Kimi K3 (max) CritPt | 23.4% |
| Kimi K3 (max) SciCode | 59.5% |
| Kimi K3 (max) AA-Omniscience Accuracy | 47.6% |
| Kimi K3 (max) AA-Omniscience Non-Hallucination Rate | 46.8% |
| Kimi K3 (low) Coding Index | 72.0 |
| Kimi K3 (low) GPQA Diamond | 84.2% |
| Kimi K3 (low) HLE | 25.0% |
| Kimi K3 (low) AA-LCR | 79.3% |
| Kimi K3 (low) CritPt | 3.1% |
| Kimi K3 (low) SciCode | 52.7% |
| Kimi K3 (low) AA-Omniscience Accuracy | 45.8% |
| Kimi K3 (low) AA-Omniscience Non-Hallucination Rate | 22.9% |
How ours was measured: GPQA Diamond (simple-evals CSV, 198 questions, options reshuffled per run), simple-evals prompt ending in 'Answer: $LETTER', 4 independent runs against https://api.costplusiq.com at the model's defaults (thinking on, default effort, no sampling overrides), max_tokens 65,536; unanswered or failed calls count as wrong. projects/inference/experiments/2026-09/2026-09-25_gpqa-kimi-k3-costplusiq
Charts are drawn from a nightly snapshot — ours from this endpoint's production requests over a rolling 7-day window, refreshed nightly, the others from what OpenRouter published for the other providers of this model. Each carries the date it was taken.