API

OpenAI-compatible chat completions with native full-rate video input. If your client already speaks the OpenAI API, the only change is the base URL.

Everything on this page is behaviour you can hold us to — including the parts most providers leave unsaid, like what happens when we are full. Keys are self-serve: sign in at the console, create a key, and top up by card.

Base URL and authentication

POST https://api.costplusiq.com/v1/chat/completions
Authorization: Bearer $COSTPLUSIQ_API_KEY
Content-Type: application/json

Keys are issued per caller, so one can be rotated without disturbing anyone else. GET /models and GET /health are public; everything else needs the bearer token.

Usage is prepaid: requests draw down the credit balance on the account that owns the key, at the prices /models publishes. A balance more than $0.50 overdrawn refuses further requests with 429 insufficient_quota until credits are added in the console — the grace exists because usage is metered after a response, so a stream in flight can overshoot zero without being cut off mid-answer.

Chat completions

A minimal request needs only model and messages. Messages carry a role — system, user, or assistant — and the full history is sent on every call; append each reply to continue a conversation.

curl https://api.costplusiq.com/v1/chat/completions \
  -H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-27B",
    "messages": [
      {"role": "user", "content": "What are three differences between TCP and UDP?"}
    ]
  }'

The same call from the OpenAI Python SDK:

from openai import OpenAI

client = OpenAI(base_url="https://api.costplusiq.com/v1",
                api_key=os.environ["COSTPLUSIQ_API_KEY"])

response = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=[{"role": "user",
               "content": "What are three differences between TCP and UDP?"}])
print(response.choices[0].message.content)

The response is the OpenAI shape, token counts included:

{
  "id": "chatcmpl-9f2…",
  "object": "chat.completion",
  "created": 1755043200,
  "model": "Qwen/Qwen3.8-27B",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "1. TCP is connection-oriented…"},
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 19, "completion_tokens": 164, "total_tokens": 183}
}

Add "stream": true for server-sent events — see streaming below. The request body is forwarded to the engine after authentication, admission and reasoning normalization. All three models support native function calling and structured JSON output. Reproducible sampling with seed is not supported on the current fleet: the field is accepted, but repeated requests can produce different outputs. logit_bias remains unqualified.

Request parameters

parameterrangewhat it does
temperature0 – 2sampling temperature; the published recipe is 1
top_p0 – 1nucleus sampling; recipe 0.95
top_k1 – 100top-k sampling; recipe 20
min_p0 – 1minimum-probability cutoff
presence_penalty−2 – 2recipe 0
frequency_penalty−2 – 2
repetition_penalty0 – 21 is off
max_tokens1 – 32,768output cap, reasoning included; max_completion_tokens is accepted as an alias
stop≤ 4 strings
stream, stream_options.include_usageboolusage is always metered; include_usage only decides whether you see the frame
chat_template_kwargs.enable_thinkingboolfalse skips the reasoning phase entirely — the fastest answers
chat_template_kwargs.reasoning_effortlow · medium · high · xhighQwen's native levels are low, medium and xhigh (default). high is an alias for xhigh, not a separate budget. Top-level reasoning_effort accepts the same alias.
reasoning_effort on nvidia/Kimi-K3-NVFP4none · low · high · maxKimi K3 has exactly these four levels. The two lines above work on it too: minimal runs as low, medium as high, xhigh as max, and enable_thinking: false as none. Any other value is a 400. Unset, Kimi thinks at max.
video_configobjectper-request frame sampling — see video sampling

Unset sampling fields fall back to the checkpoint's own generation_config. When engines disagree on a default, the recipe on each model page is what our published numbers were measured with.

Function calling

Define functions with tools. The model returns calls in choices[0].message.tool_calls; your application executes them. tool_choice accepts auto, required, none, or a named function. Streaming responses carry call fragments in delta.tool_calls; assemble them by index.

{
  "model": "Qwen/Qwen3.8-27B",
  "messages": [{"role": "user", "content": "Find the red plate."}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "locate_object",
      "description": "Find the center pixel of an object.",
      "parameters": {
        "type": "object",
        "properties": {"name": {"type": "string"}},
        "required": ["name"],
        "additionalProperties": false
      }
    }
  }],
  "tool_choice": "required",
  "max_tokens": 512,
  "chat_template_kwargs": {"enable_thinking": false}
}

To continue, append the complete assistant message containing tool_calls, then one role: "tool" message per call. Set tool_call_id to the returned call ID and content to the tool result as a string. Send the updated messages with the tool definitions in the next request. For a named function, use {"type":"function","function":{"name":"locate_object"}} as tool_choice.

Structured output

Use response_format to constrain the assistant's JSON answer. json_schema supplies a schema; json_object requests a JSON object without a supplied schema. Include the desired fields in your prompt. For example, add this to a request asking for an integer count:

"response_format": {
  "type": "json_schema",
  "json_schema": {
    "name": "count",
    "strict": true,
    "schema": {
      "type": "object",
      "properties": {"count": {"type": "integer"}},
      "required": ["count"],
      "additionalProperties": false
    }
  }
}

Use function calling for tool arguments and response_format for the final answer in separate turns. Their combination in a single request is not qualified. A schema constrains the output format; it does not guarantee correct answers or prevent truncation at the token limit.

Video input

Three input methods, all through the standard content parts array:

A video clip is decoded at the model's native 2 fps across its whole duration, so prompt tokens scale with both length and resolution — one temporal frame per second contributes (H∕32)×(W∕32) tokens (≈333/s at a 576² grid). How to size a clip is in the video token budget below.

curl https://api.costplusiq.com/v1/chat/completions \
  -H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3.8-27B",
    "stream": true,
    "messages": [{"role": "user", "content": [
      {"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}},
      {"type": "text", "text": "Describe every step the operator performs."}
    ]}]
  }'

The same call from the OpenAI Python SDK:

from openai import OpenAI

client = OpenAI(base_url="https://api.costplusiq.com/v1",
                api_key=os.environ["COSTPLUSIQ_API_KEY"])

stream = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B", stream=True,
    messages=[{"role": "user", "content": [
        {"type": "video_url", "video_url": {"url": clip_url}},
        {"type": "text", "text": "Describe every step the operator performs."},
    ]}])
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Video sampling, per request

The defaults above — 2 fps, up to 1,560 frames, whole-clip pixel budget — are the serving contract, and they are what every request gets when it says nothing. A video_config object on the request changes the sampling for that call only: no engine restart, no separate deployment. Fewer or smaller frames mean fewer prompt tokens, a smaller bill and a faster first token; the model was trained at 2 fps, so sparser sampling trades temporal detail for cost.

curl https://api.costplusiq.com/v1/chat/completions \
  -H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/Qwen3.8-Flash-Next-NVFP4",
    "video_config": {"fps": 0.5, "max_frames": 240},
    "messages": [{"role": "user", "content": [
      {"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}},
      {"type": "text", "text": "Summarize what happens."}
    ]}]
  }'
keyrangemeaning
fps0.1 – 8frames sampled per second of video; default 2. Exclusive with nframes
nframes1 – 1,560sample exactly this many frames, spread over the clip, whatever its length
min_frames / max_frames1 – 1,560bounds applied after fps; a long clip is sampled sparser than fps once it hits max_frames
min_pixels / max_pixels≤ 1,000,000,000per-frame pixel bounds before the whole-clip budget
total_pixels≤ 1,000,000,000whole-clip pixel budget (frames × height × width); lower it to downscale every frame together
resized_height / resized_width≤ 4,096force a frame size (rounded to the model's 32-pixel grid); send both

Frames are counted in pairs (the model merges two frames into one temporal frame), so odd values round down. A request may sample sparser or smaller than the contract, never larger: anything above the ranges, an unknown key, or a non-numeric value is refused with a 400 before it reaches a GPU. The 262,144-token window still bounds the request, so raising fps on a long clip is refused by the engine once video tokens plus your prompt exceed it. The same object is accepted on both Qwen models and is ignored by a request that carries no video. nvidia/Kimi-K3-NVFP4 reads video under its own native contract (below) and takes no video_config: sending one is answered with a 400 rather than silently ignored.

Kimi K3 video

Kimi K3 reads video with its own vision stack. Send it exactly as for the Qwen models, a video_url content part with a URL or a base64 data URL; images and videos can be mixed in any order.

The video token budget

Video is turned into prompt tokens by the native Qwen processor, and those tokens are both what you pay for and what the context window bounds. Two things set the count:

A single whole-clip pixel budget governs downscaling: size.longest_edge = 351,462,400 is the maximum T·H·W (frames × height × width) across the clip — sized as 1,560 frames × 640×352, the serving contract. Below it, frames keep native resolution; above it every frame is scaled down together, so tokens plateau at ~160k instead of overflowing — an oversized clip is served at reduced resolution, never refused for its pixel count.

Keep the whole request — video tokens + your prompt text + the output you want back — inside the 262,144-token context window. The measured envelope:

knobrecommendednote
resolution≤ 640×360~220 tokens/temporal-frame; higher-res input is downscaled for you past the budget
frame rate2 fps (native)one temporal frame per second — the rate the model reads; video_config lowers it per request
clip length≤ 13 min (1,560 frames)full-rate to 780 s; longer clips are sampled below 2 fps
longest_edge351,462,400whole-video T·H·W budget; native resolution across the whole contract
Video tokens plateau at ~160k, leaving ~80k of the 262,144-token window for your prompt and the full 32k output — measured, not theoretical: a 13-minute maximum-size clip serves at 160,584 prompt tokens in ~137 s (~86 s warm), and back-to-back maximum-size requests are stable.

Streaming, and the silence before the first token

Long video is prefill-dominated: on a 10-minute clip the first output token can be tens of seconds away, and until then a normal SSE connection carries nothing at all. Intermediaries drop connections that look idle, so we send SSE comment frames while the prefill runs:

: keep-alive

: keep-alive

data: {"id":"chatcmpl-…","choices":[{"delta":{"content":"The operator"}}]}
Every conformant SSE client ignores comment lines — the OpenAI SDKs included — so this needs no handling on your side. If you wrote your own parser, make sure a line beginning with : is skipped rather than treated as a frame.
Streaming is the contract for video. A non-streaming response cannot begin until the prefill finishes, and our edge closes a silent connection after about 100 seconds with HTTP 524. Prefill time scales with clip length, so a non-streaming request with a clip beyond roughly 90 seconds will time out at the edge even while the service is healthy. This is a hard behavioural boundary, not advice: send stream: true for video. Keep-alive comment frames hold a streaming connection open for the whole prefill, however long the clip.

If the upstream fails after a stream has started, the error arrives as an SSE frame — data: {"error":{"code":502,…}} followed by data: [DONE] — because the HTTP status was already sent.

Limits

limitvaluewhy
clip duration~13 min full-rate1,560 frames at native 2 fps; longer clips sample sparser (token budget)
inline media256 MBa base64 body of ~341 MB; larger clips go by URL
request body384 MiBhard ceiling; past it you get 413
context262,144 tokensvideo plateaus at ~160k, leaving ~80k for prompt + output
output32,768 tokensper response
concurrency16 in flightpublished in /models, and enforced

What happens when we are full

We refuse rather than queue. A request that arrives with every slot busy gets 429 with Retry-After immediately, instead of sitting in a queue and returning slowly.

Capacity is judged by what the GPUs report running, not by what we have dispatched, so a busy box is never sold as an idle one.

statusmeaning
401missing or unknown bearer token
404that model is not served here — see /models
413body past the inline-media ceiling; send a URL instead
429 type rate_limit_errorat declared concurrency; retry after the header says
429 type insufficient_quotacredit balance exhausted (more than $0.50 overdrawn); add credits in the console. Retrying without topping up will not succeed
502upstream failure. Transport failures are already retried once on another GPU before you see this
503no healthy capacity at all
524a non-streaming request whose prefill exceeded the ~100 s edge ceiling — use stream: true (streaming)

Errors use the OpenAI shape: {"error":{"type":…,"message":…}}.

Key tiers, and zero data retention

Retention is a property of your API key, chosen when the key is created in the console and fixed for the key's life. There is one model id — no -zdr suffix; the key that authorizes a request determines how its data is handled.

tierdata handlingprice
Standard (default) You grant us a licence to train on your prompts Lower input and output rates than ZDR; cached input the same on both tiers
ZDR Zero data retention: prompts, media and completions are processed in memory and never persisted, never logged, and never trained on the rates published in /models

Same fleet, same measured performance on both tiers. What we keep for every request, on either tier, is the metadata a bill is made of: which key called, when, the model, the status, and token counts. On the ZDR tier:

Need both behaviours? Create two keys — one per tier — and route each request with the key whose guarantee it needs.

Discovery endpoints

routeauthreturns
GET /modelspublicthe provider document: models, pricing, published capacity, compliance
GET /v1/modelsbearerthe OpenAI-shaped model list
GET /healthpublicreadiness of the endpoint
© 2026 CostPlusIQ, a product of Ensoul Inc · overview · models, measured throughput and TTFT · Terms · Privacy · Refunds