OpenAI-compatible chat completions with native full-rate video input. If your client already speaks the OpenAI API, the only change is the base URL.
Everything on this page is behaviour you can hold us to — including the parts most providers leave unsaid, like what happens when we are full. Keys are self-serve: sign in at the console, create a key, and top up by card.
POST https://api.costplusiq.com/v1/chat/completions
Authorization: Bearer $COSTPLUSIQ_API_KEY
Content-Type: application/json
Keys are issued per caller, so one can be rotated without
disturbing anyone else. GET /models and GET /health
are public; everything else needs the bearer token.
Usage is prepaid: requests draw down the credit balance on
the account that owns the key, at the prices /models
publishes. A balance more than $0.50 overdrawn refuses further requests
with 429 insufficient_quota until credits are added in the
console — the grace
exists because usage is metered after a response, so a stream in flight
can overshoot zero without being cut off mid-answer.
A minimal request needs only model and
messages. Messages carry a role —
system, user, or assistant — and the
full history is sent on every call; append each reply to continue a
conversation.
curl https://api.costplusiq.com/v1/chat/completions \
-H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.8-27B",
"messages": [
{"role": "user", "content": "What are three differences between TCP and UDP?"}
]
}'
The same call from the OpenAI Python SDK:
from openai import OpenAI
client = OpenAI(base_url="https://api.costplusiq.com/v1",
api_key=os.environ["COSTPLUSIQ_API_KEY"])
response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[{"role": "user",
"content": "What are three differences between TCP and UDP?"}])
print(response.choices[0].message.content)
The response is the OpenAI shape, token counts included:
{
"id": "chatcmpl-9f2…",
"object": "chat.completion",
"created": 1755043200,
"model": "Qwen/Qwen3.8-27B",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "1. TCP is connection-oriented…"},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 19, "completion_tokens": 164, "total_tokens": 183}
}
Add "stream": true for server-sent events — see
streaming below. The request body is forwarded to
the engine after authentication, admission and reasoning normalization.
All three models support native function calling and
structured JSON output.
Reproducible sampling with seed is not supported on the current
fleet: the field is accepted, but repeated requests can produce different
outputs. logit_bias remains unqualified.
| parameter | range | what it does |
|---|---|---|
temperature | 0 – 2 | sampling temperature; the published recipe is 1 |
top_p | 0 – 1 | nucleus sampling; recipe 0.95 |
top_k | 1 – 100 | top-k sampling; recipe 20 |
min_p | 0 – 1 | minimum-probability cutoff |
presence_penalty | −2 – 2 | recipe 0 |
frequency_penalty | −2 – 2 | |
repetition_penalty | 0 – 2 | 1 is off |
max_tokens | 1 – 32,768 | output cap, reasoning included; max_completion_tokens is accepted as an alias |
stop | ≤ 4 strings | |
stream, stream_options.include_usage | bool | usage is always metered; include_usage only decides whether you see the frame |
chat_template_kwargs.enable_thinking | bool | false skips the reasoning phase entirely — the fastest answers |
chat_template_kwargs.reasoning_effort | low · medium · high · xhigh | Qwen's native levels are low, medium and xhigh (default). high is an alias for xhigh, not a separate budget. Top-level reasoning_effort accepts the same alias. |
reasoning_effort on nvidia/Kimi-K3-NVFP4 | none · low · high · max | Kimi K3 has exactly these four levels. The two lines above work on it too:
minimal runs as low, medium as high, xhigh as
max, and enable_thinking: false as none. Any other value is a
400. Unset, Kimi thinks at max. |
video_config | object | per-request frame sampling — see video sampling |
Unset sampling fields fall back to the checkpoint's own
generation_config. When engines disagree on a default, the
recipe on each model page is what our published numbers were measured
with.
Define functions with tools. The model returns calls in
choices[0].message.tool_calls; your application executes them.
tool_choice accepts auto, required,
none, or a named function. Streaming responses carry call fragments
in delta.tool_calls; assemble them by index.
{
"model": "Qwen/Qwen3.8-27B",
"messages": [{"role": "user", "content": "Find the red plate."}],
"tools": [{
"type": "function",
"function": {
"name": "locate_object",
"description": "Find the center pixel of an object.",
"parameters": {
"type": "object",
"properties": {"name": {"type": "string"}},
"required": ["name"],
"additionalProperties": false
}
}
}],
"tool_choice": "required",
"max_tokens": 512,
"chat_template_kwargs": {"enable_thinking": false}
}
To continue, append the complete assistant message containing
tool_calls, then one role: "tool" message per call.
Set tool_call_id to the returned call ID and content
to the tool result as a string. Send the updated messages with the tool
definitions in the next request. For a named function, use
{"type":"function","function":{"name":"locate_object"}}
as tool_choice.
Use response_format to constrain the assistant's JSON answer.
json_schema supplies a schema; json_object requests
a JSON object without a supplied schema. Include the desired fields in your
prompt. For example, add this to a request asking for an integer count:
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "count",
"strict": true,
"schema": {
"type": "object",
"properties": {"count": {"type": "integer"}},
"required": ["count"],
"additionalProperties": false
}
}
}
Use function calling for tool arguments and response_format
for the final answer in separate turns. Their combination in a single
request is not qualified. A schema constrains the output format; it does
not guarantee correct answers or prevent truncation at the token limit.
Three input methods, all through the standard content
parts array:
video_url part with an
https URL we fetch.video_url part with
the clip as a data: URL (256 MB media limit; larger
clips go by URL).image_url
parts, one per frame, in order. Frames may be hosted URLs or inline
data: URLs. Use this when you have already sampled the
frames yourself; the model reads them in the order sent.A video clip is decoded at the model's native 2 fps across its
whole duration, so prompt tokens scale with both length and
resolution — one temporal frame per second contributes
(H∕32)×(W∕32) tokens (≈333/s at a 576² grid). How to size a
clip is in the video token budget below.
curl https://api.costplusiq.com/v1/chat/completions \
-H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.8-27B",
"stream": true,
"messages": [{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}},
{"type": "text", "text": "Describe every step the operator performs."}
]}]
}'
The same call from the OpenAI Python SDK:
from openai import OpenAI
client = OpenAI(base_url="https://api.costplusiq.com/v1",
api_key=os.environ["COSTPLUSIQ_API_KEY"])
stream = client.chat.completions.create(
model="Qwen/Qwen3.8-27B", stream=True,
messages=[{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": clip_url}},
{"type": "text", "text": "Describe every step the operator performs."},
]}])
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
The defaults above — 2 fps, up to 1,560 frames, whole-clip pixel budget
— are the serving contract, and they are what every request gets when it
says nothing. A video_config object on the request changes the
sampling for that call only: no engine restart, no separate
deployment. Fewer or smaller frames mean fewer prompt tokens, a smaller
bill and a faster first token; the model was trained at 2 fps, so sparser
sampling trades temporal detail for cost.
curl https://api.costplusiq.com/v1/chat/completions \
-H "Authorization: Bearer $COSTPLUSIQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Qwen3.8-Flash-Next-NVFP4",
"video_config": {"fps": 0.5, "max_frames": 240},
"messages": [{"role": "user", "content": [
{"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}},
{"type": "text", "text": "Summarize what happens."}
]}]
}'
| key | range | meaning |
|---|---|---|
fps | 0.1 – 8 | frames sampled per second of video; default 2. Exclusive with nframes |
nframes | 1 – 1,560 | sample exactly this many frames, spread over the clip, whatever its length |
min_frames / max_frames | 1 – 1,560 | bounds applied after fps; a long clip is sampled sparser than fps once it hits max_frames |
min_pixels / max_pixels | ≤ 1,000,000,000 | per-frame pixel bounds before the whole-clip budget |
total_pixels | ≤ 1,000,000,000 | whole-clip pixel budget (frames × height × width); lower it to downscale every frame together |
resized_height / resized_width | ≤ 4,096 | force a frame size (rounded to the model's 32-pixel grid); send both |
Frames are counted in pairs (the model merges two frames
into one temporal frame), so odd values round down. A request may sample
sparser or smaller than the contract, never larger: anything above the
ranges, an unknown key, or a non-numeric value is refused with a
400 before it reaches a GPU. The 262,144-token window still
bounds the request, so raising fps on a long clip is refused
by the engine once video tokens plus your prompt exceed it. The same
object is accepted on both Qwen models and is ignored by a request that
carries no video. nvidia/Kimi-K3-NVFP4 reads video under its
own native contract (below) and takes no
video_config: sending one is answered with a 400
rather than silently ignored.
Kimi K3 reads video with its own vision stack. Send it exactly as for the
Qwen models, a video_url content part with a URL or a base64
data URL; images and videos can be mixed in any order.
400,
never cut short. Audio tracks are ignored.Video is turned into prompt tokens by the native Qwen processor, and those tokens are both what you pay for and what the context window bounds. Two things set the count:
(H∕32)×(W∕32) tokens — 324 at a 576×576 grid, so ≈333 tokens
per second of video once the per-second timestamp is added.A single whole-clip pixel budget governs downscaling:
size.longest_edge = 351,462,400 is the maximum
T·H·W (frames × height × width) across the clip — sized as
1,560 frames × 640×352, the serving contract. Below it, frames keep native
resolution; above it every frame is scaled down together, so tokens
plateau at ~160k instead of overflowing — an oversized clip is served
at reduced resolution, never refused for its pixel count.
Keep the whole request — video tokens + your prompt text + the output you want back — inside the 262,144-token context window. The measured envelope:
| knob | recommended | note |
|---|---|---|
| resolution | ≤ 640×360 | ~220 tokens/temporal-frame; higher-res input is downscaled for you past the budget |
| frame rate | 2 fps (native) | one temporal frame per second — the rate the model reads; video_config lowers it per request |
| clip length | ≤ 13 min (1,560 frames) | full-rate to 780 s; longer clips are sampled below 2 fps |
longest_edge | 351,462,400 | whole-video T·H·W budget; native resolution across the whole contract |
Long video is prefill-dominated: on a 10-minute clip the first output token can be tens of seconds away, and until then a normal SSE connection carries nothing at all. Intermediaries drop connections that look idle, so we send SSE comment frames while the prefill runs:
: keep-alive
: keep-alive
data: {"id":"chatcmpl-…","choices":[{"delta":{"content":"The operator"}}]}
: is skipped
rather than treated as a frame.524. Prefill time scales with clip length, so a non-streaming
request with a clip beyond roughly 90 seconds will time out at the
edge even while the service is healthy. This is a hard behavioural
boundary, not advice: send stream: true for video. Keep-alive
comment frames hold a streaming connection open for the whole prefill,
however long the clip.If the upstream fails after a stream has started, the error
arrives as an SSE frame — data: {"error":{"code":502,…}} followed
by data: [DONE] — because the HTTP status was already sent.
| limit | value | why |
|---|---|---|
| clip duration | ~13 min full-rate | 1,560 frames at native 2 fps; longer clips sample sparser (token budget) |
| inline media | 256 MB | a base64 body of ~341 MB; larger clips go by URL |
| request body | 384 MiB | hard ceiling; past it you get 413 |
| context | 262,144 tokens | video plateaus at ~160k, leaving ~80k for prompt + output |
| output | 32,768 tokens | per response |
| concurrency | 16 in flight | published in /models, and enforced |
We refuse rather than queue. A request that arrives with every slot busy
gets 429 with Retry-After immediately, instead of
sitting in a queue and returning slowly.
Capacity is judged by what the GPUs report running, not by what we have dispatched, so a busy box is never sold as an idle one.
| status | meaning |
|---|---|
401 | missing or unknown bearer token |
404 | that model is not served here — see /models |
413 | body past the inline-media ceiling; send a URL instead |
429 type rate_limit_error | at declared concurrency; retry after the header says |
429 type insufficient_quota | credit balance exhausted (more than $0.50 overdrawn); add credits in the console. Retrying without topping up will not succeed |
502 | upstream failure. Transport failures are already retried once on another GPU before you see this |
503 | no healthy capacity at all |
524 | a non-streaming request whose prefill exceeded the ~100 s edge ceiling — use stream: true (streaming) |
Errors use the OpenAI shape:
{"error":{"type":…,"message":…}}.
Retention is a property of your API key, chosen when the key is
created in the console and
fixed for the key's life. There is one model id — no -zdr
suffix; the key that authorizes a request determines how its data is
handled.
| tier | data handling | price |
|---|---|---|
| Standard (default) | You grant us a licence to train on your prompts | Lower input and output rates than ZDR; cached input the same on both tiers |
| ZDR | Zero data retention: prompts, media and completions are processed in memory and never persisted, never logged, and never trained on | the rates published in /models |
Same fleet, same measured performance on both tiers. What we keep for every request, on either tier, is the metadata a bill is made of: which key called, when, the model, the status, and token counts. On the ZDR tier:
Need both behaviours? Create two keys — one per tier — and route each request with the key whose guarantee it needs.
| route | auth | returns |
|---|---|---|
GET /models | public | the provider document: models, pricing, published capacity, compliance |
GET /v1/models | bearer | the OpenAI-shaped model list |
GET /health | public | readiness of the endpoint |