Docs
Quickstart
SOKKAN Router speaks the OpenAI Chat Completions API. If your code already uses an OpenAI SDK, you change the base URL and the key, and keep everything else.
Base URL
https://router.sokkan.ch/v1POST /v1/chat/completionsGET /v1/modelsAuthentication
Create a key in your account. Keys start with skr_ and are shown once. Send
the key as a Bearer token:
Authorization: Bearer skr_…Requests are billed to the account that owns the key. Keep keys server-side; revoke one from the account page at any time.
First request
curl https://router.sokkan.ch/v1/chat/completions \
-H "Authorization: Bearer $SOKKAN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.8-27b",
"messages": [{"role": "user", "content": "Write a haiku about glaciers."}]
}'
# pip install openai
import os
from openai import OpenAI
client = OpenAI(base_url="https://router.sokkan.ch/v1", api_key=os.environ["SOKKAN_API_KEY"])
resp = client.chat.completions.create(
model="qwen/qwen3.8-27b",
messages=[{"role": "user", "content": "Write a haiku about glaciers."}],
)
print(resp.choices[0].message.content)
// npm install openai
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://router.sokkan.ch/v1", apiKey: process.env.SOKKAN_API_KEY });
const resp = await client.chat.completions.create({
model: "qwen/qwen3.8-27b",
messages: [{ role: "user", content: "Write a haiku about glaciers." }],
});
console.log(resp.choices[0].message.content);
Model ids are listed on Models & pricing and by GET /v1/models.
Tools (function calling), temperature, max_tokens and the other usual parameters are
passed through to the model.
Streaming
Set stream: true to receive Server-Sent Events in the OpenAI format. The last chunk carries
usage, including the cost of the request.
stream = client.chat.completions.create(
model="qwen/qwen3.8-27b",
messages=[{"role": "user", "content": "Explain hedged requests."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)Tip. Use stream: true for long generations: non-streamed responses are cut by a
100 s proxy timeout. A streamed response has no such limit as long as tokens keep flowing.
Data location
Restrict where a request may run, per request, with a header or a body field. Only upstreams in the allowed region take part.
| Value | Where the request may run |
|---|---|
CH | Switzerland only |
EU | The EU / EEA or Switzerland |
any | Any region where the model runs, including the US (default) |
X-Data-Location: CH
{
"model": "moonshotai/kimi-k2.6",
"messages": [{"role": "user", "content": "Hello"}],
"provider": { "data_location": "CH" }
}
client.chat.completions.create(
model="moonshotai/kimi-k2.6",
messages=[{"role": "user", "content": "Hello"}],
extra_headers={"X-Data-Location": "CH"},
# or: extra_body={"provider": {"data_location": "CH"}}
)
If the header and the body field are both set, the header wins. A model that does not run in the requested
location returns 400 with code no_provider: nothing is sent elsewhere.
Response headers
| Header | Meaning |
|---|---|
x-sokkan-region | Region where this response was generated: CH, EU or US |
x-request-id | Request id, also the id of the completion. Quote it when you contact support. |
When a model runs in several places, the request races them: if the first one has not started streaming after a short delay, the next one starts, and the first token wins. The switch happens before the first byte, so a response always comes from a single upstream, in the region named in the headers.
Usage and cost
Every response includes the usual usage object, plus usage.cost in USD. Cached prompt
tokens appear in usage.prompt_tokens_details.cached_tokens and are billed at the cached price.
Prices are on Models & pricing.
"usage": {
"prompt_tokens": 812, "completion_tokens": 164, "total_tokens": 976,
"prompt_tokens_details": { "cached_tokens": 512 },
"cost": 0.000246
}Errors
Errors use the OpenAI shape: {"error": {"message", "type", "code"}}.
| Status | Code | What happened |
|---|---|---|
400 | secret_detected | The prompt contains credentials (API keys, tokens, private keys). Nothing was sent to the model. Remove the secret and retry. |
400 | no_provider | This model does not run in the requested data location. |
400 | Invalid request: malformed JSON, empty messages, invalid data_location or max_tokens. | |
401 | Missing, invalid or revoked API key. | |
402 | insufficient_balance | Your balance does not cover the request (or the requested max_tokens). Top up and retry. |
403 | committed_throughput | This model is offered with Committed throughput only and your account has no active commitment for it. See Committed throughput. |
404 | model_not_found | Unknown model id. See GET /v1/models. |
429 | Too many requests for this key. Wait for the number of seconds in Retry-After. | |
502 | Every upstream failed, or the stream was interrupted upstream. Safe to retry. | |
503 | no_provider | The model is listed but not live yet (“coming soon”). |
504 | No upstream produced a first token in time. Safe to retry. |
Requests rejected before reaching a model, and requests no upstream answered, are not billed. If a response is cut midway, you pay for what was processed up to that point.
Limits and timeouts
- Each key has a request rate and a concurrency limit; over the limit you get
429withRetry-After. Need more? Ask support. - Non-streamed responses: 100 s proxy timeout. Use
stream: truefor long outputs. - If
max_tokensis not set, the output is capped by what your balance can pay for.
Committed throughput
Some models are not sold per token but with a throughput commitment on our API, prepaid by the week. In
GET /v1/models they have "availability": "committed" and a
committed_throughput object: guaranteed_output_tps (output tokens per second guaranteed to
your account), max_streams (concurrent requests covered), price_usd_per_hour,
min_weeks, activation_hours (time from payment to activation) and zone
(where requests are served).
- With an active commitment, you call the model like any other, with any of your keys. Requests are not charged to
your balance (
"cost": 0); up tomax_streamsrequests run at the same time across your keys, above that you get429. - Without a commitment, the model answers
403committed_throughput. - Models with
"availability": "opportunistic"are billed per token, best effort, and are available only while their capacity is running. - To request a commitment, use the button on the models page or write to [email protected]. Terms: section 5A.
Support
Email [email protected] with the x-request-id of the request if you
have one. Legal: Terms of Service, Data Processing Agreement,
Privacy Policy.