Docs

Quickstart

SOKKAN Router speaks the OpenAI Chat Completions API. If your code already uses an OpenAI SDK, you change the base URL and the key, and keep everything else.

Base URL

Base URLhttps://router.sokkan.ch/v1
ChatPOST /v1/chat/completions
ModelsGET /v1/models

Authentication

Create a key in your account. Keys start with skr_ and are shown once. Send the key as a Bearer token:

Authorization: Bearer skr_…

Requests are billed to the account that owns the key. Keep keys server-side; revoke one from the account page at any time.

First request

curl https://router.sokkan.ch/v1/chat/completions \
  -H "Authorization: Bearer $SOKKAN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.8-27b",
    "messages": [{"role": "user", "content": "Write a haiku about glaciers."}]
  }'

Model ids are listed on Models & pricing and by GET /v1/models. Tools (function calling), temperature, max_tokens and the other usual parameters are passed through to the model.

Streaming

Set stream: true to receive Server-Sent Events in the OpenAI format. The last chunk carries usage, including the cost of the request.

stream = client.chat.completions.create(
    model="qwen/qwen3.8-27b",
    messages=[{"role": "user", "content": "Explain hedged requests."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)

Tip. Use stream: true for long generations: non-streamed responses are cut by a 100 s proxy timeout. A streamed response has no such limit as long as tokens keep flowing.

Data location

Restrict where a request may run, per request, with a header or a body field. Only upstreams in the allowed region take part.

ValueWhere the request may run
CHSwitzerland only
EUThe EU / EEA or Switzerland
anyAny region where the model runs, including the US (default)
X-Data-Location: CH

If the header and the body field are both set, the header wins. A model that does not run in the requested location returns 400 with code no_provider: nothing is sent elsewhere.

Response headers

HeaderMeaning
x-sokkan-regionRegion where this response was generated: CH, EU or US
x-request-idRequest id, also the id of the completion. Quote it when you contact support.

When a model runs in several places, the request races them: if the first one has not started streaming after a short delay, the next one starts, and the first token wins. The switch happens before the first byte, so a response always comes from a single upstream, in the region named in the headers.

Usage and cost

Every response includes the usual usage object, plus usage.cost in USD. Cached prompt tokens appear in usage.prompt_tokens_details.cached_tokens and are billed at the cached price. Prices are on Models & pricing.

"usage": {
  "prompt_tokens": 812, "completion_tokens": 164, "total_tokens": 976,
  "prompt_tokens_details": { "cached_tokens": 512 },
  "cost": 0.000246
}

Errors

Errors use the OpenAI shape: {"error": {"message", "type", "code"}}.

StatusCodeWhat happened
400secret_detectedThe prompt contains credentials (API keys, tokens, private keys). Nothing was sent to the model. Remove the secret and retry.
400no_providerThis model does not run in the requested data location.
400Invalid request: malformed JSON, empty messages, invalid data_location or max_tokens.
401Missing, invalid or revoked API key.
402insufficient_balanceYour balance does not cover the request (or the requested max_tokens). Top up and retry.
403committed_throughputThis model is offered with Committed throughput only and your account has no active commitment for it. See Committed throughput.
404model_not_foundUnknown model id. See GET /v1/models.
429Too many requests for this key. Wait for the number of seconds in Retry-After.
502Every upstream failed, or the stream was interrupted upstream. Safe to retry.
503no_providerThe model is listed but not live yet (“coming soon”).
504No upstream produced a first token in time. Safe to retry.

Requests rejected before reaching a model, and requests no upstream answered, are not billed. If a response is cut midway, you pay for what was processed up to that point.

Limits and timeouts

  • Each key has a request rate and a concurrency limit; over the limit you get 429 with Retry-After. Need more? Ask support.
  • Non-streamed responses: 100 s proxy timeout. Use stream: true for long outputs.
  • If max_tokens is not set, the output is capped by what your balance can pay for.

Committed throughput

Some models are not sold per token but with a throughput commitment on our API, prepaid by the week. In GET /v1/models they have "availability": "committed" and a committed_throughput object: guaranteed_output_tps (output tokens per second guaranteed to your account), max_streams (concurrent requests covered), price_usd_per_hour, min_weeks, activation_hours (time from payment to activation) and zone (where requests are served).

  • With an active commitment, you call the model like any other, with any of your keys. Requests are not charged to your balance ("cost": 0); up to max_streams requests run at the same time across your keys, above that you get 429.
  • Without a commitment, the model answers 403 committed_throughput.
  • Models with "availability": "opportunistic" are billed per token, best effort, and are available only while their capacity is running.
  • To request a commitment, use the button on the models page or write to [email protected]. Terms: section 5A.

Support

Email [email protected] with the x-request-id of the request if you have one. Legal: Terms of Service, Data Processing Agreement, Privacy Policy.