Skip to content

API reference

Authentication, endpoints, streaming, reasoning and errors on the OpenAI surface, the native Anthropic APIs, and the per-endpoint compatibility matrix.

Authentication

Every request to /v1/* authenticates with a virtual key in the Authorization: Bearer header. Mint keys in the Console; the secret is shown once and only its hash is stored. A key is a credential for a workload. The workload owns rate limits, budget, allowed models and policy overrides; requests also remain subject to organisation policy.

A request whose model is not on the workload's non-empty allow-list is refused with a sealed 403 permission_error before anything is dispatched; the refusal itself lands in the audit chain.

Minting a production key requires a verified email address.

Workload credentials: manage API keys and their access settings.
Workload credentials: manage API keys and their access settings.English interface · illustrative demo data. Open the image for full size.
curl https://api.sluis.ai/v1/models \
  -H "Authorization: Bearer $SLUIS_KEY"

# a model outside the key's allow-list never dispatches:
# → 403 permission_error · the refusal is sealed in the audit chain

Endpoints

Sluis exposes the OpenAI-compatible surface below, plus the native Anthropic paths detailed further down. Endpoints without a first-class handler are proxied verbatim to the routed provider, streaming included, so OpenAI-compatible upstreams keep full fidelity. A browser that opens a path the API does not serve, such as the bare API host, is redirected to these docs; API clients still get a plain 404.

EndpointPurpose
POST /v1/chat/completionsChat completions, the primary surface: routing, data protection, caching, streaming.
POST /v1/completionsLegacy text completions.
POST /v1/embeddingsEmbeddings.
POST /v1/moderationsModeration classification.
POST /v1/responsesThe OpenAI Responses API.
POST /v1/systemoneNative System One typed decisions, not OpenAI chat messages.
POST /v1/ocrDocument OCR (Mistral, or fully local with model sluis/ocr) · billed per page. With the dlp_documents policy on, inline documents are anonymized before they leave.
POST /v1/documents/anonymizeDocument anonymization · PII swapped for «MERGE_TAG»s, images blurred · billed per page/image.
POST /v1/documents/anonymize/jobsAsync anonymization jobs · enqueue large documents, poll status, fetch the result via a time-limited signed URL that needs no API key.
GET /v1/modelsThe models your policy and credentials can actually reach, nothing hypothetical.
GET /v1/models/{id}One model's metadata.
POST /v1/audio/*Transcription, translation, speech · proxied to the routed provider.
POST /v1/images/*Image generation and edits · proxied. mistral/image-generation is the exception: Mistral serves image generation only through its agent connector, so the gateway translates the OpenAI request and answers with b64_json, billed per image.
POST /v1/video/generationsVideo generation · proxied.
/v1/filesFile operations · proxied on your organisation's own key (BYOK or a custom provider), never on managed keys.
POST /v1/messagesNative Anthropic Messages ingress · point any Anthropic-SDK tool at Sluis; any connected model.
POST /v1/messages/count_tokensLocal token estimate for Anthropic-SDK context management; nothing is dispatched.
GET /anthropic/v1/modelsNative Anthropic Models discovery · list and get in Anthropic's shape, from the in-process catalog. At /v1/models the anthropic-version header selects the same shape.
HEAD /api/helloWarm-up probe for Anthropic-SDK clients such as Claude Code · 200 with an empty body.
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest",
        "messages": [{ "role": "user", "content": "Say hi" }] }'
Playground with a selected model and a completed example request.
Playground with a selected model and a completed example request.English interface · illustrative demo data. Open the image for full size.

Every JSON body above may also carry the gateway extension sluis.remove, naming terms that must be pseudonymized out of the prompt before it is dispatched. It is stripped before the request leaves the gateway; see the Data protection reference.

# Document OCR, billed per page. Returns pages[] markdown + usage_info.pages_processed.
# model "sluis/ocr" runs the gateway's own OCR engine: fully local, no provider egress (inline data: URLs only).
curl https://api.sluis.ai/v1/ocr \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-ocr-latest", "document": { "type": "document_url", "document_url": "https://example.com/invoice.pdf" } }'
# Document anonymization, billed per page/image. PII becomes «MERGE_TAG»s; images are blurred.
curl https://api.sluis.ai/v1/documents/anonymize \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -F file=@contract.docx \
  -F 'options={ "entities": ["person_name", "email", "iban"], "remove": ["Project Nightingale"], "include_mapping": true }'
Playground Documents: upload a document for anonymization.
Playground Documents: upload a document for anonymization.English interface · illustrative demo data. Open the image for full size.

System One typed decisions

# The router behind sluis/auto: send state only, not questions.
curl https://api.sluis.ai/v1/systemone \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"sluis/typed/router",
       "state":{"request":"Compare these two supplier contracts clause by clause and list the risks.",
                "previous":"Here are both contracts.",
                "attachments":["contract_a.pdf (application/pdf)","contract_b.pdf (application/pdf)"]}}'
Local-pack contractBehaviour
Model idssluis/typed answers your own questions on a local open model (GLiNER2.5-Decide). sluis/typed/router is the router Sluis uses for sluis/auto, with its own calibrated questions.
sluis/typed/routerSend the text a request carries: state.request (required), and optional previous, system and attachments (a string or a list of "name (mime)" strings); a plain string is the request. It answers reasoning (noul, P(needs multi-step reasoning)), complexity (score, 0 trivial to 3 hard, with legend and probabilities per level) and task (choice over code, writing, translation, extraction, document, support, smalltalk, action). decision holds the route and demand sluis/auto would take from this text alone, or null. Files, images, JSON mode, tools, prompt size and your organisation's policy are not part of this input and can change a real sluis/auto route. answers.reasoning keeps the keys it always had; the rest is additional.
Plain sluis/typedAlways answered on Sluis infrastructure, never by an outside provider. When the local model is unavailable the call returns 503. GET /v1/models lists it only when it can answer.
TypeSafeBring your own key only; no sluis/* alias reaches it. Name the model explicitly, for example typesafe/jev-latest. The normal gate applies: it is a US provider, so your residency policy must allow US.
AvailabilityEvery answer comes from verified model weights, reported as revision. If the local model is unavailable or cannot be verified, the call fails closed with 503.
StateRequired JSON object, list (array) or string, at most 128 KiB (131072 bytes) of serialized JSON before DLP. Missing or invalid state returns 422; oversized state returns 413. The open pack reads the first 384 words.
Questionssluis/typed/router supplies its own calibrated questions: any questions field, including null or an empty object, returns 400. Plain sluis/typed requires 1 to 16 questions of type noul, choice (2 to 20 options) or score (2 to 20 levels), each with instructions; missing questions return 422, a malformed set 400, more than 64 KiB 413. The characters [ ] ( ) and backticks are removed before the model reads them, so options that then collide return 400, as does a set longer than about 1,000 tokens.
ResponseJSON with model, answers and revision (weights SHA256); sluis/typed/router adds decision. Each answer has its type, its typed value (choice, score or noul), confidence and, where the pack reports them, probabilities, legend and answer_confidence. On sluis/typed gate on answer_confidence; its probabilities are uncalibrated. No chat completion or streaming envelope.
Latencysluis/typed/router usually answers well under 200 ms; when busy it returns 503 with Retry-After: 2. Plain sluis/typed runs all questions in one CPU pass: about 2 s for a prompt-sized state, up to about 7 s for a long one, with a 10 s budget.
Data protectionTenant DLP applies locally to the state and to the questions' instructions and criteria: block refuses; tokenize or mask rewrites what the pack is shown. Per-request removal terms are honoured.
ResidencyProcessed on Sluis infrastructure with no outside provider and no upstream residency routing step.
MeteringEach successful local decision is sealed and billed per input token at €0.01 per 1M tokens; output is free. The pack reports no token count, so input is estimated at about 4 characters per token over the state and questions it was shown. Normal limits, budgets and model allow-lists still apply.
AliasesAn existing tenant alias with the same name takes precedence over the local pack.

Document input

Attach a PDF, Word, image, or plain-text file to a chat request as a "type": "file" content part carrying a base64 file_data data URL. Any routed model accepts it: providers that understand file parts natively receive them unchanged, and for every other provider the gateway converts the document before dispatch. Provider-side file_id references are not resolved; inline the bytes as file_data.

Vision-capable models get each PDF page rendered as an image_url part (48-page budget per request, shared across all attached documents); models without image input get the document's extractable text inlined as prompt text, which the data-protection scan then covers like any other text. Longer or heavier documents belong on /v1/ocr. With the dlp_documents policy on, the document is anonymized or refused before anything leaves the gateway. When it is off and the organisation's DLP mode is protective (tokenize, mask or block), documents that would travel as uninspectable images are refused; under allow_log they pass with an explicit unscanned marker in the audit trail.

# Attach a document as an OpenAI-style `file` content part. The gateway
# normalizes it for the routed provider, so this works on any model.
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "scaleway/gemma-4-26b-a4b-it",
        "messages": [{ "role": "user", "content": [
          { "type": "file",
            "file": { "filename": "drawing.pdf",
                      "file_data": "data:application/pdf;base64,'"$(base64 -w0 drawing.pdf)"'" } },
          { "type": "text", "text": "Which discipline does this drawing document?" }
        ] }] }'

Streaming

Set stream: true and the response arrives as server-sent events: each frame is a chat.completion.chunk delta and the stream ends with data: [DONE]. Streams are never buffered in the gateway: they tee through it, and the audit seal and metering happen even if the client hangs up early.

Ask for stream_options.include_usage and the final frame carries exact token usage, the same numbers the gateway meters and bills.

After 15 seconds without a frame the gateway writes an SSE comment line (: keepalive) between events, so proxies with idle timeouts keep the connection open while a model thinks; SSE clients ignore it. Streams carry Cache-Control: no-cache and X-Accel-Buffering: no. When a gateway replica restarts it keeps serving open streams for up to five minutes; a stream still open after that ends with an error event with code server_restarting, so retry the request.

stream = client.chat.completions.create(
    model="mistral/mistral-large-latest",
    messages=[{"role": "user", "content": "Write a haiku"}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Reasoning

Pass the OpenAI reasoning_effort parameter (minimal | low | medium | high) on any thinking-capable model. Sluis translates it per provider (Gemini's thinking level, Claude's adaptive thinking effort) and omits it where a model would reject it, so one parameter works across the whole catalog.

resp = client.chat.completions.create(
    model="vertex/claude-opus-4-8",
    messages=[{"role": "user", "content": "Prove it step by step…"}],
    reasoning_effort="high",  # minimal | low | medium | high
)

Prompt caching

Repeated prompt prefixes — a long system prompt, tool definitions, a growing conversation — can be cached at the provider so they cost less and answer faster. Sluis speaks the industry-standard markers: Anthropic-style cache_control blocks (with an optional 5m or 1h TTL) and OpenAI's top-level prompt_cache_key. Send them in one shape; Sluis translates or strips them per provider, so the same request works across the whole catalog.

ProviderPrompt-cache behaviour
AnthropicExplicit cache_control breakpoints on system blocks, message content and tool definitions, with a 5m or 1h TTL. Cache reads and cache writes are reported separately.
Bedrockcache_control is translated to Converse cachePoint blocks in system, messages and toolConfig — same request, nothing to rewrite on your side.
OpenAI · AzureAutomatic caching on long prefixes. Markers are stripped; an optional prompt_cache_key is forwarded verbatim so repeat calls keep cache affinity.
Gemini · VertexGemini models get implicit caching inside the provider — nothing to send, and the reported cached-token share is passed through to your usage. Claude on Vertex takes explicit cache_control, TTL included, exactly as on Anthropic.
xAI · DeepSeek · Mistral · Qwen · Moonshot · ZhipuAutomatic where the provider offers it. Explicit markers are sanitized out, so a request written for Claude is never rejected here.
Scaleway · customScaleway has no provider prompt cache: directives are sanitized out and the gateway's own opt-in response cache is the lever there. A custom endpoint gets both directives unmodified — that behaviour is the operator's contract.

Explicit prompt caching is organisation-level opt-in and off by default: a cached prefix is data at rest at the provider — DLP-filtered by Sluis first, but stored outside it for the cache TTL. Turn it on in the Console's Data protection view. Some providers (OpenAI, Gemini, xAI, DeepSeek) cache automatically inside their own infrastructure regardless; the toggle governs what Sluis explicitly forwards (cache_control, cachePoint, prompt_cache_key), not provider-internal behaviour. Without your own key for a provider, calls run on a Sluis credential that other organisations share: on OpenAI and Azure, Sluis sets a prompt_cache_key derived from your organisation in place of yours, and on every other provider the cached-token counts in your usage show 0. Billing still applies the real cache discount.

Cache reads are billed at the provider's discounted rate; Anthropic-style cache writes add the provider's write premium to the input rate. Both are visible per call in the audit drawer and totalled in Usage & budget. On a Sluis-managed key, a request that selects a surcharge Sluis cannot meter — a cache_control ttl other than 5m where the provider honours it (Anthropic, Bedrock, Claude on Vertex), a service_tier other than auto, default, flex or standard_only, or the Anthropic long-context beta (context-1m) — is refused with 400 before anything is sent; your own provider key forwards all three unchanged.

Error codes

Errors use the OpenAI error envelope; the error.type value mirrors the HTTP status, so your SDK's error handling keeps working unchanged.

CodeWhen
400 invalid_request_errorMalformed request body or parameters.
400 invalid_request_errorModel id missing its provider prefix. Every callable id is provider/model, e.g. mistral/mistral-large-latest; the body reads: model must be provider-prefixed.
400 invalid_request_errorRequest carries the removed x-sluis-dlp header. Per-request overrides were replaced by workload option overrides; configure them in Console → Workloads.
401 authentication_errorMissing or unknown API key.
402 insufficient_quotaNo active plan or budget reached. Activate a plan or raise the budget; the request never reaches a provider.
403 permission_errorThe key lacks permission, for example the model is not on its allow-list.
422 invalid_request_errorRefused before dispatch, for example data protection in block mode matched the request.
422 document_too_large_to_scanA document is too large for the path it took: a page still over the render pixel cap at the minimum scan resolution, or a file over the text-extraction byte cap. Large-format pages (A0/A1 drawings) render downscaled automatically; the dedicated code lets a client shrink or split and resubmit.
429 rate_limit_errorRate limit reached. Enforced at the gateway; the request never hits a provider.
403 permission_errorBlocked by residency policy: no allowed jurisdiction serves the request. The body includes the reason.
403 scope_violationRefused by the workload's scope guard in block mode: the call falls outside the scope written for that workload. The message carries the guard's reason, and the request never reaches a provider.
503 pdf_renderer_unavailableAn inline PDF had to be sent as page images because the routed model reads image input, but this deployment has no PDF renderer. It is refused rather than silently answered from extracted text, which would discard everything a drawing or a scan shows.
503 scope_guard_unavailableThe workload's scope guard is in block mode but could not reach a verdict, so the call is refused instead of let through. Retry with backoff.
503 auth_unavailableThe gateway could not read its key store, so it could not verify the key. The key is not known to be invalid; retry with backoff and honour Retry-After.
504 api_errorThe provider did not answer within the workload's answer timeout: 300 seconds unless the workload sets 10 to 600, shorter if the request sent x-sluis-upstream-timeout. error.sluis.cause is timeout. The gateway did not replay the call because the provider may still be running it; shorten the prompt or ask for a longer timeout before resending.
5xx api_errorUpstream provider failure after retries; the circuit breaker steers traffic around unhealthy providers. The body carries the provider's own message alongside an error.sluis object naming the cause and every attempt made, with its upstream status or transport error and how long it took.
{
  "error": {
    "message": "model must be provider-prefixed, e.g. mistral/mistral-large-latest",
    "type": "invalid_request_error",
    "param": null,
    "code": "invalid_request"
  }
}

Anthropic Messages API

Sluis speaks the Anthropic Messages API natively. Point Claude Code or any Anthropic SDK at the gateway and call POST /v1/messages: the request is translated into the internal shape, clears exactly the same Inspect → Route → Seal → Meter pipeline as every other call, and comes back translated into Anthropic's format, streaming included. /anthropic/v1/messages is the unambiguous prefixed path to the same handler.

Authentication

Send your Sluis key in Anthropic's x-api-key header or in the usual Authorization: Bearer header; both are accepted. The anthropic-version header is read but never required, and only its shape is validated. Parsing is deliberately open-world: unknown body fields, unknown content-block types, unknown anthropic-beta values and the ?beta=true query Claude Code sends are tolerated rather than refused. Beta flags travel byte-for-byte to Anthropic-family upstreams and are ignored anywhere else.

Endpoints

EndpointPurpose
POST /v1/messagesMessages, buffered or streaming · system prompts as a string or as blocks, tools and tool results (including error results and image blocks inside them), images, document blocks, history thinking and redacted_thinking blocks, and cache_control markers on every cacheable position.
POST /v1/messages/count_tokensLocal token estimate for SDK context management: nothing is dispatched, retained or metered.
GET /v1/models · GET /v1/models/{id}Native Models discovery in Anthropic's shape (data, has_more, first_id, last_id, display_name, created_at, best-effort before_id/after_id/limit cursors), answered from the in-process catalog with no database round-trip. At the root path this shape is chosen by the anthropic-version header — without it you get the OpenAI shape; under /anthropic it always applies. Ids stay Sluis-callable provider/model ids, and created_at is a fixed placeholder because the catalog carries no per-model release date.
HEAD /api/helloThe warm-up probe Claude Code sends before a session · HEAD only, answered with 200 and an empty body.

Model routing

The wire format never restricts the model: the ingress is a codec, and the model id owns the route. A bare id such as claude-sonnet-5 resolves through this ingress's default provider — anthropic unless your organisation points it elsewhere, for example vertex for EU-hosted Claude — after which the residency gate still decides whether that target may serve the request. Pin a target explicitly with a provider-prefixed id (vertex/claude-opus-4-8), or pass a sluis/* alias: an Anthropic-SDK tool then runs on any connected provider, inside your policy.

Streaming

With "stream": true the answer is the full Anthropic event grammar: message_start carrying the real input token count, content_block_start / content_block_delta / content_block_stop per block, message_delta with cumulative usage including cache read and creation tokens, then message_stop. Text arrives as text_delta, tool arguments as input_json_delta, and reasoning as thinking_delta plus signature_delta when the routed provider produced it. A ping event follows message_start and repeats after roughly 15 seconds of upstream silence, so a long thinking turn never trips a client watchdog, and a failure mid-stream is re-encoded as an error event instead of a truncated stream. stop_sequence carries the matched sequence when the answer came from an Anthropic-family provider, and is null otherwise.

# native Anthropic Messages — the model id owns the route, not the wire format
curl https://api.sluis.ai/v1/messages \
  -H "x-api-key: $SLUIS_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{ "model": "vertex/claude-opus-4-8",
        "max_tokens": 256,
        "system": "Be concise.",
        "messages": [{ "role": "user", "content": "Say hi" }] }'

# /anthropic/v1/messages is the unambiguous prefixed path for the same handler

Errors

Every failure on these paths uses Anthropic's envelope — a top-level "type": "error" with the vendor error type (invalid_request_error, authentication_error, billing_error, permission_error, not_found_error, conflict_error, request_too_large, rate_limit_error, api_error, overloaded_error) — including the refusals that never reach a model, such as an unknown key or a data-protection block. The gateway's request id is on the response as the request-id header and in the body's request_id field, and it is the same id that identifies the call in the audit trail.

# every failure on /v1/messages* is Anthropic-shaped — auth and
# data-protection refusals included, never an OpenAI envelope
{
  "type": "error",
  "error": {
    "type": "rate_limit_error",
    "message": "rate limit exceeded"
  },
  "request_id": "req_9f2e…"
}

One request field is refused rather than forwarded: mcp_servers returns 400 invalid_request_error. A model calling an external MCP server itself would sit outside the Sluis gate, so register the server in the Console and use the MCP gateway at /v1/mcp instead, where every tool call is inspected, routed, sealed and metered.

Compatibility matrix

What the native ingresses cover today, endpoint by endpoint. This table is the claim: what is not listed here as supported is on the roadmap, deliberately refused, or a different product surface — never quietly implied.

AnthropicStatusNotes
POST /v1/messages✓ nativeBuffered and streaming, tools, images, document blocks, thinking round-trip and prompt-cache markers.
POST /v1/messages/count_tokens✓ local estimateAnswered locally; nothing is dispatched.
GET /v1/models · /v1/models/{id}✓ nativeNative list and get shape from the in-process catalog; header-negotiated at the root path, unconditional under /anthropic.
HEAD /api/hello✓Claude Code's warm-up probe, HEAD only.
Message Batches✗ roadmapMessage Batches are not implemented; batch work runs as ordinary live calls today.
Files✗ roadmapThe Files API is not bridged; inline document blocks and file_data cover the common path.
mcp_servers✗ rejectedRefused with 400 instead of forwarded — register the server and call the Sluis MCP gateway at /v1/mcp.
Skills · Agents · Admin✗ out of scopeSkills, managed agents and the admin, usage and cost APIs are a different product surface, not inference compatibility.

Cross-protocol routing. Vendor-only fields survive as long as the request lands on a provider that speaks the same protocol: an Anthropic request served by an Anthropic-family provider keeps its thinking blocks, cache markers, beta flags and matched stop_sequence. Routed cross-protocol, those fields are dropped at dispatch — never rejected at ingress — and the audit row records both the wire format and the resolved provider, so any degradation is attributable. Opaque values are never rewritten: thinking signatures and redacted_thinking payloads are forwarded verbatim or not at all, which is also why a parked thinking block that the data-protection scan flags is dropped rather than masked — a rewritten block would carry an invalid signature.