API reference
Authentication, endpoints, streaming, reasoning and errors on the OpenAI surface, the native Anthropic APIs, and the per-endpoint compatibility matrix.
Authentication
Every request to /v1/* authenticates with a virtual key in the Authorization: Bearer header. Mint keys in the Console; the secret is shown once and only its hash is stored. A key is a credential for a workload. The workload owns rate limits, budget, allowed models and policy overrides; requests also remain subject to organisation policy.
A request whose model is not on the workload's non-empty allow-list is refused with a sealed 403 permission_error before anything is dispatched; the refusal itself lands in the audit chain.
Minting a production key requires a verified email address.

curl https://api.sluis.ai/v1/models \ -H "Authorization: Bearer $SLUIS_KEY" # a model outside the key's allow-list never dispatches: # → 403 permission_error · the refusal is sealed in the audit chain
{
"object": "list",
"data": [
{ "id": "mistral/mistral-large-latest", "object": "model", "owned_by": "mistral" },
{ "id": "vertex/claude-opus-4-8", "object": "model", "owned_by": "vertex" },
{ "id": "sluis/auto", "object": "model", "owned_by": "sluis" }
]
}Endpoints
Sluis exposes the OpenAI-compatible surface below, plus the native Anthropic paths detailed further down. Endpoints without a first-class handler are proxied verbatim to the routed provider, streaming included, so OpenAI-compatible upstreams keep full fidelity. A browser that opens a path the API does not serve, such as the bare API host, is redirected to these docs; API clients still get a plain 404.
| Endpoint | Purpose |
|---|---|
| POST / | Chat completions, the primary surface: routing, data protection, caching, streaming. |
| POST / | Legacy text completions. |
| POST / | Embeddings. |
| POST / | Moderation classification. |
| POST / | The OpenAI Responses API. |
| POST / | Native System One typed decisions, not OpenAI chat messages. |
| POST / | Document OCR (Mistral, or fully local with model sluis/ocr) · billed per page. With the dlp_documents policy on, inline documents are anonymized before they leave. |
| POST / | Document anonymization · PII swapped for «MERGE_TAG»s, images blurred · billed per page/image. |
| POST / | Async anonymization jobs · enqueue large documents, poll status, fetch the result via a time-limited signed URL that needs no API key. |
| GET / | The models your policy and credentials can actually reach, nothing hypothetical. |
| GET / | One model's metadata. |
| POST / | Transcription, translation, speech · proxied to the routed provider. |
| POST / | Image generation and edits · proxied. mistral/image-generation is the exception: Mistral serves image generation only through its agent connector, so the gateway translates the OpenAI request and answers with b64_json, billed per image. |
| POST / | Video generation · proxied. |
| / | File operations · proxied on your organisation's own key (BYOK or a custom provider), never on managed keys. |
| POST / | Native Anthropic Messages ingress · point any Anthropic-SDK tool at Sluis; any connected model. |
| POST / | Local token estimate for Anthropic-SDK context management; nothing is dispatched. |
| GET / | Native Anthropic Models discovery · list and get in Anthropic's shape, from the in-process catalog. At /v1/models the anthropic-version header selects the same shape. |
| HEAD / | Warm-up probe for Anthropic-SDK clients such as Claude Code · 200 with an empty body. |
curl https://api.sluis.ai/v1/chat/completions \ -H "Authorization: Bearer $SLUIS_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "mistral/mistral-large-latest", "messages": [{ "role": "user", "content": "Say hi" }] }'
{
"id": "chatcmpl-9f2e…",
"object": "chat.completion",
"model": "mistral-large-latest",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Hi! How can I help?" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 9, "completion_tokens": 8, "total_tokens": 17 }
}
Every JSON body above may also carry the gateway extension sluis.remove, naming terms that must be pseudonymized out of the prompt before it is dispatched. It is stripped before the request leaves the gateway; see the Data protection reference.
# Document OCR, billed per page. Returns pages[] markdown + usage_info.pages_processed. # model "sluis/ocr" runs the gateway's own OCR engine: fully local, no provider egress (inline data: URLs only). curl https://api.sluis.ai/v1/ocr \ -H "Authorization: Bearer $SLUIS_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "mistral/mistral-ocr-latest", "document": { "type": "document_url", "document_url": "https://example.com/invoice.pdf" } }'
{
"pages": [{
"index": 0,
"markdown": "# Invoice 2026-118\nAcme BV · Keizersgracht 1…",
"images": [],
"dimensions": { "dpi": 150, "height": 1754, "width": 1240 }
}],
"model": "mistral-ocr-latest",
"usage_info": { "pages_processed": 1, "doc_size_bytes": 48213 }
}# Document anonymization, billed per page/image. PII becomes «MERGE_TAG»s; images are blurred. curl https://api.sluis.ai/v1/documents/anonymize \ -H "Authorization: Bearer $SLUIS_KEY" \ -F file=@contract.docx \ -F 'options={ "entities": ["person_name", "email", "iban"], "remove": ["Project Nightingale"], "include_mapping": true }'
{
"filename": "contract.docx",
"content_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
"content_base64": "UEsDBBQABgAIA…",
"summary": { "pages": 4, "images": 1, "categories": ["EMAIL", "IBAN", "PERSON_NAME"], "downgraded": false },
"mapping": [
{ "token": "«PERSON_NAME_1»", "original": "Jan de Vries" },
{ "token": "«IBAN_1»", "original": "NL91ABNA0417164300" }
]
}
System One typed decisions
# The router behind sluis/auto: send state only, not questions. curl https://api.sluis.ai/v1/systemone \ -H "Authorization: Bearer $SLUIS_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"sluis/typed/router", "state":{"request":"Compare these two supplier contracts clause by clause and list the risks.", "previous":"Here are both contracts.", "attachments":["contract_a.pdf (application/pdf)","contract_b.pdf (application/pdf)"]}}'
{
"model": "sluis/typed/router",
"answers": {
"reasoning": { "type": "noul", "noul": 0.9412,
"probabilities": { "false": 0.0588, "true": 0.9412 },
"confidence": 0.9412, "answer_confidence": 0.9412 },
"complexity": { "type": "score", "score": 2.31,
"legend": { "0": "trivial", "1": "simple", "2": "moderate", "3": "hard" },
"probabilities": { "0": 0.01, "1": 0.12, "2": 0.42, "3": 0.45 },
"confidence": 0.2612, "answer_confidence": 0.45 },
"task": { "type": "choice", "choice": "document",
"probabilities": { "code": 0.01, "writing": 0.03, "translation": 0.01, "extraction": 0.05,
"document": 0.86, "support": 0.02, "smalltalk": 0.01, "action": 0.01 },
"confidence": 0.6888, "answer_confidence": 0.86 }
},
"decision": { "route": "sluis/reasoning", "demand": "strong" },
"revision": ""
} | Local-pack contract | Behaviour |
|---|---|
| Model ids | sluis/typed answers your own questions on a local open model (GLiNER2.5-Decide). sluis/typed/router is the router Sluis uses for sluis/auto, with its own calibrated questions. |
| sluis/ | Send the text a request carries: state.request (required), and optional previous, system and attachments (a string or a list of "name (mime)" strings); a plain string is the request. It answers reasoning (noul, P(needs multi-step reasoning)), complexity (score, 0 trivial to 3 hard, with legend and probabilities per level) and task (choice over code, writing, translation, extraction, document, support, smalltalk, action). decision holds the route and demand sluis/auto would take from this text alone, or null. Files, images, JSON mode, tools, prompt size and your organisation's policy are not part of this input and can change a real sluis/auto route. answers.reasoning keeps the keys it always had; the rest is additional. |
| Plain sluis/ | Always answered on Sluis infrastructure, never by an outside provider. When the local model is unavailable the call returns 503. GET /v1/models lists it only when it can answer. |
| TypeSafe | Bring your own key only; no sluis/* alias reaches it. Name the model explicitly, for example typesafe/jev-latest. The normal gate applies: it is a US provider, so your residency policy must allow US. |
| Availability | Every answer comes from verified model weights, reported as revision. If the local model is unavailable or cannot be verified, the call fails closed with 503. |
| State | Required JSON object, list (array) or string, at most 128 KiB (131072 bytes) of serialized JSON before DLP. Missing or invalid state returns 422; oversized state returns 413. The open pack reads the first 384 words. |
| Questions | sluis/typed/router supplies its own calibrated questions: any questions field, including null or an empty object, returns 400. Plain sluis/typed requires 1 to 16 questions of type noul, choice (2 to 20 options) or score (2 to 20 levels), each with instructions; missing questions return 422, a malformed set 400, more than 64 KiB 413. The characters [ ] ( ) and backticks are removed before the model reads them, so options that then collide return 400, as does a set longer than about 1,000 tokens. |
| Response | JSON with model, answers and revision (weights SHA256); sluis/typed/router adds decision. Each answer has its type, its typed value (choice, score or noul), confidence and, where the pack reports them, probabilities, legend and answer_confidence. On sluis/typed gate on answer_confidence; its probabilities are uncalibrated. No chat completion or streaming envelope. |
| Latency | sluis/typed/router usually answers well under 200 ms; when busy it returns 503 with Retry-After: 2. Plain sluis/typed runs all questions in one CPU pass: about 2 s for a prompt-sized state, up to about 7 s for a long one, with a 10 s budget. |
| Data protection | Tenant DLP applies locally to the state and to the questions' instructions and criteria: block refuses; tokenize or mask rewrites what the pack is shown. Per-request removal terms are honoured. |
| Residency | Processed on Sluis infrastructure with no outside provider and no upstream residency routing step. |
| Metering | Each successful local decision is sealed and billed per input token at €0.01 per 1M tokens; output is free. The pack reports no token count, so input is estimated at about 4 characters per token over the state and questions it was shown. Normal limits, budgets and model allow-lists still apply. |
| Aliases | An existing tenant alias with the same name takes precedence over the local pack. |
Document input
Attach a PDF, Word, image, or plain-text file to a chat request as a "type": "file" content part carrying a base64 file_data data URL. Any routed model accepts it: providers that understand file parts natively receive them unchanged, and for every other provider the gateway converts the document before dispatch. Provider-side file_id references are not resolved; inline the bytes as file_data.
Vision-capable models get each PDF page rendered as an image_url part (48-page budget per request, shared across all attached documents); models without image input get the document's extractable text inlined as prompt text, which the data-protection scan then covers like any other text. Longer or heavier documents belong on /v1/ocr. With the dlp_documents policy on, the document is anonymized or refused before anything leaves the gateway. When it is off and the organisation's DLP mode is protective (tokenize, mask or block), documents that would travel as uninspectable images are refused; under allow_log they pass with an explicit unscanned marker in the audit trail.
# Attach a document as an OpenAI-style `file` content part. The gateway # normalizes it for the routed provider, so this works on any model. curl https://api.sluis.ai/v1/chat/completions \ -H "Authorization: Bearer $SLUIS_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "scaleway/gemma-4-26b-a4b-it", "messages": [{ "role": "user", "content": [ { "type": "file", "file": { "filename": "drawing.pdf", "file_data": "data:application/pdf;base64,'"$(base64 -w0 drawing.pdf)"'" } }, { "type": "text", "text": "Which discipline does this drawing document?" } ] }] }'
import OpenAI from "openai"; import { readFileSync } from "node:fs"; const client = new OpenAI({ baseURL: "https://api.sluis.ai/v1", apiKey: process.env.SLUIS_KEY, }); const pdf = readFileSync("drawing.pdf").toString("base64"); const reply = await client.chat.completions.create({ model: "vertex/gemini-3.1-pro", messages: [{ role: "user", content: [ { type: "file", file: { filename: "drawing.pdf", file_data: `data:application/pdf;base64,${pdf}` } }, { type: "text", text: "Which discipline does this drawing document?" }, ], }], });
Streaming
Set stream: true and the response arrives as server-sent events: each frame is a chat.completion.chunk delta and the stream ends with data: [DONE]. Streams are never buffered in the gateway: they tee through it, and the audit seal and metering happen even if the client hangs up early.
Ask for stream_options.include_usage and the final frame carries exact token usage, the same numbers the gateway meters and bills.
After 15 seconds without a frame the gateway writes an SSE comment line (: keepalive) between events, so proxies with idle timeouts keep the connection open while a model thinks; SSE clients ignore it. Streams carry Cache-Control: no-cache and X-Accel-Buffering: no. When a gateway replica restarts it keeps serving open streams for up to five minutes; a stream still open after that ends with an error event with code server_restarting, so retry the request.
stream = client.chat.completions.create(
model="mistral/mistral-large-latest",
messages=[{"role": "user", "content": "Write a haiku"}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")const stream = await client.chat.completions.create({ model: "mistral/mistral-large-latest", messages: [{ role: "user", content: "Write a haiku" }], stream: true, stream_options: { include_usage: true }, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); }
# the raw event stream on the wire
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Water"}}]}
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" finds a way."}}]}
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21}}
data: [DONE]Reasoning
Pass the OpenAI reasoning_effort parameter (minimal | low | medium | high) on any thinking-capable model. Sluis translates it per provider (Gemini's thinking level, Claude's adaptive thinking effort) and omits it where a model would reject it, so one parameter works across the whole catalog.
resp = client.chat.completions.create(
model="vertex/claude-opus-4-8",
messages=[{"role": "user", "content": "Prove it step by step…"}],
reasoning_effort="high", # minimal | low | medium | high
)Prompt caching
Repeated prompt prefixes — a long system prompt, tool definitions, a growing conversation — can be cached at the provider so they cost less and answer faster. Sluis speaks the industry-standard markers: Anthropic-style cache_control blocks (with an optional 5m or 1h TTL) and OpenAI's top-level prompt_cache_key. Send them in one shape; Sluis translates or strips them per provider, so the same request works across the whole catalog.
| Provider | Prompt-cache behaviour |
|---|---|
| Anthropic | Explicit cache_control breakpoints on system blocks, message content and tool definitions, with a 5m or 1h TTL. Cache reads and cache writes are reported separately. |
| Bedrock | cache_control is translated to Converse cachePoint blocks in system, messages and toolConfig — same request, nothing to rewrite on your side. |
| OpenAI · Azure | Automatic caching on long prefixes. Markers are stripped; an optional prompt_cache_key is forwarded verbatim so repeat calls keep cache affinity. |
| Gemini · Vertex | Gemini models get implicit caching inside the provider — nothing to send, and the reported cached-token share is passed through to your usage. Claude on Vertex takes explicit cache_control, TTL included, exactly as on Anthropic. |
| xAI · DeepSeek · Mistral · Qwen · Moonshot · Zhipu | Automatic where the provider offers it. Explicit markers are sanitized out, so a request written for Claude is never rejected here. |
| Scaleway · custom | Scaleway has no provider prompt cache: directives are sanitized out and the gateway's own opt-in response cache is the lever there. A custom endpoint gets both directives unmodified — that behaviour is the operator's contract. |
Explicit prompt caching is organisation-level opt-in and off by default: a cached prefix is data at rest at the provider — DLP-filtered by Sluis first, but stored outside it for the cache TTL. Turn it on in the Console's Data protection view. Some providers (OpenAI, Gemini, xAI, DeepSeek) cache automatically inside their own infrastructure regardless; the toggle governs what Sluis explicitly forwards (cache_control, cachePoint, prompt_cache_key), not provider-internal behaviour. Without your own key for a provider, calls run on a Sluis credential that other organisations share: on OpenAI and Azure, Sluis sets a prompt_cache_key derived from your organisation in place of yours, and on every other provider the cached-token counts in your usage show 0. Billing still applies the real cache discount.
Cache reads are billed at the provider's discounted rate; Anthropic-style cache writes add the provider's write premium to the input rate. Both are visible per call in the audit drawer and totalled in Usage & budget. On a Sluis-managed key, a request that selects a surcharge Sluis cannot meter — a cache_control ttl other than 5m where the provider honours it (Anthropic, Bedrock, Claude on Vertex), a service_tier other than auto, default, flex or standard_only, or the Anthropic long-context beta (context-1m) — is refused with 400 before anything is sent; your own provider key forwards all three unchanged.
Error codes
Errors use the OpenAI error envelope; the error.type value mirrors the HTTP status, so your SDK's error handling keeps working unchanged.
| Code | When |
|---|---|
| 400 invalid_request_error | Malformed request body or parameters. |
| 400 invalid_request_error | Model id missing its provider prefix. Every callable id is provider/model, e.g. mistral/mistral-large-latest; the body reads: model must be provider-prefixed. |
| 400 invalid_request_error | Request carries the removed x-sluis-dlp header. Per-request overrides were replaced by workload option overrides; configure them in Console → Workloads. |
| 401 authentication_error | Missing or unknown API key. |
| 402 insufficient_quota | No active plan or budget reached. Activate a plan or raise the budget; the request never reaches a provider. |
| 403 permission_error | The key lacks permission, for example the model is not on its allow-list. |
| 422 invalid_request_error | Refused before dispatch, for example data protection in block mode matched the request. |
| 422 document_too_large_to_scan | A document is too large for the path it took: a page still over the render pixel cap at the minimum scan resolution, or a file over the text-extraction byte cap. Large-format pages (A0/A1 drawings) render downscaled automatically; the dedicated code lets a client shrink or split and resubmit. |
| 429 rate_limit_error | Rate limit reached. Enforced at the gateway; the request never hits a provider. |
| 403 permission_error | Blocked by residency policy: no allowed jurisdiction serves the request. The body includes the reason. |
| 403 scope_violation | Refused by the workload's scope guard in block mode: the call falls outside the scope written for that workload. The message carries the guard's reason, and the request never reaches a provider. |
| 503 pdf_renderer_unavailable | An inline PDF had to be sent as page images because the routed model reads image input, but this deployment has no PDF renderer. It is refused rather than silently answered from extracted text, which would discard everything a drawing or a scan shows. |
| 503 scope_guard_unavailable | The workload's scope guard is in block mode but could not reach a verdict, so the call is refused instead of let through. Retry with backoff. |
| 503 auth_unavailable | The gateway could not read its key store, so it could not verify the key. The key is not known to be invalid; retry with backoff and honour Retry-After. |
| 504 api_error | The provider did not answer within the workload's answer timeout: 300 seconds unless the workload sets 10 to 600, shorter if the request sent x-sluis-upstream-timeout. error.sluis.cause is timeout. The gateway did not replay the call because the provider may still be running it; shorten the prompt or ask for a longer timeout before resending. |
| 5xx api_error | Upstream provider failure after retries; the circuit breaker steers traffic around unhealthy providers. The body carries the provider's own message alongside an error.sluis object naming the cause and every attempt made, with its upstream status or transport error and how long it took. |
{
"error": {
"message": "model must be provider-prefixed, e.g. mistral/mistral-large-latest",
"type": "invalid_request_error",
"param": null,
"code": "invalid_request"
}
}{
"error": {
"message": "model is not on this key's allow-list",
"type": "permission_error",
"param": null,
"code": "permission_denied"
}
}{
"error": {
"message": "provider `openai` (jurisdiction `US`) is not permitted by the tenant's residency policy",
"type": "permission_error",
"param": null,
"code": "permission_denied"
}
}Anthropic Messages API
Sluis speaks the Anthropic Messages API natively. Point Claude Code or any Anthropic SDK at the gateway and call POST /v1/messages: the request is translated into the internal shape, clears exactly the same Inspect → Route → Seal → Meter pipeline as every other call, and comes back translated into Anthropic's format, streaming included. /anthropic/v1/messages is the unambiguous prefixed path to the same handler.
Authentication
Send your Sluis key in Anthropic's x-api-key header or in the usual Authorization: Bearer header; both are accepted. The anthropic-version header is read but never required, and only its shape is validated. Parsing is deliberately open-world: unknown body fields, unknown content-block types, unknown anthropic-beta values and the ?beta=true query Claude Code sends are tolerated rather than refused. Beta flags travel byte-for-byte to Anthropic-family upstreams and are ignored anywhere else.
Endpoints
| Endpoint | Purpose |
|---|---|
| POST / | Messages, buffered or streaming · system prompts as a string or as blocks, tools and tool results (including error results and image blocks inside them), images, document blocks, history thinking and redacted_thinking blocks, and cache_control markers on every cacheable position. |
| POST / | Local token estimate for SDK context management: nothing is dispatched, retained or metered. |
| GET / | Native Models discovery in Anthropic's shape (data, has_more, first_id, last_id, display_name, created_at, best-effort before_id/after_id/limit cursors), answered from the in-process catalog with no database round-trip. At the root path this shape is chosen by the anthropic-version header — without it you get the OpenAI shape; under /anthropic it always applies. Ids stay Sluis-callable provider/model ids, and created_at is a fixed placeholder because the catalog carries no per-model release date. |
| HEAD / | The warm-up probe Claude Code sends before a session · HEAD only, answered with 200 and an empty body. |
Model routing
The wire format never restricts the model: the ingress is a codec, and the model id owns the route. A bare id such as claude-sonnet-5 resolves through this ingress's default provider — anthropic unless your organisation points it elsewhere, for example vertex for EU-hosted Claude — after which the residency gate still decides whether that target may serve the request. Pin a target explicitly with a provider-prefixed id (vertex/claude-opus-4-8), or pass a sluis/* alias: an Anthropic-SDK tool then runs on any connected provider, inside your policy.
Streaming
With "stream": true the answer is the full Anthropic event grammar: message_start carrying the real input token count, content_block_start / content_block_delta / content_block_stop per block, message_delta with cumulative usage including cache read and creation tokens, then message_stop. Text arrives as text_delta, tool arguments as input_json_delta, and reasoning as thinking_delta plus signature_delta when the routed provider produced it. A ping event follows message_start and repeats after roughly 15 seconds of upstream silence, so a long thinking turn never trips a client watchdog, and a failure mid-stream is re-encoded as an error event instead of a truncated stream. stop_sequence carries the matched sequence when the answer came from an Anthropic-family provider, and is null otherwise.
# native Anthropic Messages — the model id owns the route, not the wire format curl https://api.sluis.ai/v1/messages \ -H "x-api-key: $SLUIS_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "vertex/claude-opus-4-8", "max_tokens": 256, "system": "Be concise.", "messages": [{ "role": "user", "content": "Say hi" }] }' # /anthropic/v1/messages is the unambiguous prefixed path for the same handler
{
"id": "msg_01f7…",
"type": "message",
"role": "assistant",
"model": "claude-opus-4-8",
"content": [{ "type": "text", "text": "Hi! How can I help?" }],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 9,
"output_tokens": 8,
"cache_read_input_tokens": 0,
"cache_creation_input_tokens": 0
}
}# "stream": true — the full event grammar, ping keep-alives included
event: message_start
data: {"type":"message_start","message":{"id":"msg_01f7…","role":"assistant","model":"claude-opus-4-8","usage":{"input_tokens":9,"output_tokens":0}}}
event: ping
data: {"type":"ping"}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hi!"}}
event: ping
data: {"type":"ping"}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":8,"cache_read_input_tokens":0}}
event: message_stop
data: {"type":"message_stop"}Errors
Every failure on these paths uses Anthropic's envelope — a top-level "type": "error" with the vendor error type (invalid_request_error, authentication_error, billing_error, permission_error, not_found_error, conflict_error, request_too_large, rate_limit_error, api_error, overloaded_error) — including the refusals that never reach a model, such as an unknown key or a data-protection block. The gateway's request id is on the response as the request-id header and in the body's request_id field, and it is the same id that identifies the call in the audit trail.
# every failure on /v1/messages* is Anthropic-shaped — auth and # data-protection refusals included, never an OpenAI envelope { "type": "error", "error": { "type": "rate_limit_error", "message": "rate limit exceeded" }, "request_id": "req_9f2e…" }
# mcp_servers is refused, not quietly passed through: an external MCP # server called by the model would sit outside the Sluis gate { "type": "error", "error": { "type": "invalid_request_error", "message": "mcp_servers is not supported: register the server in Sluis and call the MCP gateway at /v1/mcp" }, "request_id": "req_4c81…" }
One request field is refused rather than forwarded: mcp_servers returns 400 invalid_request_error. A model calling an external MCP server itself would sit outside the Sluis gate, so register the server in the Console and use the MCP gateway at /v1/mcp instead, where every tool call is inspected, routed, sealed and metered.
Compatibility matrix
What the native ingresses cover today, endpoint by endpoint. This table is the claim: what is not listed here as supported is on the roadmap, deliberately refused, or a different product surface — never quietly implied.
| Anthropic | Status | Notes |
|---|---|---|
| POST / | ✓ native | Buffered and streaming, tools, images, document blocks, thinking round-trip and prompt-cache markers. |
| POST / | ✓ local estimate | Answered locally; nothing is dispatched. |
| GET / | ✓ native | Native list and get shape from the in-process catalog; header-negotiated at the root path, unconditional under /anthropic. |
| HEAD / | ✓ | Claude Code's warm-up probe, HEAD only. |
| Message Batches | ✗ roadmap | Message Batches are not implemented; batch work runs as ordinary live calls today. |
| Files | ✗ roadmap | The Files API is not bridged; inline document blocks and file_data cover the common path. |
| mcp_servers | ✗ rejected | Refused with 400 instead of forwarded — register the server and call the Sluis MCP gateway at /v1/mcp. |
| Skills · Agents · Admin | ✗ out of scope | Skills, managed agents and the admin, usage and cost APIs are a different product surface, not inference compatibility. |
Cross-protocol routing. Vendor-only fields survive as long as the request lands on a provider that speaks the same protocol: an Anthropic request served by an Anthropic-family provider keeps its thinking blocks, cache markers, beta flags and matched stop_sequence. Routed cross-protocol, those fields are dropped at dispatch — never rejected at ingress — and the audit row records both the wire format and the resolved provider, so any degradation is attributable. Opaque values are never rewritten: thinking signatures and redacted_thinking payloads are forwarded verbatim or not at all, which is also why a parked thinking block that the data-protection scan flags is dropped rather than masked — a rewritten block would carry an invalid signature.