# Sluis — Full documentation (LLM edition)

> AI for your people. Control for your agents.
>
> Sluis gives employees a secure AI workspace and sends AI-agent requests through
> one governed gateway, with SSO, one policy and one audit trail. Sluis Workspace combines
> conversations, configured agents, versioned Skills and private or shared Knowledge.
> The gateway supports OpenAI and Anthropic Messages clients.
>
> The managed service is EU-hosted. API routing defaults to EU-only; organisations
> explicitly decide whether to allow more jurisdictions or require EU-owned providers.
> Data protection detects and reversibly pseudonymizes sensitive values before
> dispatch, and a hash-chained audit ledger records what policy ran. Detection is not
> perfect, and using Sluis does not by itself establish GDPR or HIPAA compliance.
> Agent Harness model traffic is the residency exception: the provider subscription
> chooses the serving region, which Sluis records as unverified rather than enforcing.
>
> This file is the complete public documentation as plain markdown, for AI answer
> engines and code assistants. Human version: https://www.sluis.ai/docs
> Site overview: https://www.sluis.ai/llms.txt

Base URL: `https://api.sluis.ai/v1` (OpenAI-compatible; use any OpenAI SDK unchanged).

Sluis Workspace is available in the browser at https://sluis.ai/workspace. Its public
product page is https://www.sluis.ai/sluis-workspace; https://www.sluis.ai/download lists desktop alpha
builds and mobile distribution options with availability per platform. The Agent
Gateway product page is https://www.sluis.ai/gateway. Sluis Edge is the optional
enterprise self-hosted data plane: https://www.sluis.ai/edge.

Pricing: https://www.sluis.ai/pricing. Self-service is prepaid only, with no free usage
tier: buy at least €25 of credit before paid requests. Optional auto-top-up needs a
credit-card mandate; it defaults to €25 when the balance reaches €5 and can be
disabled. Available one-off payment methods are shown at checkout. Managed model
usage costs provider list price + 10%. BYOK and custom OpenAI-compatible providers
cost €0.50 per 1M input + output tokens in Sluis fees; provider charges are separate.
No platform subscription or seat fee. The €25 purchase minimum is credit, not a
monthly usage commitment. Managed OCR costs €0.01 per page, floored at provider list
price; document anonymization defaults to €0.01 per page/image. AI name recognition
(NER) costs €0.005 (Swift) or €0.01 (Deep, GLiNER2) per inspected customer request;
the pattern detectors, directories and pseudonymisation are included. A completed
prompt-injection scan costs €0.01; workload scope-guard checks and hosted privacy
inspection are billed at their model's token price; Workspace Knowledge uploads are
billed like `/v1/embeddings` calls. Payment-method
surcharges apply. Invoice and manual billing are superadmin-approved exceptions,
not self-service alternatives; approved bank-transfer invoicing runs weekly with a
€75 invoice surcharge. The top-level organisation pays for its sub-organisations;
operator-enabled agency billing adds the configured child margin to future usage
without making the child a separate payer. Existing self-service accounts migrate
to prepaid without forgiving unbilled usage or outstanding charges. Buying credit
does not repay an older failed charge: use the outstanding-payment action in Billing.
Agent Harness adds €16.50 per active managed provider account per calendar month,
plus VAT and the payment-method surcharge, separate from metered usage. Edge has
a flat annual licence per gateway.

Public marketing and docs pages are maintained in English, French, German, Spanish,
Italian, Polish and Dutch: English uses bare paths, the others a `/fr`, `/de`, `/es`,
`/it`, `/pl` or `/nl` prefix with translated path segments (for example
`/nl/documentatie/api`). The application at `/workspace` has stable URLs without
locale prefixes. The complete public page map is at https://www.sluis.ai/llms.txt.

## Quickstart (https://www.sluis.ai/docs)

Sluis exposes an OpenAI-compatible API. Point your SDK's `base_url` at Sluis, use a
Sluis key and select a provider-prefixed model id or `sluis/*` alias. Requests then
flow through your residency policy and into the tamper-evident ledger. No Sluis SDK
is required; supported endpoints and cross-protocol limitations are documented below.

### 1. Get a key

Create a key in the Console's Workloads view (https://sluis.ai/console/workloads).
Workloads own rate limits, budgets, model allow-lists and policy settings; keys
are credentials under a workload. Residency is organisation policy. Fund prepaid credit before sending paid requests.

```shell
# keep it in your environment, never in source
export SLUIS_KEY="sluis-7Qk2vN4xR…"
```

### 2. Point base_url at Sluis

Swap the host. Everything downstream (models, streaming, tools, function calling)
works unchanged because Sluis proxies the same API surface.

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.sluis.ai/v1",
    api_key=os.environ["SLUIS_KEY"],
)
```

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.sluis.ai/v1",
  apiKey: process.env.SLUIS_KEY,
});
```

```shell
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest", "messages": [...] }'
```

Never used an OpenAI client? The same one-line move works for the Anthropic SDK
(Claude Code included); the native ingress is documented
under API reference.

```python
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.sluis.ai",   # the SDK appends /v1/messages
    api_key=os.environ["SLUIS_KEY"],
)
```

```shell
# Claude Code, in one line — no plugin, no config file
export ANTHROPIC_BASE_URL="https://api.sluis.ai" ANTHROPIC_AUTH_TOKEN="$SLUIS_KEY"
```


### 3. Send a request

Call it exactly as you would call the provider. Sluis inspects the request, routes it
through your policy, seals it in the audit chain, and returns the model's response
unchanged. Calls to a `sluis/*` alias disclose the concrete route in `x-sluis`
response headers.

```python
# Anthropic Claude, served from Google's EU multi-region
resp = client.chat.completions.create(
    model="vertex/claude-opus-4-8",
    messages=[{"role": "user", "content": "Summarise this chart…"}],
)
print(resp.choices[0].message.content)
```

Response headers on a `sluis/*` alias call (example):

```
x-sluis-route: sluis/auto
x-sluis-model: vertex/claude-opus-4-8
```

### 4. Set residency

Residency is organisation policy, set in the Console's Residency view
(https://sluis.ai/console/residency) and enforced at dispatch on API requests.
The default allows only the EU jurisdiction: a request that would reach a provider
outside the allowed set is refused with 403 before dispatch. EU ownership is a
separate restriction, not implied by an EU serving region. Broadening where data
may go is an explicit, recorded decision: the Console requires a transfer-terms
acknowledgement before non-EU providers unlock, and the change lands in the audit
trail. Agent Harness model traffic does not enforce residency; see its section below.

```shell
# default policy is EU-only — a US provider is refused at dispatch
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "openai/gpt-5.6", "messages": [...] }'
# → 403 permission_error
# "provider `openai` (jurisdiction `US`) is not permitted by the tenant's residency policy"
```

### 5. Verify the seal

Every call appends an entry to the hash chain; each entry's hash is
`sha256(prev_hash + record)`, so any altered field downstream breaks every link
after it. Read the audit metadata over the admin API. A full export and a chain
verification run through the gateway CLI.

In the Console's Audit log each request is one row. Expand it to see the checks it
ran: name recognition, LLM privacy inspection, the workload scope guard and the
security scan. Each check remains its own sealed entry in the chain, with its own
cost. `GET /admin/audit` returns them in the request's `checks` list; filtering by
`request_id`, `model`, `provider` or `decision=security.scan` lists them as rows.

```shell
# audit metadata, admin token with scope audit:read
curl "https://api.sluis.ai/admin/audit?from=2026-09-01" \
  -H "Authorization: Bearer $SLUIS_ADMIN_TOKEN"

# full export and chain verification are operator CLI subcommands
sluis-gateway audit verify-chain --tenant $TENANT_ID
# → chain intact
```

### Agent skill for coding assistants

An agent skill packages this reference for an AI coding assistant: one SKILL.md
plus reference files covering every endpoint, the exact error codes,
residency, data protection, keys and workloads. Download the archive from the
Quickstart page (https://www.sluis.ai/docs#agent-skill) and unzip it into your
assistant's skills folder.

## Videos (https://www.sluis.ai/docs/videos)

Films of Sluis in real use, recorded on the product with a fictional company. The picture is English; the voice-over and captions are available in English, Dutch, French, German, Spanish, Italian and Polish. Captions are on by default and can be switched off. Each film plays in 1080p or 1440p, and both files can be downloaded.

- How Sluis works: one gate (inspect, route, seal, meter) for every AI request, then what it means for Workspace users, compliance officers, IT and developers.
- Sluis Workspace, Monday at Northbridge: a customer email with personal data becomes a reply, a filed document and a scheduled agent.
- Sluis Gateway, Friday at Northbridge: a customer service agent goes live with its own workload and budget, the OpenAI SDK pointed at Sluis with the model set to Automatic, EU routing and a sealed audit trail.
- The console, feature by feature: every console area in turn (workloads, models and routing, tools and MCP, cache, Agent Harness, audit log, usage, residency, data protection, security, playground, Workspace administration, organisation settings), what each setting does and what changes when it is switched.
- Sluis Workspace, feature by feature: chat and Automatic model choice, attachments and generated files, documents and templates, knowledge and upload screening, agents, schedules and skills, memory and personal settings, and the admin switches that turn capabilities on or off for everyone.

## Sluis Workspace (https://www.sluis.ai/docs/workspace)

Use Chat, agents, Skills and Knowledge. Configure access, models, capabilities and
budgets for your organisation.

The illustrated guides show the English application interface with illustrative
demo data. Some service responses, including model availability, NER and guard
readiness, and Chat replies, are simulated for documentation. No production data
is used. Screenshots open at full size.

Open [Sluis Workspace](https://sluis.ai/workspace) or the
[Workspace console](https://sluis.ai/console/workspace).

### Start a conversation

Open Sluis Workspace and sign in with your Sluis account. Check the active
organisation in the account menu. Workspace must be enabled for that organisation;
ask an owner or admin if access is unavailable. The browser, desktop and mobile apps
use the same workspace.

The sidebar leads to Chat, Agents, Skills and Knowledge. Start a new chat, select an
available model and send your question. The model picker explains policy
restrictions; choosing another model never bypasses them. On desktop, Enter sends
when enabled and Shift+Enter adds a line; on touch keyboards Enter adds a line.
Reopen or search earlier conversations in the chat list or the command palette (⌘K).
The right column shows the context attached to a chat: its agent, Skills, Knowledge
and attachments.

While an answer is being written, the send button becomes Stop generating; on
desktop, Esc does the same. Stopping ends the turn and keeps the partial answer. In
the iOS app, Workspace notifies you when an answer or a scheduled agent run finishes
while the app is closed or in the background. The notification names the
conversation, never its content.

### Attachments and generated files

Use the paper clip in the composer to attach a file of up to 16 MiB: PDF, Word, Excel
(.xlsx, .xls), PowerPoint (.pptx), text, Markdown, CSV, JSON, PNG or JPEG. Inspect or
remove attachments before sending and resolve any upload error first. Availability
depends on organisation settings, private storage and the selected model.

Workspace reads attached documents, spreadsheets and presentations, sheet by sheet or
slide by slide when a file is large. Ask in ordinary language for a downloadable Word,
PDF, Excel or PowerPoint file, document anonymisation or an image. Created files
appear as Deliverable cards with a Download action. Attach a .docx or .pptx template
to the chat and Workspace can fill in its placeholders. Ask Workspace to save an
attachment to Knowledge: it stays private unless an owner or admin saves it for the
organisation.

These tools must be enabled for your organisation and supported by the deployment.
Image generation also needs an available image model permitted by policy. The steps
shown with an answer say what ran, for example Read a document, Searched the
knowledge base or Created a document. Generated files remain subject to retention and
access controls.

Canvas and Documents is on for every organisation; an owner or admin can switch
Documents off in the console under Settings › Features. Ask for a letter, email, slides or code and
the model writes it into a canvas in the chat that you can edit too. Every change,
yours or the model's, becomes a version; pick an earlier one under Versions to restore
it. Each canvas is saved in a folder in Documents, where a type filter narrows the
list; Open in Documents shows the document with its chat beside it. Download a canvas
as PDF, Word, PowerPoint, Markdown, text or an email file, depending on its kind.
Sending an email opens it in your own mail program.

Templates set a document's structure, styling and output, and house styles its look:
colours, font, page, logo and letterhead. Workspace picks a fitting template itself.
Open Templates from Documents, or use Save as template on a canvas to make your own.
Owners and admins manage the organisation's templates and house styles in the console
under Execute › Templates; members see them read-only and can duplicate one to change
it.

### Agents, Skills and Knowledge

Open Agents, choose an agent and start a chat, or type @ in the composer to select
one. An agent bundles instructions, an optional model, tools, Skills and Knowledge.
Create a private agent for yourself; organisation agents are shared and only owners
or admins may edit them.

Choose New in Agents to build an agent by describing the work in a chat. Workspace
asks a few questions and fills in the Configuration beside the conversation: name,
purpose, instructions, Skills, tools, Knowledge, model, schedule and scope. Choose
Create agent when it is right, or Set up the rest yourself. Edit on an agent's page
opens a conversation to change it; Workspace shows what changes before you apply it.
Skills are built the same way. In an ordinary chat, Workspace may offer to turn a
repeated task into a private Skill or agent, and creates it only after you agree.

The agent editor saves every change as you make it; there is no Save button. Under
Schedule, an agent runs once, daily or weekly at a set time and timezone. Each run
starts a fresh conversation with the agent's instructions and runs as the person who
set the schedule, under their budget and policy. The agent's page shows the next run
and recent scheduled and manual runs with their result.

Skills are reusable instructions with published versions: editing a Skill publishes
its next version, and every use records which version ran. Link the Skills an agent
may use in its editor. Knowledge holds private or organisation documents, including
spreadsheets and presentations: upload a document, inspect its status and link it to
an agent. If Knowledge screening refuses a document, its page names the categories
found; remove that information and upload a new version. No linked Knowledge
documents means the agent may search all documents visible to you; no linked Skills
means no Skill catalog. Check the sources under answers.

### Preferences, policy and receipts

Open Settings from the account menu. General sets language, dark, light or system
appearance and the 24-hour clock; Chat sets the default model, Enter behaviour and
whether replies show their route. Your data shows the current organisation policy
read-only. Personal preferences and agent settings cannot enable capabilities your
organisation has disabled.

Memory shows what Workspace remembers about your work, grouped by topic. Keep memory
decides whether new facts are stored; Bring into chats decides whether they are used
as context. Edit, add or delete a line, or clear the memory. When Workspace remembers
something new, it merges points that say the same thing and removes a point the new
one makes untrue; what you type yourself is kept as written. Agents also remember your
corrections, separately for each member: view and edit them under Your memory with
this agent on the agent's page. The same two switches apply. Connections links your
own account to MCP tools your organisation set up for personal sign-in. When a tool
needs that connection, the chat asks you to connect your account and links to
Connections.

Model and tool calls pass through the gateway controls. Inspect the route and receipt
when available; a seal records an audit event, not a guarantee that an answer is
correct. Data protection, permitted jurisdictions and retained content depend on the
effective policy. Review important output before using it.

### Administration and access

Owners and admins manage Sluis Workspace in the console. Enable the Workspace feature
in organisation settings, then open Workspace configuration. General controls whether
members can start conversations and sets a monthly euro budget per member; blank
means no per-member cap. Organisation budgets and billing still apply.

In Models, allow the full eligible catalog or select a restricted set and a default
from that set. Residency and provider availability still apply. Save settings to
apply changes, or Cancel to restore the saved configuration. Use Members to invite
colleagues and manage roles; being a member does not grant administration rights.

### Capabilities and shared resources

In the Features tab, configure attachments, document creation, Knowledge upload
screening and web search. Attachments need private storage; document creation depends
on attachments. Screening checks Knowledge uploads for secrets and personal data and
refuses a document in which it finds them. Web queries leave the residency boundary,
so keep web search off for an EU-only policy.

The same tab switches MCP tools in Workspace conversations and sub-agents. Turning
MCP tools off keeps connections and grants and does not affect external agents or
Playground. Sub-agents are off by default; only while they are on may an agent
delegate parts of a task to helper agents, whose usage counts toward the budget.

Manage shared agents, Skills and Knowledge in their console sections. Register MCP
servers and grant Workspace access under Tools & MCP. An agent only receives the
intersection of its selected tools and the organisation grants. Disabling an
organisation capability cannot be undone in an agent editor.

### Usage and troubleshooting

Check Usage & budget and the audit log in the active organisation. Self-serve usage
is prepaid and billed through the top-level paying organisation; a sub-organisation
does not become a separate payer. Owners manage credit in Billing. Operator-approved
invoice or manual arrangements are exceptions. Adding a document to Knowledge is
metered too: its embedding is priced like an embeddings call and counts toward your
monthly Workspace budget.

For access or policy refusals, ask an owner or admin to check Workspace access, model
restrictions, residency and tool grants. The code workspace_disabled means Workspace
conversations are off for the organisation; workspace_only_account means your account
has Workspace-only access, so the console is not available. For budget or credit
refusals, check the applicable cap and payer balance. On upload or connection
failure, keep your draft, resolve the reported cause and try again. To end an answer
that is still arriving, use Stop rather than sending again.

## API reference (https://www.sluis.ai/docs/api)

### Authentication

Every request to `/v1/*` authenticates with a virtual key in the
`Authorization: Bearer` header. Mint keys in the Console; the secret is shown once and
only its hash is stored. A key is a credential for a workload. The workload owns
rate limits, budget, allowed models and policy overrides. Requests remain subject
to organisation policy; an empty workload model allow-list inherits the
organisation's allowed scope.

A request whose model is not on the workload's non-empty allow-list is refused with a
sealed `403 permission_error` before anything is dispatched; the refusal itself lands
in the audit chain. Minting a production key requires a verified email address.

```shell
curl https://api.sluis.ai/v1/models \
  -H "Authorization: Bearer $SLUIS_KEY"
```

```json
{
  "object": "list",
  "data": [
    { "id": "mistral/mistral-large-latest", "object": "model", "owned_by": "mistral" },
    { "id": "vertex/claude-opus-4-8", "object": "model", "owned_by": "vertex" },
    { "id": "sluis/auto", "object": "model", "owned_by": "sluis" }
  ]
}
```

### Endpoints

Sluis exposes the OpenAI-compatible surface below. Endpoints without a first-class
handler are proxied verbatim to the routed provider, streaming included, so
OpenAI-compatible upstreams keep full fidelity. A browser that opens a path the API
does not serve (a `GET` or `HEAD` whose `Accept` includes `text/html`, such as the
bare API host) is redirected to these docs; API clients, including SDKs and
curl sending `Accept: */*`, still get a plain `404`.

| Endpoint | Purpose |
|---|---|
| POST /v1/chat/completions | Chat completions, the primary surface: routing, data protection, caching, streaming. |
| POST /v1/completions | Legacy text completions. |
| POST /v1/embeddings | Embeddings. |
| POST /v1/moderations | Moderation classification. |
| POST /v1/responses | The OpenAI Responses API. |
| POST /v1/systemone | Native typed decisions: `sluis/typed` answers your own questions locally; `sluis/typed/router` is the router behind `sluis/auto`. |
| POST /v1/ocr | Document OCR (Mistral, or fully local with model `sluis/ocr` — the gateway's own in-process OCR engine, inline `data:` URLs only, no provider egress) · billed per page. Returns `pages[]` markdown and `usage_info.pages_processed`. With the `dlp_documents` policy on, inline documents are anonymized before they leave the gateway, and the local `sluis/ocr` model applies the tenant DLP mode to its extracted text (block refuses, pseudonymize/mask rewrite). A busy local engine answers 503 after waiting up to 30 s; retry with backoff. |
| POST /v1/documents/anonymize | Document anonymization · docx, pdf, images, and text in; the same document out with PII swapped for «MERGE_TAG»s and images blurred · billed per page/image. |
| POST /v1/documents/anonymize/jobs | Async anonymization: enqueue large documents (`202` + job id, `Idempotency-Key` honoured; a job belongs to the enqueuing key's workload, so other workloads' keys get `404` and never replay it), poll `GET …/jobs/{id}`, download via `GET …/jobs/{id}/result` or hand a third-party app the time-limited signed `download_url` (`/v1/documents/deliverables/{id}` — MAC-authenticated, no API key). Redacted PDFs additionally carry an invisible searchable text layer built from the anonymized OCR text. |
| GET /v1/models | The models your policy and credentials can actually reach, nothing hypothetical. |
| GET /v1/models/{id} | One model's metadata. |
| POST /v1/audio/* | Transcription, translation, speech · proxied to the routed provider. |
| POST /v1/images/* | Image generation and edits · proxied. |
| POST /v1/video/generations | Video generation · proxied. |
| /v1/files | File operations · proxied on your organisation's own key (BYOK or a custom provider), never on managed keys. |
| POST /v1/messages | Native Anthropic Messages ingress · any Anthropic-SDK tool, any connected model. |
| POST /v1/messages/count_tokens | Local token estimate for Anthropic-SDK context management. |
| GET /anthropic/v1/models | Native Anthropic Models discovery · list and get in Anthropic's shape, from the in-process catalog. At `/v1/models` the same shape is selected by the `anthropic-version` header. |
| HEAD /api/hello | Warm-up probe for Anthropic-SDK clients such as Claude Code · 200 with an empty body (HEAD only). |

A canonical chat request and response:

```shell
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest",
        "messages": [{ "role": "user", "content": "Say hi" }] }'
```

```json
{
  "id": "chatcmpl-9f2e…",
  "object": "chat.completion",
  "model": "mistral-large-latest",
  "choices": [{
    "index": 0,
    "message": { "role": "assistant", "content": "Hi! How can I help?" },
    "finish_reason": "stop"
  }],
  "usage": { "prompt_tokens": 9, "completion_tokens": 8, "total_tokens": 17 }
}
```

Document OCR (`POST /v1/ocr`) returns `pages[].markdown` plus
`usage_info.pages_processed`:

```json
{
  "pages": [{
    "index": 0,
    "markdown": "# Invoice 2026-118\nAcme BV · Keizersgracht 1…",
    "images": [],
    "dimensions": { "dpi": 150, "height": 1754, "width": 1240 }
  }],
  "model": "mistral-ocr-latest",
  "usage_info": { "pages_processed": 1, "doc_size_bytes": 48213 }
}
```

Document anonymization (`POST /v1/documents/anonymize`, JSON delivery) returns the
rewritten document plus a summary; the mapping appears only when `include_mapping`
was requested and is never stored:

```json
{
  "filename": "contract.docx",
  "content_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
  "content_base64": "UEsDBBQABgAIA…",
  "summary": { "pages": 4, "images": 1, "categories": ["EMAIL", "IBAN", "PERSON_NAME"], "downgraded": false },
  "mapping": [
    { "token": "«PERSON_NAME_1»", "original": "Jan de Vries" },
    { "token": "«IBAN_1»", "original": "NL91ABNA0417164300" }
  ]
}
```

### System One typed decisions

`POST /v1/systemone` accepts native System One typed decisions, not OpenAI chat
messages. Plain `sluis/typed` answers your own `questions` on a local open model
(GLiNER2.5-Decide). `sluis/typed/router` is the router Sluis uses for `sluis/auto`:
send the text a request carries and read back what the router sees in it, plus the route
`sluis/auto` would take. Its questions are fixed, so send `state` only:

```bash
curl https://api.sluis.ai/v1/systemone \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"sluis/typed/router",
       "state":{"request":"Compare these two supplier contracts clause by clause and list the risks.",
                "previous":"Here are both contracts.",
                "attachments":["contract_a.pdf (application/pdf)","contract_b.pdf (application/pdf)"]}}'
```

```json
{
  "model": "sluis/typed/router",
  "answers": {
    "reasoning":  {"type": "noul", "noul": 0.9412,
                   "probabilities": {"false": 0.0588, "true": 0.9412},
                   "confidence": 0.9412, "answer_confidence": 0.9412},
    "complexity": {"type": "score", "score": 2.31,
                   "legend": {"0": "trivial", "1": "simple", "2": "moderate", "3": "hard"},
                   "probabilities": {"0": 0.01, "1": 0.12, "2": 0.42, "3": 0.45},
                   "confidence": 0.2612, "answer_confidence": 0.45},
    "task":       {"type": "choice", "choice": "document",
                   "probabilities": {"code": 0.01, "writing": 0.03, "translation": 0.01, "extraction": 0.05,
                                     "document": 0.86, "support": 0.02, "smalltalk": 0.01, "action": 0.01},
                   "confidence": 0.6888, "answer_confidence": 0.86}
  },
  "decision": {"route": "sluis/reasoning", "demand": "strong"},
  "revision": "<sha256 of the router weights>"
}
```

| Local-pack contract | Behaviour |
|---|---|
| Model ids | `sluis/typed` (open pack) and `sluis/typed/router` (the router). |
| `sluis/typed/router` | `state.request` (required), and optional `previous`, `system` and `attachments` (a string or a list of `"name (mime)"` strings); a plain string is the request. Answers `reasoning` (`noul`, P(needs multi-step reasoning)), `complexity` (`score`, 0 `trivial` to 3 `hard`, with `legend` and `probabilities` per level) and `task` (`choice` over `code`, `writing`, `translation`, `extraction`, `document`, `support`, `smalltalk`, `action`). `decision` is the route and demand `sluis/auto` would take from this text alone, or `null`. Files, images, JSON mode, tools, prompt size and your organisation's policy are not part of this input and can change a real `sluis/auto` route. `answers.reasoning` keeps the keys it always had; the rest is additional. |
| Plain `sluis/typed` | Always answered on Sluis infrastructure, never by an outside provider. When the local model is unavailable the call returns 503. `GET /v1/models` lists it only when it can answer. |
| TypeSafe | Bring your own key only; no `sluis/*` alias reaches it. Name the model explicitly, for example `typesafe/jev-latest`. The normal gate applies: it is a US provider, so your residency policy must allow `US`. |
| Availability | Every answer comes from verified model weights, reported as `revision`. If the local model is unavailable or cannot be verified, the call fails closed with HTTP 503. |
| State | Required JSON object, list (array) or string. Maximum 128 KiB (131072 bytes) of serialized JSON, checked before DLP. Missing or invalid state returns 422; oversized state returns 413. The open pack reads the first 384 words. |
| Questions | `sluis/typed/router` supplies its own calibrated questions: any `questions` field, including null or an empty object, returns 400. Plain `sluis/typed` requires 1 to 16 questions of type `noul`, `choice` (2 to 20 options) or `score` (2 to 20 levels), each with `instructions`; missing questions return 422, a malformed set 400, more than 64 KiB 413. The characters `[ ] ( )` and backticks are removed before the model reads them, so options that then collide return 400, as does a set longer than about 1,000 tokens. |
| Response | JSON with `model`, an `answers` map and `revision` (weights SHA256); `sluis/typed/router` adds `decision`. Each answer has its `type`, its typed value (`choice`, `score` or `noul`), `confidence` and, where the pack reports them, `probabilities`, `legend` and `answer_confidence`. On `sluis/typed` gate on `answer_confidence`; its probabilities are uncalibrated. No chat completion or streaming envelope. |
| Latency | `sluis/typed/router` usually answers well under 200 ms; when busy it returns 503 with `Retry-After: 2`. Plain `sluis/typed` runs all questions in one CPU pass: about 2 s for a prompt-sized state, up to about 7 s for a long one, with a 10 s budget. |
| Data protection | Tenant DLP applies locally to the state and to the questions' `instructions` and `criteria`: block refuses; tokenize or mask rewrites what the pack is shown. Per-request removal terms are honoured. |
| Residency | Processed on Sluis infrastructure with no outside provider and no upstream residency routing step. |
| Metering | Each successful local decision is sealed and billed per input token at €0.01 per 1M tokens; output is free. The pack reports no token count, so input is estimated at about 4 characters per token over the state and questions it was shown. Normal limits, budgets and model allow-lists still apply. |
| Aliases | An existing tenant alias with the same name takes precedence over the local pack. |

### Document input (`file` content parts)

Attach a PDF, Word, image, or plain-text file to a chat request as an OpenAI-style
`file` content part carrying a base64 `file_data` data URL. Any routed model accepts
it: providers that understand file parts natively receive them unchanged, and for
every other provider the gateway converts the document before dispatch. Provider-side
`file_id` references are not resolved; inline the bytes as `file_data`. Vision-capable
models get each PDF page rendered as an `image_url` part (a 48-page budget per request,
shared across all attached documents); models without image input get the document's
extractable text inlined as prompt text, which the data-protection scan then covers
like any other text. Longer or heavier documents belong on `/v1/ocr`. With the
`dlp_documents` policy on, the document is anonymized or refused before anything
leaves the gateway. When it is off and the organisation's DLP mode is protective
(`tokenize`, `mask` or `block`), documents that would travel as uninspectable
images are refused; under `allow_log` they pass with an explicit unscanned
marker in the audit trail.

```shell
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "scaleway/gemma-4-26b-a4b-it",
        "messages": [{ "role": "user", "content": [
          { "type": "file",
            "file": { "filename": "drawing.pdf",
                      "file_data": "data:application/pdf;base64,'"$(base64 -w0 drawing.pdf)"'" } },
          { "type": "text", "text": "Which discipline does this drawing document?" }
        ] }] }'
```

### API compatibility: OpenAI, vendor prefixes, and native Anthropic

The OpenAI-compatible surface is the native default at `/v1/*`. For drop-in use
with clients that expect a vendor path prefix, the identical surface is also served
under `/openai/v1/*` and `/langchain/v1/*` (the LangChain OpenAI client is
OpenAI-shaped, so these are aliases, not different wire formats).

Sluis additionally speaks the **Anthropic Messages** API natively. It is a codec in front of one pipeline: the request is translated
into the internal OpenAI shape, clears the same Inspect → Route → Seal → Meter
gate (routing, residency, data protection, audit, metering), and the response —
streaming included — is translated back into the vendor's format. There is no
pass-through mode: fidelity comes from park-and-restore, not bypass.

The wire format never restricts the model: the model id owns the route. Any
provider-prefixed id (`mistral/mistral-large-latest`, `deepseek/deepseek-chat`,
`moonshot/kimi-k2`) or `sluis/*` alias (`sluis/auto`, `sluis/code`, or your own
tenant alias) routes to that target, so an Anthropic-SDK tool can
run on any connected provider, subject to your residency policy. A bare vendor id
resolves through the tenant's per-ingress default provider (`anthropic` for
`/v1/messages`, configurable in the Console),
after which the residency gate can still reroute or refuse it; send a prefixed id
(`vertex/claude-opus-4-8`) to pin the target explicitly.

### Anthropic Messages API (native)

Endpoints: `POST /v1/messages` (buffered and streaming), `POST
/v1/messages/count_tokens` (local estimate — nothing dispatched, retained or
metered), `GET /v1/models` and `GET /v1/models/{id}` in Anthropic's shape, and
`HEAD /api/hello` (Claude Code's warm-up probe: 200 with an empty body, HEAD
only). All of them are also served unconditionally under `/anthropic/...`.

Authentication: `x-api-key: <key>` or `Authorization: Bearer <key>`. The
`anthropic-version` header is read but never required, and only its shape is
validated. Parsing is open-world: unknown body fields, unknown content-block
types, unknown `anthropic-beta` values and the `?beta=true` query Claude Code
sends are tolerated, never a 400. Beta flags travel byte-for-byte to
Anthropic-family upstreams and are ignored elsewhere.

Request coverage: system prompts as a string or as blocks; `text`, `image`,
`document`, `tool_use` and `tool_result` blocks (including `is_error` and image
blocks inside a tool result); history `thinking` / `redacted_thinking` blocks;
`tools` / `tool_choice`; `cache_control` on every cacheable position; `top_k`,
`metadata`, `output_config`, `service_tier`, `container`; `thinking` (which also
derives the internal `reasoning_effort`).

Response coverage: `content` blocks including `thinking` plus its `signature`
when the routed provider produced them, `stop_reason`, `stop_sequence`
(non-null only same-protocol), and usage with `cache_read_input_tokens` /
`cache_creation_input_tokens`. Every response carries a `request-id` header and a
`request_id` body field — the same id as its audit row.

Streaming grammar: `message_start` (with the real `input_tokens`) →
`content_block_start` / `content_block_delta` / `content_block_stop` per block →
`message_delta` (cumulative usage incl. cache tokens) → `message_stop`. Deltas are
`text_delta`, `input_json_delta`, `thinking_delta` and `signature_delta`. A `ping`
event follows `message_start` and repeats after roughly 15 s of upstream silence,
so a long thinking turn never trips a client watchdog; a failure mid-stream is
re-encoded as an `error` event instead of a truncated stream.

Errors are always Anthropic-shaped — authentication and data-protection refusals
included: `{"type":"error","error":{"type":…,"message":…},"request_id":…}`, with
`invalid_request_error` (400), `authentication_error` (401), `billing_error`
(402), `permission_error` (403), `not_found_error` (404), `conflict_error` (409),
`request_too_large` (413), `rate_limit_error` (429), `api_error` (500) and
`overloaded_error` (529).

`mcp_servers` is **rejected** with 400 `invalid_request_error`: a model calling an
external MCP server itself would bypass the Sluis gate. Register the server in the
Console and use the MCP gateway at `/v1/mcp`, where every tool call is inspected,
routed, sealed and metered.

```shell
curl https://api.sluis.ai/v1/messages \
  -H "x-api-key: $SLUIS_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{ "model": "vertex/claude-opus-4-8", "max_tokens": 256,
        "messages": [{ "role": "user", "content": "Say hi" }] }'
```

```json
{
  "id": "msg_01f7…",
  "type": "message",
  "role": "assistant",
  "model": "claude-opus-4-8",
  "content": [{ "type": "text", "text": "Hi! How can I help?" }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": { "input_tokens": 9, "output_tokens": 8,
             "cache_read_input_tokens": 0, "cache_creation_input_tokens": 0 }
}
```

### Compatibility matrix

What the native ingresses cover today, endpoint by endpoint. This table is the
claim: what is not listed here as supported is on the roadmap, deliberately
refused, or a different product surface — never quietly implied.

| Anthropic | Status | Notes |
|---|---|---|
| POST /v1/messages | ✓ native | Buffered and streaming, tools, images, document blocks, thinking round-trip, prompt-cache markers. |
| POST /v1/messages/count_tokens | ✓ local estimate | Answered locally; nothing is dispatched. |
| GET /v1/models · /v1/models/{id} | ✓ native | Anthropic list/get shape from the in-process catalog; header-negotiated at the root path, unconditional under /anthropic. |
| HEAD /api/hello | ✓ | Claude Code's warm-up probe, HEAD only. |
| Message Batches | ✗ roadmap | Not implemented; batch work runs as ordinary live calls today. |
| Files | ✗ roadmap | Not bridged; inline document blocks and file_data cover the common path. |
| mcp_servers | ✗ rejected | 400 invalid_request_error — register the server and use the Sluis MCP gateway at /v1/mcp. |
| Skills · Agents · Admin | ✗ out of scope | Skills, managed agents and the admin/usage/cost APIs are a different product surface, not inference compatibility. |

**Cross-protocol routing.** Vendor-only fields survive as long as the request
lands on a provider that speaks the same protocol: an Anthropic request served by
an Anthropic-family provider keeps its thinking blocks, cache markers, beta flags
and matched `stop_sequence`. Routed cross-protocol, those fields
are dropped at dispatch — never rejected at ingress — and the audit row records
both the wire format and the resolved provider, so any degradation is
attributable. Opaque values are never rewritten: thinking signatures and
`redacted_thinking` payloads are forwarded verbatim or not at all, which is also
why a parked thinking block that the data-protection scan flags is dropped rather
than masked (a rewritten block would carry an invalid signature).

### Streaming

Set `stream: true` and the response arrives as server-sent events: each frame is a
`chat.completion.chunk` delta and the stream ends with `data: [DONE]`. Streams are
never buffered in the gateway: they tee through it, and the audit seal and metering
happen even if the client hangs up early. Ask for `stream_options.include_usage` and
the final frame carries exact token usage, the same numbers the gateway meters and
bills.

After 15 seconds without a frame the gateway writes an SSE comment line
(`: keepalive`) between events, so proxies with idle timeouts keep the connection
open while a model thinks; SSE clients ignore it. Streams carry
`Cache-Control: no-cache` and `X-Accel-Buffering: no`. When a gateway replica
restarts it keeps serving open streams for up to five minutes; a stream still open
after that ends with an error event with code `server_restarting` (or, if the
upstream stalled mid-event at that moment, just ends), so retry the request.

```python
stream = client.chat.completions.create(
    model="mistral/mistral-large-latest",
    messages=[{"role": "user", "content": "Write a haiku"}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
```

The raw event stream on the wire:

```
data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Water"}}]}

data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" finds a way."}}]}

data: {"object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21}}

data: [DONE]
```

### Reasoning

Pass the OpenAI `reasoning_effort` parameter (`minimal | low | medium | high`) on any
thinking-capable model. Sluis translates it per provider (Gemini's thinking level,
Claude's adaptive thinking effort) and omits it where a model would reject it, so one
parameter works across the whole catalog.

```python
resp = client.chat.completions.create(
    model="vertex/claude-opus-4-8",
    messages=[{"role": "user", "content": "Prove it step by step…"}],
    reasoning_effort="high",  # minimal | low | medium | high
)
```

### Prompt caching

Repeated prompt prefixes — a long system prompt, tool definitions, a growing
conversation — can be cached at the provider so they cost less and answer faster.
Sluis speaks the industry-standard markers: Anthropic-style `cache_control` blocks
(with an optional 5m or 1h TTL) and OpenAI's top-level `prompt_cache_key`. Send them
in one shape; Sluis translates or strips them per provider, so the same request works
across the whole catalog.

| Provider | Prompt-cache behaviour |
|---|---|
| Anthropic | Explicit cache_control breakpoints on system blocks, message content and tool definitions, with a 5m or 1h TTL. Cache reads and cache writes are reported separately. |
| Bedrock | cache_control is translated to Converse cachePoint blocks in system, messages and toolConfig — same request, nothing to rewrite on your side. |
| OpenAI · Azure | Automatic caching on long prefixes. Markers are stripped; an optional prompt_cache_key is forwarded verbatim so repeat calls keep cache affinity. |
| Gemini · Vertex | Gemini models get implicit caching inside the provider — nothing to send, and the reported cached-token share is passed through to your usage. Claude on Vertex takes explicit cache_control, TTL included, exactly as on Anthropic. |
| xAI · DeepSeek · Mistral · Qwen · Moonshot · Zhipu | Automatic where the provider offers it. Explicit markers are sanitized out, so a request written for Claude is never rejected here. |
| Scaleway · custom | Scaleway has no provider prompt cache: directives are sanitized out and the gateway's own opt-in response cache is the lever there. A custom endpoint gets both directives unmodified — that behaviour is the operator's contract. |

Explicit prompt caching is organisation-level opt-in and off by default: a cached
prefix is data at rest at the provider — DLP-filtered by Sluis first, but stored
outside it for the cache TTL. Turn it on in the Console's Data protection view. Some
providers (OpenAI, Gemini, xAI, DeepSeek) cache automatically inside their own
infrastructure regardless; the toggle governs what Sluis explicitly forwards
(cache_control, cachePoint, prompt_cache_key), not provider-internal behaviour.
Without your own key for a provider, calls run on a Sluis credential that other
organisations share: on OpenAI and Azure, Sluis sets a prompt_cache_key derived from
your organisation in place of yours, and on every other provider the cached-token
counts in your usage show 0. Billing still applies the real cache discount.

Cache reads are billed at the provider's discounted rate; Anthropic-style cache
writes add the provider's write premium to the input rate. Both are visible per call
in the audit drawer and totalled in Usage & budget. On a Sluis-managed key, a request
that selects a surcharge Sluis cannot meter — a cache_control ttl other than 5m where
the provider honours it (Anthropic, Bedrock, Claude on Vertex), a service_tier other
than auto, default, flex or standard_only, or the Anthropic long-context beta
(context-1m) — is refused with 400 before anything is sent; your own provider key
forwards all three unchanged.

### Error codes

Errors use the OpenAI error envelope (`{"error": {"message", "type", "param",
"code"}}`); the `error.type` value mirrors the HTTP status, so SDK error handling
keeps working unchanged. The same conditions on the native Anthropic
path are returned in the vendor envelope instead (see the protocol section
above), including the auth and data-protection refusals that never reach a model.

| Code | When |
|---|---|
| 400 invalid_request_error | Malformed request body or parameters. |
| 400 invalid_request_error | Model id missing its provider prefix on the OpenAI-shaped surface. Use provider/model or a sluis/* alias; the native Anthropic ingress also resolves bare vendor ids through their configured default provider. |
| 400 invalid_request_error | Request carries the removed x-sluis-dlp header. Per-request overrides were replaced by workload option overrides; configure them on the workload in Console → Workloads. |
| 401 authentication_error | Missing or unknown API key. |
| 402 insufficient_quota | No active plan or budget reached. Activate a plan or raise the budget; the request never reaches a provider. |
| 403 permission_error | The key lacks permission, for example the model is not on its allow-list. |
| 422 invalid_request_error | Refused before dispatch, for example data protection in block mode matched the request. |
| 422 document_too_large_to_scan | A document is too large for the path it took: a page still over the render pixel cap at the minimum scan resolution, or a file over the text-extraction byte cap. Large-format pages (A0/A1 drawings) render downscaled automatically; the dedicated code lets a client shrink or split and resubmit. |
| 429 rate_limit_error | Rate limit reached. Enforced at the gateway; the request never hits a provider. |
| 403 permission_error | Blocked by residency policy: no allowed jurisdiction serves the request. The body includes the reason. |
| 403 scope_violation | Refused by the workload's scope guard in block mode: the call falls outside the scope written for that workload. The message names the rule the call breaks and the guard's reason (`request refused: outside this workload's scope (<rule>): <reason>`), and the request never reaches a provider. |
| 503 pdf_renderer_unavailable | An inline PDF had to be sent as page images because the routed model reads image input, but this deployment has no PDF renderer. It is refused rather than silently answered from extracted text, which would discard everything a drawing or a scan shows. |
| 503 auth_unavailable | The gateway could not read its key store, so it could not verify the key. The key is not known to be invalid; retry with backoff and honour Retry-After. |
| 503 scope_guard_unavailable | The workload's scope guard is in block mode but could not reach a verdict, so the call is refused instead of let through. Retry with backoff. |
| 504 api_error | The provider did not answer within the workload's answer timeout (`upstream_timeout_secs`): 300 seconds unless the workload sets 10 to 600, shorter if the request sent `x-sluis-upstream-timeout`. `error.sluis.cause` is `timeout`. The gateway did not replay the call because the provider may still be running it; shorten the prompt or ask for a longer timeout before resending. |
| 5xx api_error | Upstream provider failure after retries; the circuit breaker steers traffic around unhealthy providers. The body includes the provider's message and an error.sluis object naming the cause and attempts, their upstream status or transport error, and duration. |

Example envelopes:

```json
{ "error": { "message": "model must be provider-prefixed, e.g. mistral/mistral-large-latest",
             "type": "invalid_request_error", "param": null, "code": "invalid_request" } }
```

```json
{ "error": { "message": "model is not on this key's allow-list",
             "type": "permission_error", "param": null, "code": "permission_denied" } }
```

```json
{ "error": { "message": "blocked by residency policy: provider jurisdiction US is not in the allowed set [EU]",
             "type": "permission_error", "param": null, "code": "permission_denied" } }
```

## Residency & models (https://www.sluis.ai/docs/residency)

### Response headers

Calls to a `sluis/*` alias disclose the resolved route via `x-sluis-*` response
headers, and any call that carried an inline document discloses how that
document reached the model. A plain provider/model call with no document carries
neither. Residency itself is organisation policy, enforced at dispatch — never a
request header.

| Header | Description |
|---|---|
| x-sluis-route | Response, on sluis/* alias calls · the managed or tenant alias that resolved the request, e.g. sluis/auto. |
| x-sluis-model | Response, on sluis/* alias calls · the concrete provider/model target selected after policy and routing. |
| x-sluis-document-route | Response, on any call that carried an inline PDF or Word document · how it reached the model: images (rasterised to one page image per page), text (extracted text only — the pages themselves were not sent), native (forwarded to a provider that reads the document itself), or mixed. Absent when the request carried no document. |


### Managed aliases

Managed aliases are 13 reserved `sluis/...` model names: `sluis/auto`, `sluis/code`,
`sluis/chat`, `sluis/support`, `sluis/fast`, `sluis/cheap`, `sluis/agents`,
`sluis/extract`, `sluis/docs`, `sluis/longcontext`, `sluis/vision`,
`sluis/translate`, and `sluis/sovereign`. `sluis/auto` resolves per request among the
routes the tenant's policy allows: a classifier inside the Sluis deployment can pick the
reasoning route without any text leaving it; otherwise a hosted classifier reads three
short excerpts of the prompt after data protection, and only ever on a model from an
EU-owned provider processing in the EU (a GLOBAL provider is never used for it). The
use-case aliases are refreshed daily from Artificial Analysis benchmark snapshots and
curated by Sluis before they become active. Resolution is tenant-aware and
residency-aware: the same alias may resolve to a different target for an EU-only tenant
than for a tenant that allows US providers, and Sluis never routes a tenant to a
non-compliant provider just because that provider ranks higher.

| Alias | Use case |
|---|---|
| sluis/auto | Per-request routing. Sluis picks the best compliant target for the prompt. |
| sluis/code | Coding and code review. |
| sluis/chat | General conversation. |
| sluis/support | Customer-support answers. |
| sluis/fast | Low latency. |
| sluis/cheap | Best value for bulk work. |
| sluis/agents | Tool use and agentic workflows. |
| sluis/extract | Structured extraction. |
| sluis/docs | Document extraction and parsing. |
| sluis/longcontext | Long-context analysis. |
| sluis/vision | Image input. |
| sluis/translate | Multilingual translation. |
| sluis/sovereign | Open-weights only. |

`sluis/sovereign` restricts selection to open-weight models within the tenant's
residency policy; it does not itself require EU ownership. The local `sluis/ocr`
engine is not one of these routing aliases.

```shell
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "sluis/auto", "messages": [{ "role": "user", "content": "Plan a GDPR-safe rollout" }] }'
# response headers:
# x-sluis-route: sluis/auto
# x-sluis-model: mistral/mistral-large-latest
```

Tenant aliases shadow managed names. If an organisation creates `sluis/code`, that
route wins for that tenant. Key allow-lists may include reserved aliases such as
`sluis/auto`; the allow-list is checked before the alias resolves.

If the provider refuses the credential for a managed route's model (401, 403 or 404,
e.g. a model the account is not entitled to), the call moves on to the route's next
compliant model, and `x-sluis-model` names the one that answered. A concrete
`provider/model` or a tenant alias is never swapped.

### Supported models

Sluis exposes built-in providers through managed keys when a platform credential is
configured; connect your own key (BYOK) or add a custom OpenAI-compatible provider in
the Console to override or extend the catalog. On the OpenAI-shaped surface, pass a
provider-prefixed id (e.g. `mistral/mistral-large-latest`) or a `sluis/*` alias.
The native Anthropic ingress also resolves bare vendor ids as described in
API reference. The catalog is credential- and policy-aware, uses live provider
catalogs where available, and can fall back to declared models during an upstream
catalog failure. Vertex candidates are separately probe-verified for EU availability.
Nebius Token Factory is EU-owned, but its public `nebius/...` models have no region
guarantee: they count as `GLOBAL` and need an explicit GLOBAL opt-in. Only
`nebius-eu`, the operator's dedicated endpoints in an EU region, counts as `EU`.

The model list on https://www.sluis.ai/docs/residency comes directly from the models
callable through Sluis platform credentials. It follows the gateway's ten-minute
cache. The same data is public at https://api.sluis.ai/public/models.json (no auth);
per workload, `GET /v1/models` additionally applies its policy and allow-list.

```shell
# models callable through Sluis platform credentials
curl https://api.sluis.ai/public/models.json | jq '.models[] | "\(.provider)/\(.model)"'

# per key: the models your policy, credentials and allow-list actually reach
curl https://api.sluis.ai/v1/models -H "Authorization: Bearer $SLUIS_KEY"
```


## Data protection (https://www.sluis.ai/docs/data-protection)

Data protection runs on every request before dispatch. The default mode is
`tokenize`: every detected value is swapped for a stable typed token like `«EMAIL_1»`.
Detected values reach the provider only as tokens, and the response is restored to the
real values on its way back to you, streamed or not. The token map lives in memory for
the life of the request and is never persisted. Other modes: `mask` rewrites detected
values irreversibly, `block` refuses the request with 422, and `allow_log` passes it
through while flagging the audit row.

Pseudonymization is one way: it protects what you send out. In Sluis Workspace
(tokenize mode), public web search results and the model's own output from earlier in
the same turn get no fresh detection, no new tokens and no NER scan; a value the
conversation already tokenized stays replaced by its token inside them, so a result
cannot reveal what a token stands for. Prompt Guard and the workload scope guard still
read them, search queries stay tokenized, and the audit row records the one-way
exception. `/v1` callers cannot set this marker.

These modes act on detected values, not on a guarantee of complete PII recognition.
Coverage depends on enabled detectors, optional entity layers, input format and
policy overrides. Use the measured NER results below to assess detection limits.

What you send:

```shell
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "sluis/auto", "messages": [{ "role": "user",
        "content": "Mail j.devries@acme.nl that IBAN NL91ABNA0417164300 is active." }] }'
```

What the model provider receives (only stable typed tokens):

```json
{ "role": "user",
  "content": "Mail «EMAIL_1» that IBAN «IBAN_1» is active." }
```

What you get back, restored at the single client egress (streaming included):

```json
{ "choices": [{ "message": { "role": "assistant",
    "content": "Draft: Dear j.devries@acme.nl, your account NL91ABNA0417164300 is active…" } }] }
```

60 built-in detectors ship out of the box: 28 for personal data, from the US Social
Security number to a national-ID pack that checksum-validates 12 EU countries, and 32
for secrets and credentials. Each one
can be toggled per organisation, and custom terms (plain or regex) cover anything
specific to your business.

Personal data (28): Email, Phone number, IBAN, Credit card, IPv4 address, IPv6
address, MAC address, US Social Security number, Dutch BSN, Portuguese NIF, German
Steuer-ID, Polish PESEL, Belgian rijksregisternummer, French NIR (INSEE), Spanish
DNI/NIE, Italian codice fiscale, Swedish personnummer, Danish CPR number, Finnish
henkilötunnus, UK National Insurance number, EU VAT number, BIC/SWIFT code, Dutch
license plate, Dutch address, Passport number, Date of birth, GPS coordinates, Vehicle
identification number.

Secrets & credentials (32): API key (generic), AWS access key, AWS secret access key,
Private key (PEM), GitHub token, GitLab token, Slack token, Slack webhook URL, Discord
webhook URL, Google API key, Google OAuth refresh token, Stripe key, Mollie API key,
Anthropic API key, OpenAI API key, Sluis key, Hugging Face token, npm token, SendGrid
key, Twilio key, Shopify token, Vault token, Databricks token, Docker Hub token,
Telegram bot token, JSON Web Token, Credentials in URL, .env file dump, Azure storage
key / SAS, Password assignment, Confidentiality marker, High-entropy token (generic).

### Entity detection

Persons, organizations and locations are hard to catch by pattern alone. Sluis
combines context heuristics, email correlation, a tenant name directory, shipped
dictionaries and optional NER. The NER model runs as a network-internal sidecar:
text stays inside the deployment perimeter. These layers also inspect extracted
document and OCR text when document protection is enabled. Tokenize mode maps
detected entities to «PERSON_NAME_n», «ORGANIZATION_n» and «LOCATION_n».

With NER on, one request may name up to 100,000 distinct entities and carry up to
6.4 million characters of inspected text; past a ceiling the scan counts as incomplete.
Repeated mentions do not count toward the cap. A name the model detects once is replaced
wherever it recurs in the request, as the same whole word with the same casing and the
same token, including messages where the model did not report it ("Amsterdam" never
matches inside "Amsterdammers"). Lone generic words and short acronyms, and hits the
context rules reject, only cover the place the model reported.

The name dictionary is the deterministic counterpart to NER: given names and
surnames compiled from government open data into a compiled dictionary that
ships inside the gateway. It matches full names and honorific-anchored names
in-process, with nothing leaving your perimeter. Names that are also ordinary words
are only matched with name-shaped context around them, so rare names remain the job
of the directory or NER.

### NER model options

Choose Basic, Expanded or Deep in Data protection. Basic uses built-in detectors
and six deterministic name layers without a model call; new organisations start
here. Expanded adds Swift (spaCy); Deep adds GLiNER2-PII instead. Both model
profiles require complete inspection and block model failures or incomplete scans.
Deep costs more processing time, not a guarantee of higher recall. Custom allows
individual tuning. Existing policies, exclusions, directories and workload
overrides remain unchanged; unknown-word detection stays off in all three
profiles. A workload cannot enable NER when the organisation has it off.

AI name recognition is billed once per customer request it inspects (an API call, an
`/agent/mcp` tool call, a Playground request or a Workspace turn): €0.005 with Swift,
€0.01 with Deep; the Basic layer is included. Follow-up calls Sluis makes within that
request (agent loop iterations, sub-agents, tool calls, titles, retries, the router
classifier) are not billed again, and an incomplete inspection costs €0. Each charge is
a linked child audit receipt (`sluis/ner-swift` or `sluis/ner-deep`) that counts toward
workload, organisation and member budgets and prepaid credit.

Deep adds substantial latency before your chosen model starts generating a response.
Swift and GLiNER2 scan time depends on input length, message count and hardware.
The deadline covers the total NER scan across all messages, attachments and windows in
one request, not each window: 30 seconds for up to 320,000 characters, plus 30 seconds
per further 320,000 characters, at most 10 minutes. It is a maximum, not typical latency.
Optional hosted LLM inspection adds a separate wait before response generation,
depending on input length, provider load and retries; it is outside this NER deadline.

Any profile can enable hosted LLM inspection with Nemotron 3 Super
(`nebul/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16`).
Only the eligible text remaining after local protection is sent to EU-hosted Nebul
under the residency policy. Returned terms must match that text and are replaced
locally; an incomplete inspection is refused after at most three attempts.
The current Console report (2026-09-15) separates successful-case detail removal,
whole-document identifiability, harmless-item removal and technical reliability.
No removal percentage certifies anonymity or permits an international transfer.

| Model | Coverage | Measured accuracy | Latency |
|---|---|---|---|
| swift · spaCy xx_ent_wiki_sm | multilingual, basic | all-entity F1 0.58 · person F1 0.72 · recall ~0.53 | Depends on input and hardware |
| deep · GLiNER2-PII (mDeBERTa-v3, Apache-2.0) | 7 trained languages (EN, FR, ES, DE, IT, PT, NL) + multilingual backbone transfer | all-entity F1 0.61 · person F1 0.76 · recall ~0.70 | Depends on input and hardware |

Measured 2026-08-12 on an internal benchmark: WikiANN test sentences in six
languages (NL, EN, DE, FR, ES, IT; 150 per language), scored as micro-F1 over
(label, text) pairs at text level with the production false-positive filter
applied, on an Apple-silicon dev machine at FP32. WikiANN is Wikipedia-domain,
silver-standard data and spaCy's home turf, so read the numbers as relative
guidance between the tiers, not absolute field accuracy.

### Security scanning

Opt-in prompt-injection and jailbreak detection uses one fixed model:
**Llama Prompt Guard 2 86M** (`sluis/prompt-guard-2-86m`), served by a local
sidecar inside the Sluis deployment. There is no security model selector and no
hosted security alternative; the optional Nebul Nemotron 3 Super pass above is
the separate privacy layer, not a security scanner.
Prompt Guard is **Built with Llama**; its distribution retains the Llama 4
Community License and attribution notice.
There is no automatic model fallback. Three modes:
`off | log | block`. `off` skips scanning; `log` records findings without
blocking traffic; `block` refuses detected attacks and fails closed if
classification fails or remains incomplete.

The model inspects outgoing text after data protection and returns a binary
safe/unsafe verdict, not a calibrated confidence score. Off or logging-only data
protection can retain personal data. Security has its own bounded deadline,
separate from NER.

The total scan budget is 4096 to 65536 tokens (default 8192) and is an aggregate
excerpt budget. Prompt Guard scans selected segments in overlapping 512-token
windows, not a 512-token total budget. Selection uses a conservative byte bound.
Long inputs send selected excerpts (beginning, middle and end, plus
instruction-like cues and distributed samples), not all content. Sampling can miss
attacks: a sampled-safe result means only that no threat was found in inspected
excerpts. Audit records inspected and total bytes.

Each scan has a separate, linked `security.scan` audit entry. Each completed
logical Prompt Guard scan costs €0.01, independent of internal window count;
technical failures, Off mode and preflight cost zero. Charges can apply when the
parent is blocked or served from cache. Per-key overrides can change the scan mode.
Historical measurements do not establish current model performance.

The matched 2026-09-15 held-out benchmark is weak. At the default 8192-token
budget, Prompt Guard detected 31/126 completed attacks with 568/1008 benign
false alarms and no technical failures. Successful wrong verdicts remain
counted; failures and all-attempt costs are reported separately. The Console
links the full report, regenerated Prompt-Guard-only from that existing measured
lane; the historical two-model matrix stays archived unchanged. The fixed model
is a product setting, not a quality guarantee.

Separately, opt-in key-behaviour anomaly detection runs as a background job
(zero request latency): per-key baselines from robust statistics with
hour-of-week seasonality, plus a multivariate isolation-forest layer. Alerts
are explainable — never a bare score — and land in the Console's Security view,
optionally by email.

### Model notice

When tokenization rewrote a request, Sluis injects a leading system message telling
the model the «…» tokens are opaque placeholders it must keep intact; that is what
keeps the restore reliable. On by default; customise or disable it per organisation.

### Retention & audit fidelity

Content retention (request and response bodies for the audit log) is on by default
and encrypted at rest; audit fidelity chooses whether retained content stores the
tokens or the original values. Turn retention off for a metadata-only ledger; that
also disables the response cache.

### Document anonymization

Send a docx, pdf, image, or text file to `POST /v1/documents/anonymize` and the
same document comes back with detected PII and secrets replaced by merge tags like
«PERSON_NAME_1» in text, and blurred out of images and PDF pages. Redacted PDFs are
re-rendered and keep an invisible searchable layer built from the anonymized text;
this does not guarantee that detection finds every sensitive value.
OCR and deterministic/NER rewriting run inside the deployment. With hosted LLM
inspection enabled, the eligible text remaining after local protection is sent
to EU-hosted Nebul under the residency policy; incomplete inspection refuses delivery.
The operation is sealed in the tenant audit chain and metered
per page/image, and the token mapping is returned only on request
(`include_mapping`), never stored. Scope the detectors per request with
`entities` (category tags plus `person_name`/`organization`/`location`) and
pick `image_mode` `blur`, `strip`, or `skip`. The same protection works in
transit: with the `dlp_documents` policy on, files uploaded through
`/v1/files` and inline OCR documents are anonymized before they leave toward a
provider, refused under `block`, and scanned under `allow_log`. In Sluis Workspace,
attach a document and ask the agent to anonymize it: the `doc.anonymize` tool
runs the same pipeline and returns a downloadable deliverable.

Scans are never guessed at. Lines the OCR cannot read reliably (signatures, illegible
handwriting) and stroke ink outside every recognised text line (initials in a page
corner, stamps, line art such as charts) are destroyed whole. They stand as
«UNREADABLE» in the searchable text layer, and the response discloses the
`UNREADABLE_REDACTED` category (`summary.categories` in JSON; the binary response's
`x-sluis-anonymize-categories` header lists categories lower-cased). Ruled lines,
checkboxes and solid graphics such as logos and photos are left alone; a noisy blank
page passes, while a page with no readable text and only a solid graphic (a QR code)
is refused. Under `allow_log` the in-transit path forwards the original and records
it as unscanned. `sluis/ocr` under a rewriting document policy returns «UNREADABLE»
for such lines instead of text that may be a misread.

Async jobs for large documents:

```shell
# enqueue (202 + job id; Idempotency-Key honoured)
curl https://api.sluis.ai/v1/documents/anonymize/jobs \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -F file=@archive.pdf

# poll until succeeded; the signed download_url then needs no API key
curl https://api.sluis.ai/v1/documents/anonymize/jobs/9c31… \
  -H "Authorization: Bearer $SLUIS_KEY"
```

```json
{
  "id": "9c31…",
  "state": "succeeded",
  "filename": "archive.pdf",
  "summary": { "pages": 12, "images": 3, "categories": ["PERSON_NAME", "EMAIL"], "downgraded": false },
  "download_url": "https://api.sluis.ai/v1/documents/deliverables/9c31…?tenant=…&exp=…&sig=…"
}
```

### Embeddings

Pseudonymized tokens are stable within one request, not across requests, so
embeddings of tokenized text may not match between calls. The `dlp_embeddings`
policy setting controls whether the request-side scan covers `/v1/embeddings`:
it is on by default, and setting it to `off` sends embedding inputs to the
provider unscanned. Every exempted call is recorded in the audit trail.

### Per-request removal terms

Any JSON request on the data plane — `/v1/chat/completions`, `/v1/completions`,
`/v1/embeddings`, `/v1/responses`, and the native Anthropic `/v1/messages` ingress — may carry a top-level `sluis` extension whose `remove` list
names the terms the gateway must remove from the prompt before dispatch: a name
your detectors cannot know, an unreleased project, a customer.

```shell
# name the terms the gateway must remove — the top-level "sluis" object never leaves the gateway
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest",
        "messages": [{ "role": "user",
          "content": "Bas Alderding did a good job, send an email explaining how happy you are with that" }],
        "sluis": { "remove": [{ "text": "Bas Alderding", "kind": "person" }, "Project Nightingale"] } }'

# provider receives  → "«PERSON_NAME_1» did a good job, send an email explaining how happy you are with that"
# you receive back   → "Dear Bas Alderding, I am delighted with your work on Project Nightingale…"
# audit note         → dlp:tokenize:person_name(request)|dlp:tokenize:term(request)
```

```python
# the OpenAI SDKs forward the gateway extension through extra_body
resp = client.chat.completions.create(
    model="mistral/mistral-large-latest",
    messages=[{"role": "user", "content": "Bas Alderding did a good job…"}],
    extra_body={"sluis": {"remove": [{"text": "Bas Alderding", "kind": "person"}]}},
)
# the originals are restored on the way out — the provider only ever saw «PERSON_NAME_1»
print(resp.choices[0].message.content)
```

An entry is either a bare string — equivalent to `{"text": "…", "kind":
"term"}` — or an object with `text` plus a `kind` out of `person`,
`organization`, `location`, `term`. At most 128 entries; each `text` is
trimmed, must be non-empty, and is at most 256 characters. Matching is
word-boundary and case-insensitive (ASCII case folding). The `sluis` object
itself is a gateway extension and is always stripped before dispatch — it never
reaches a provider, the cache, or retained content.

Every match is replaced by a reversible token («PERSON_NAME_1»,
«ORGANIZATION_1», «LOCATION_1», «TERM_1») before dispatch. Under `tokenize` —
the default — the terms join the normal pseudonymization pass and the original
values are restored in the response, streaming included; under `mask` they are
masked as `[REDACTED:<kind>]`; under `block` a request carrying a listed term is
refused with 422.

The instruction is always honoured. Even when the key sets `dlp: off`, the
tenant mode is `allow_log`, or `/v1/embeddings` is exempted with
`dlp_embeddings=off`, the listed terms are still tokenized, and the audit trail
discloses exactly that
(`dlp:tokenize:person_name(request)|scan:off:key_override`). The directive can
only strengthen protection, never weaken it — which is why it is allowed per
request where the retired `x-sluis-dlp` header was not, and why the listed term
never reaches the provider.

A malformed directive is refused with 422, never silently ignored: an unknown
`kind`, an entry over the limits, or an unknown key inside `sluis`. Audit notes
attribute caller-supplied terms with the `(request)` layer, distinct from
`(ner)` and `(directory)`, so the trail shows where each removal came from.

`POST /v1/documents/anonymize` — and its async jobs, and the Console
playground — accepts the same list as a `remove` field alongside `entities`.

### Workload option overrides

Overrides belong to the workload and apply to all its credentials. An owner or
admin configures `option_overrides` in Console → Workloads or with
`POST /admin/workloads` when creating a workload. Unset options inherit organisation
policy. To edit an existing workload, `PATCH /admin/workloads/{workload_id}` requires
its complete mutable settings, including name, status, limits and budget; it
replaces them rather than merging individual fields.

```shell
# create a workload with governed overrides; keys inherit its settings
curl -X POST https://api.sluis.ai/admin/workloads \
  -H "Authorization: Bearer $SLUIS_ADMIN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "name": "document-review", "option_overrides": {
        "dlp": { "mode": "off" },
        "ner_model": "deep",
        "security": { "prompt_injection": { "mode": "block" } } } }'
```

The `x-sluis-dlp` request header has been removed. Requests carrying it receive 400;
configure the workload instead. Effective overrides are recorded in the sealed
audit trail.

## Budgets & caching (https://www.sluis.ai/docs/limits)

### Budgets & limits

Each workload owns rate limits (requests and tokens per minute) and a spend budget
shared across its credentials. Choose a budget period: `total`, `daily` or `monthly`. An organisation-wide aggregate cap sits above
all workloads. Enforcement happens at the gateway, in real money: every request is priced
from the live price book and debited before dispatch. Past the rate limit the call
returns 429; past the budget, 402. The request never reaches a provider. If the
counters are ever unreachable, the check fails closed and reconciles from the
immutable audit ledger instead of guessing.

Separately metered checks count toward the same budgets: AI name recognition,
prompt-injection scans, workload scope-guard checks and hosted privacy inspection each
seal a linked receipt that is debited like any usage. Workspace Knowledge uploads are
metered the same way: the embeddings a document needs (sent in batches) are priced like
a `/v1/embeddings` call and charged to the organisation, its prepaid credit and the
uploading member's budget. An upload is refused with 402 before anything is embedded
when the organisation budget or credit is exhausted.

```json
{ "error": { "message": "rate limit exceeded",
             "type": "rate_limit_error", "param": null, "code": "rate_limit_exceeded" } }
```

```json
{ "error": { "message": "budget exceeded",
             "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } }
```

### Caching

An opt-in exact response cache matches normalized requests. It is strictly isolated
per organisation, encrypted at rest, and requires content retention. Enable it in
the Console's Cache view. New audit rows record how the call was served:
`none | exact`.

```shell
# identical request twice — the second is served from the exact cache
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest", "messages": [{ "role": "user", "content": "Define GDPR" }] }'
# audit row of call #2: cache_hit: exact · no provider dispatch, no token cost
```

## MCP tools (https://www.sluis.ai/docs/mcp)

Register external MCP tool servers (HTTP, SSE, or stdio) in the Console with residency
tags; auth secrets are sealed in the credential vault and never echoed back. Every
tool call clears the same gate as a chat completion: Inspect (data protection on the
arguments), Route (residency on the server), Seal (audit chain), Meter (rate and
budget). Tool grants are per-key and deny-by-default. Sluis also answers
`POST /v1/mcp` as an MCP server itself, exposing exactly the tools the calling key is
granted.

```shell
# list the tools this key is granted (deny-by-default)
curl https://api.sluis.ai/v1/mcp \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -d '{ "jsonrpc": "2.0", "id": 1, "method": "tools/list" }'

# invoke a granted tool — the call clears Inspect → Route → Seal → Meter
curl https://api.sluis.ai/v1/mcp \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -d '{ "jsonrpc": "2.0", "id": 2, "method": "tools/call",
        "params": { "name": "crm.lookup", "arguments": { "email": "j.devries@acme.nl" } } }'
```

```json
{
  "jsonrpc": "2.0",
  "id": 2,
  "result": { "content": [{ "type": "text", "text": "1 match: account 8817, status active" }] }
}
```

### MCP authentication: two directions

#### Static server credentials

Use headers for fixed headers. The legacy oauth value means a static OAuth bearer token stored in the vault, not browser sign-in.

#### Account mode

For interactive OAuth, Organisation (account_mode=organization) uses a shared identity explicitly connected by an admin with OAuth. It never reuses a personal credential. Each user (account_mode=user) uses only the acting member’s connection. Authentication is separate: interactive OAuth uses user_oauth; the legacy oauth value remains a static bearer token.

#### Personal upstream connections

In Each user mode, sign in to the upstream service in Workspace → Settings → Connections. Admin-added servers must first be granted to Workspace. A connection does not bypass tool grants, residency, data protection or budgets. Admins can view status and revoke connections, but cannot use another member’s identity. Calls without a user actor, including organisation-owned keys, are refused in Each user mode. GET/PUT /chat/settings exposes mcp_enabled (boolean, default true for new organisations, legacy settings and an absent response field). Omitting it from PUT or sending null preserves the current value, including false; the BFF preserves that distinction too. When false, Workspace does not discover MCP tools and rejects their execution, including previously discovered tools in existing conversations, Agents and subagents. Registrations, approvals, grants and credentials are preserved. External /agent/mcp, /v1/mcp and console Playground are unaffected. Personal server registration policy remains independent and organisation-wide. Owners and admins switch it with MCP tools in Workspace conversations in the console's Workspace configuration. When a Workspace tool call reaches an Each user server you have not connected, or the service no longer accepts your sign-in, the tool returns connect_required and the answer shows a Connect action for that server; nothing runs under another identity.

#### Personal MCP suggestions

Personal MCP suggestions are Off by default. Require admin approval lets members submit private, requester-only servers. Admins review the tools, residency and authentication explicitly; approval never grants wildcard access or automatically promotes a server to the organisation. Private servers support none (no upstream sign-in) or interactive OAuth, not personal static secrets. Off pauses existing personal servers, but does not block sign-in to admin-added servers. A member's request names only the server and its endpoint; while pending it neither contacts the server nor grants tools. Approval sets the reviewed configuration and an explicit list of tool names, available to the requester alone, who then signs in under Connections if the server uses OAuth. Removing an approved request revokes the server and its connections.

#### Sign in to Sluis from an MCP client

Inbound OAuth at /agent/mcp signs the client in to Sluis through the browser. It is separate from upstream user_oauth sign-in. Approve access in the intended organisation. Sluis creates or reuses a personal tool-plane credential for this client automatically. Existing plugin bindings, tool grants and policy still apply. No key is pasted into this OAuth tool-plane setup; model-plane key configuration is unchanged. The client finds Sluis through the published discovery documents, registers itself and opens your browser on the consent page, /console/mcp-authorize, which shows the client, the organisation and the tools it can reach; approve or deny there. The organisation needs Agent Harness enabled. Access tokens last one hour and refresh tokens rotate on every use. Revoke a client under Agent Harness › Connected clients, which disconnects its tokens and personal agent key.

```shell
claude mcp add --transport http sluis https://<gateway>/agent/mcp

codex mcp add sluis --url https://<gateway>/agent/mcp
codex mcp login sluis
```

Configuration examples, not a verified interoperability matrix. External Claude, Codex and upstream provider interoperability has not been established by these examples.

## Agent Harness (https://www.sluis.ai/docs/agent-harness)

Claude Code, Codex and Cursor keep running on the developer's machine: their processes,
files, built-in tools and agent loop stay local. Sluis governs the traffic they send out,
the model requests and the tool calls, not the agent itself. The local client
authenticates to Sluis with an agent key and reaches `/agent/claude` (Anthropic API, for
Claude Code), `/agent/codex/v1` (OpenAI Responses API, for Codex), `/agent/cursor`
(Cursor's Connect/protobuf transport) and `/agent/mcp` (stateless Streamable HTTP MCP).
The harness governs *how* an agent works as well as what it may reach: capability-free
`context` and `skill` plugin versions, bound at organisation, user or key scope, are
injected into the governed model request (see "Context and skills injected by the
gateway" below).

The harness is not restricted to those named clients: Pi, T3 Code or another client
that accepts an Anthropic Messages, OpenAI Responses or MCP base URL and API key can
use the corresponding ingress.

Cursor's wire is the one that is not JSON. Sluis does not proxy it opaquely: it walks the
protobuf frames without a schema, hands every text leaf to the same detectors and tenant
scanner that protect `/v1/chat/completions`, and rewrites those leaves in place, so
`tokenize` restores the originals in the streamed response and `mask` substitutes
`[REDACTED:<category>]`. A frame the codec cannot decode is refused with `422` under
`tokenize`, `mask` or `block` rather than forwarded — Sluis does not forward what it
could not inspect — and is forwarded only under `allow_log`, sealed with
`dlp:allow_log:not_inspectable`.

Two further Cursor limits matter: the name-detection sidecar and prompt-injection
guard require a JSON request, so neither runs on Connect/protobuf. The audit row
records them as not run; mandatory name detection refuses the turn instead of
forwarding it without that scan. Cursor's transport also reports no token usage,
so the turn is sealed without token counts. Context/skill injection is not supported
on Cursor either, as detailed below.

Purpose isolation is the rule most often missed: an agent key authenticates **only** on
`/agent/*`, an ordinary API key **only** on `/v1/*`, and either one presented on the
wrong surface answers 401 without disclosing the key's purpose. Each surface also
accepts exactly one credential form: `x-api-key` on `/agent/claude/*` (what Anthropic
clients send), `Authorization: Bearer` on `/agent/codex/v1/*`, `/agent/cursor`
and `/agent/mcp`.

### Connect a subscription

An organisation connects one or more Claude and Codex subscription accounts in the
Console (Agent Harness → Connections). Claude is connected with a token from
`claude setup-token`; Codex through OpenAI's device authorization. Sluis proves the
credential with the provider, then encrypts it, stores it in the gateway vault and
refreshes it there. Provider tokens never reach the local client, the logs, the audit
records or the traces. Connecting is member self-service; owners and admins see every
account in the organisation, a member sees and disconnects only their own.

### Claude: setup token

Prerequisite: an active Claude subscription on the account being connected.

1. Run `claude setup-token` on the developer machine. It prints a long-lived token, valid
   about one year, beginning `sk-ant-oat`.
2. Paste that token into the masked Claude field in the Console (Agent Harness →
   Connections).
3. Sluis verifies the token with Anthropic before storing anything.
4. A verified token is sealed encrypted in the gateway vault: never shown again, never
   returned to a client, never sent anywhere but Anthropic.

Failure: a token Anthropic cannot verify is refused with `422` — no account is created,
nothing is stored, nothing is billed. Reconnecting the same subscription reuses the
existing account instead of creating a second one.

### Codex: device authorization

Prerequisite: a ChatGPT plan that includes Codex.

1. Sluis obtains a verification URL and a user code from OpenAI.
2. The user opens the URL, signs in, enters the user code and approves the request; MFA
   may be required.
3. The Console polls for approval and runs the token exchange in that browser tab.
   Leaving the page or switching to another tab of the view cancels the flow and stores
   nothing, exactly as Cancel does.
4. On approval the access and refresh tokens are stored encrypted gateway-side and
   refreshed automatically, with rotation serialized across replicas.

Failure: an expired user code cannot be resumed and the flow must be restarted. A
credential the upstream later rejects shows as `expired` and needs reconnecting.

### Order of operations

1. Connect the subscription (Claude setup token, or Codex device authorization).
2. Create an agent key.
3. Link that key to the connected provider account; a provider the key has no linked
   account for fails with `502 no credential configured for provider ...`.
4. Configure the local client: `SLUIS_AGENT_KEY` in the environment, then
   `/agent/claude` for Claude Code, `/agent/codex/v1` for Codex, or
   `/agent/cursor` for Cursor.
5. Register the tool endpoint `/agent/mcp` separately, because a base URL alone
   adds no tools.

Purpose isolation holds throughout: agent keys authenticate only on `/agent/*`, ordinary
API keys only on `/v1/*`, and either one on the wrong surface answers 401.

### Create an agent key

Agent keys are minted on the Agent Harness view, separately from API keys. The secret is
shown once and only its hash is stored. Each key is routed through explicitly selected
provider accounts, so connecting a second member's subscription never reroutes an
existing key. A key spends only the subscriptions it is linked to: if a request names a
provider the key has no linked account for, the dispatch fails with a terminal
`502 no credential configured for provider ...` rather than falling back to the
organisation's other credentials.

The key's tool list is the organisation's recorded ceiling on the local client's
built-in tools, written in Claude Code's tool grammar and validated at mint time. The
developer's client enforces it on the machine; Sluis never reads it on a request, so no
key holder can widen their own key, and an empty list records no ceiling. The tools
Sluis gates server-side, per request, are the MCP tools it exposes over `/agent/mcp`.

### Point the clients at Sluis

The Console generates this block with the deployment's public gateway URL filled in.

```shell
# the agent key is shown once, at creation: keep it in the environment
export SLUIS_AGENT_KEY='sluis-9f2c…'
export ANTHROPIC_BASE_URL='https://api.sluis.ai/agent/claude'
export ANTHROPIC_API_KEY="$SLUIS_AGENT_KEY"
```

`ANTHROPIC_BASE_URL` carries no plugin or MCP discovery, so setting it alone never adds
a tool. Register the MCP endpoint explicitly, either with the CLI or with a project
`.mcp.json`; both reference `${SLUIS_AGENT_KEY}` so no secret lands in a committed file.

```shell
claude mcp add --transport http sluis 'https://api.sluis.ai/agent/mcp' \
  --header 'Authorization: Bearer ${SLUIS_AGENT_KEY}'
```

```json
{
  "mcpServers": {
    "sluis": {
      "type": "http",
      "url": "https://api.sluis.ai/agent/mcp",
      "headers": { "Authorization": "Bearer ${SLUIS_AGENT_KEY}" }
    }
  }
}
```

Codex reads its provider block from `~/.codex/config.toml`:

```toml
model = "openai/gpt-5.4"
model_provider = "sluis"

[model_providers.sluis]
name = "Sluis"
base_url = "https://api.sluis.ai/agent/codex/v1"
env_key = "SLUIS_AGENT_KEY"
wire_api = "responses"

[mcp_servers.sluis]
url = "https://api.sluis.ai/agent/mcp"
bearer_token_env_var = "SLUIS_AGENT_KEY"
default_tools_approval_mode = "approve"
```

### MCP plugins

`POST /agent/mcp` is stateless: it keeps no session, and every request re-derives the
whole authorization chain, so revoking a key, a plugin or a signing key takes effect on
the next call. Two classes of plugin exist. Sluis-maintained plugins are compiled Sluis
code, offered organisation-wide, and execute no tenant code. Organisation plugins are
the tenant's own: a TOML manifest signed with Ed25519, accepted only when its signature
verifies against a signing key the organisation has enrolled, immutable per
`(organisation, id, version)`, and pinned to one governed HTTPS MCP server whose
endpoint must match the URL inside the signed manifest exactly.

That is the flow that exposes tools. The two capability-free kinds that inject text rather
than granting reach — `context` and `skill` — share the same registry and the same
immutability but are console-authored and unsigned; see the next section.

Operator workflow: the owner enrols the organisation's signing key, an owner or admin
registers the signed manifest against a governed MCP server, then binds the plugin to
the agent keys that may use it. A plugin bound to nothing is exposed to nobody. Server
URLs, manifests, signatures and upstream credentials never reach a client, which sees
only tool names and schemas.

### Context and skills injected by the gateway

An organisation's engineering standards are only policy if a developer cannot forget them.
A `CLAUDE.md` in the repository is a suggestion — editable, deletable, and nothing records
which happened. Two capability-free plugin kinds make the same text organisation policy
instead: a `context` entry carries verbatim text, a `skill` entry carries a name, its
instructions and the tool names it expects. Both cap at 16 KiB, must be non-empty after
trimming and may not carry control characters other than tab, carriage return and newline.
A skill's tool names are *intersected* with the key's recorded tool ceiling at request time
and are never a grant of their own, so no skill can widen a key; a declared tool the ceiling
does not carry is simply left out of what the model is told.

Bindings carry a scope. `organisation` reaches every agent key in the organisation,
including keys minted later; `user` reaches every key one member owns; `key` is the exact
per-key binding that already existed. Resolution is deterministic, because injected order
changes the prompt: the union of the three scopes, each plugin re-verified exactly as
`/agent/mcp` re-verifies it, deduplicated by `(plugin_id, version)`, ordered organisation →
user → key, and within one scope by `(kind, plugin_id, version)`. The general rule is read
first, the narrow exception last. A `context` block injects its text verbatim; a `skill`
injects `## <name>` followed by its instructions.

Injection sites: on `/agent/claude` the blocks are prepended to `system`, in either
documented shape (a bare string or an array of content blocks); on `/agent/codex/v1` they
are prepended to `instructions`, or inserted as a leading `developer` item when the request
carries none. Total injected bytes are capped at 32 KiB; past the cap the request is
refused with `422` naming the plugins that overflowed it, never silently truncated —
half an instruction is worse than none.

`/agent/cursor` receives no injection. Cursor's Connect/protobuf schema is proprietary:
Sluis can find and rewrite text leaves inside those frames, which is how DLP works there,
but it cannot know which leaf is the system prompt, and writing into a guessed field would
corrupt the request rather than govern it. A Cursor turn carries organisation context only
if the developer's own client sends it.

These two kinds carry no Ed25519 signature, deliberately. Signing is mandatory where a
manifest points a governed credential at an endpoint — an unsigned `mcp_http` entry could
aim the organisation's vault-sealed upstream auth at a host nobody approved. A block of
text reaches nothing, so a signing ceremony there would protect nothing. Everything else
still holds: immutable per `(tenant, plugin_id, version)`, canonical TOML hashed into
`manifest_sha256`, the acting user recorded, and a database trigger that permits only
`active -> revoked`. A manifest carrying an MCP entry keeps the mandatory signature and the
per-request enrolled-signer check.

Injected text is not DLP-rewritten, and the order is one-way: the caller's content is
scanned and rewritten first, then the policy is prepended. The blocks are the
organisation's own governed content, so there is no party to protect them from, and
tokenizing them would destroy what they encode — "escalate to security@acme.example before
pasting" rewritten to "escalate to «EMAIL_1» before pasting" is no longer actionable.

Every injected block adds the marker `context:<plugin_id>@<version>#<digest8>` to the
sealed audit row's note, where `digest8` is the first eight hex characters of that
version's `manifest_sha256`. Because a version is immutable and its digest is one of the
columns the trigger refuses to change, the ledger alone proves what the model was told.

### Governance and billing

Agent traffic uses pseudonymisation and DLP before dispatch (including inside
Cursor's protobuf frames), workload rate and budget limits, plugin and tool governance,
and audit plus metering on model and tool calls, subject to the Cursor scan and
token-reporting limits above. An agent key cannot loosen organisation policy.

The exception is residency, and it is not a gap. An agent turn dispatches on the
developer's own provider subscription, so Anthropic, OpenAI or Cursor — not Sluis —
chooses which region serves it, and Sluis can neither select nor observe that region.
Refusing a turn on jurisdiction would block it over a region Sluis never chose; rerouting
it would send the prompt to an upstream the subscription does not cover; and recording the
provider registry's jurisdiction would be a verified-looking `EU` that nothing verified.
So the residency policy is not applied on the agent plane, and the sealed row says so: no
region, the `unknown` jurisdiction disclosure, and the note marker `region_unenforced`,
which the console renders as "region not enforced" rather than "in-region". Residency is
untouched where Sluis does choose the destination — the `/v1/*` data plane, and the
governed MCP servers Sluis dials itself.

Each active Sluis-managed provider account costs €16.50 list price per calendar month,
charged at most once per account per month, billed to the organisation's top-level
billing root as its own invoice line, separate from metered usage. VAT and the
payment-method surcharge behave exactly as on the rest of the invoice. Disconnecting an
account stops future months; reconnecting the same upstream subscription reuses the
existing account instead of adding a second billable one, so a disconnect and reconnect
inside one month still bills once. Gateway usage itself stays consumption-billed.

The whole Agent Harness is behind the per-organisation `agent_harness` rollout flag,
enforced by one gateway middleware over the `/agent/*` data plane and the
`/admin/agent-harness/*` control plane alike. It fails closed: the flag resolving to
off, and any failure to resolve it at all, both answer `404`. The Console surface
appears once the flag is enabled and renders an unavailable state otherwise.

### Troubleshooting

- `401` on an `/agent/*` call: wrong surface or wrong header form (see purpose isolation
  above).
- `502 no credential configured for provider ...`: the key has no linked account for the
  provider the request named. Link one, or name a model the linked account serves.
- Subscription refresh rejected: the provider refused the grant permanently, so the
  account stops serving until a human reconnects it. Reconnecting reuses the same
  account and its existing monthly fee.
- `400` `mcp_servers` is not supported: a remote-MCP declaration in a Responses body is
  refused before dispatch. Register the server with Sluis and call it over `/agent/mcp`.
- No setup snippets in the Console: the deployment has no validated public gateway URL.
  Configure `SLUIS_GATEWAY_PUBLIC_URL`, then read the endpoints from the Connections tab.

## Admin API (https://www.sluis.ai/docs/admin-api)

Scoped, expiring tokens for administering an organisation without a browser session:
invoicing exports (Moneybird, Exact, …), provisioning client organisations, members and
API keys from a CRM or IdP sync, and audit feeds for a SIEM.

Create a token under Settings → Admin API (https://sluis.ai/console/settings/admin-api),
owner or admin with a verified email address. Name it, tick its scopes, pick an expiry (up
to 365 days, 90 by default; the console offers 30/90/180/365). The secret
`sluis_admin_<32 chars>` is shown once; only its first 16 characters are kept.
Tokens are minted by a signed-in session only — a token can never create, list or revoke
tokens. Write scopes can only be granted by an organisation owner (a write token acts as
owner on its family).

Authentication: `Authorization: Bearer sluis_admin_…`. A token is bound to the
organisation it was created in and reaches only the endpoints its scopes allow;
everything else answers `403 {"error":{"message":"token scope does not allow this
endpoint","type":"sluis_error"}}`. Scopes are `<family>:read` (GET/HEAD on the family)
and `<family>:write` (every method on it, implies read). Read-only families have no
write scope.

| Family     | Read (GET)                                                                                 | Write                                                                              |
|------------|--------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------|
| `usage`    | `/admin/v1/usage`, `/admin/usage`                                                          | —                                                                                  |
| `orgs`     | `/admin/orgs` (the bound org + its clients)                                                | `POST /admin/orgs`, `PUT /admin/orgs/{id}/budget`, `PUT /admin/orgs/{id}/margin`  |
| `members`  | `/admin/members`                                                                           | `POST /admin/members`, `PUT/DELETE /admin/members/{id}`, `POST …/{id}/invite`      |
| `keys`     | `/admin/keys*`, `/admin/workloads*`                                                        | create / update / rotate / move / revoke keys, manage workloads and their members  |
| `policy`   | `GET /admin/policy`                                                                        | —                                                                                  |
| `audit`    | `GET /admin/audit` (metadata only), `/admin/operator-audit`, `/admin/requests/live`        | —                                                                                  |
| `billing`  | `/admin/billing`, `/admin/billing/invoice/{id}`, `/admin/billing/statement/{y}/{m}`, `/admin/usage/org-spend` | — (payments are never scriptable)                           |

Session-only whatever the token carries: `/admin/api-tokens*`, `/admin/2fa/*`, provider /
BYOK credentials, edge gateways, `/admin/legal*`, `/admin/session/switch`,
`/admin/orgs/standalone`, `PUT /admin/policy` and `/admin/policy/versions*`, audit traces
and `/admin/traces*` (decrypted payloads), every `POST`/`PUT` under `/admin/billing`.

Usage export: `GET /admin/v1/usage?from=YYYY-MM-DD&to=YYYY-MM-DD` (UTC, half-open
`[from, to)`, ≤ 92 days; default = the previous calendar month). One entry per
organisation under the root, zero-usage clients included, each with per-provider/model
`lines[]`. `billable_cents` is what the root is charged for that organisation (integer
cents, excl. VAT), computed with the same query as the console billing page;
`list_cents` is provider list price (0 on BYOK). Root organisations only: a client-org
token gets `403 "usage export is root-only"`. One `organisations[]` entry = one customer
invoice; store `period.from`/`period.to` on it so a re-run of the month is idempotent.

Provisioning: with `orgs:write` a token creates client organisations (the owner who
minted the token becomes each new client's first owner, so the minter must still be an
owner of the organisation) and sets their budget (micro-euros) and margin (basis points).
`members:write` invites by email only, changes roles and removes members — a token cannot
grant the owner role, change an owner's role, set a password or remove an owner.
`keys:write` creates workloads and keys. Every mutation runs the same validation, role
and billing rules as the console (creating client orgs still requires active billing);
a token cannot grant the owner role, set passwords, or read request/response payloads.
The same ceiling holds for people: only a signed-in owner can grant the owner role, or
change or remove an owner; an admin session gets `403`. An agency admin who creates a
client organisation keeps only the admin access it inherits from the agency there and
does not become its owner.

Lifecycle: every console-minted token expires (`expires_in_days` 1–365, default 90); an
expired token answers `401` like a revoked one and is listed as `expired`. Revocation takes
effect within 30 seconds; tokens are also revoked automatically when the user who minted
them is removed, demoted below owner, or changes or resets their password. Rate limit:
600 requests per minute per token, over it `429` with `Retry-After`. Every mutation made
with a token lands in the operator audit (`GET /admin/operator-audit`, console → Audit log)
with `metadata.actor_token = {id, name}`.

Errors: `401` missing/invalid/revoked/expired token · `403` scope or role refusal · `429`
per-token rate limit (wait `Retry-After`) · `400` invalid body/query (the message names the
field). Scopes are fixed at creation: mint a new token rather than widening one.

Operator tokens minted server-side (`sluis-gateway create-admin-token --scopes …`) follow
the same rules: explicit scopes (or `all` for full owner access, operator-only) and an
expiry of at most 365 days. They are listed — flagged "full access" — and revocable under
Settings → Admin API like any other token.

## Self-hosted gateways — Sluis Edge (https://www.sluis.ai/docs/edge)

Sluis Edge is enterprise-only and requires a license: no gateway can be registered
until the organisation redeems a license key (from hello@sluis.ai) under
Settings → Gateways (https://sluis.ai/console/settings/gateways). Manage registered
gateways in the fleet view (https://sluis.ai/console/gateways).

Edge is the self-hosted data plane: one binary embedding the same router, providers,
and gate, with local state and an outbound-only control tunnel to the cloud console.
Keys, policy, and BYOK credentials sync down; usage and audit are read live from the
box and never stored in the cloud.

Prompts do not transit Sluis's cloud. External model-provider egress still depends
on your configured providers and policy; Edge alone does not make an external model
local. The enterprise licence is flat, per gateway, annual.

Install on any systemd host; configuration lives in `/etc/sluis-edge/env` and survives
re-runs and reboots.

```shell
curl -fsSL https://get.sluis.ai/edge/install.sh | sudo bash -s -- \
  --control-url wss://api.sluis.ai/edge/connect \
  --install-token sluis_edge_install_…

systemctl status sluis-edge     # service health
journalctl -u sluis-edge -f     # live logs
```

### Observability

For dedicated and self-hosted deployments the gateway exposes compliance-labelled
Prometheus metrics: requests, tokens, cost, and latency broken down by provider,
jurisdiction, and routing decision, plus cache hits and circuit-breaker state. Enable
with `SLUIS_METRICS_ENABLED` and scrape `/metrics` on a private network. Traces export
over OTLP (`SLUIS_OTLP_ENDPOINT`); every trace carries the request id that correlates
it one-to-one with its audit row.

```yaml
scrape_configs:
  - job_name: sluis-gateway
    metrics_path: /metrics
    static_configs:
      - targets: ["gateway.internal:8080"]
```

## Contact

- Sales & enterprise: hello@sluis.ai
- General: hello@sluis.ai
- Security: hello@sluis.ai

## Workloads

A workload owns shared runtime settings for a team, user, agent or application. API keys are credentials under the workload. Rotate credentials without changing its policy.

Configure requests/minute, tokens/minute, allowed models and a total, daily or monthly budget. Budget action block returns 402; notify allows calls to continue. Pausing refuses requests with 403 without revoking credentials.

An empty option_overrides object inherits organisation defaults. Overrides are validated by the gateway and cannot escape the organisation's allowed provider and residency scope. Spend is aggregated across the workload's keys in the current budget period.

Answer timeout (`upstream_timeout_secs`, Limits & budget): how long a model call waits for the provider's answer, 10 to 600 seconds, 300 by default. A call without an answer in time fails with 504 (`error.sluis.cause: timeout`) and is not replayed, because the provider may still be working on it. A caller can shorten it per request with the `x-sluis-upstream-timeout` header, never extend it. A streamed answer that has started is not cut by it; only 600 seconds of silence mid-stream end it.

Scope guard (`scope_guard_mode` off | log | block, `scope_guard_prompt` up to 4,000 characters, required when on): a small EU-hosted model (Mistral Small) checks every model call made with the workload's keys (embeddings excepted) against the scope you wrote, after data protection, on the text the provider would receive. Log judges in the background without delaying the call and records out-of-scope calls in the audit log; block waits for the verdict and refuses with `403 scope_violation`, or `503 scope_guard_unavailable` when the check cannot complete. Each check is a separate audit receipt billed at the model's token price; the organisation's residency policy must allow the model, or the guard cannot be switched on.

## Security and audit

Prompt-injection scanning supports off, log and block. Log records findings while allowing the request; block refuses detected attacks. Detection is not a guarantee and should be evaluated on your workload. The configured security classifier runs inside Sluis.

Key-behaviour anomaly detection runs in the background. Review findings in the audit trail. Audit evidence records request decisions; it is not proof that every sensitive value or attack was detected. Organisation policy and validated workload overrides determine effective controls.

The audit log shows one row per request; expand it to see the checks it ran (name recognition, LLM privacy inspection, the workload scope guard, the security scan). Each check remains its own sealed chain entry with its own cost.

Account security: sign in with email and password, or with Google, Apple or Microsoft (work or school accounts in Microsoft Entra ID). In Profile, add a passkey or an authenticator app (any TOTP app, such as Google Authenticator, Microsoft Authenticator, 1Password or Bitwarden) as your second factor, then generate one-time recovery codes. Owners and admins can require two-factor authentication for the organisation under Settings › Application security and choose which method counts: passkeys, email codes or an authenticator app; members who lack it must set it up before they can continue. A session started with Google sign-in is exempt from that requirement; Apple, Microsoft and password sign-ins are not. Sign-in events are sealed to the audit log.

## Performance

Internal benchmark, 2026-07-04. Identical development hardware; Node keep-alive mock, 60 ms delay, concurrency 256, 30 seconds, oha 1.14. Not a publicly reproducible harness or a production SLA.

| Metric | Sluis gate off | Sluis gate on | LiteLLM | Bifrost |
|---|---|---|---|---|
| Throughput RPS | 4046 | 3678 | 298 | 3148 |
| p50 ms | 62 | 63 | 743 | 80 |
| p99 ms | 77 | 130 | 4584 | 118 |
| Success | 100% | 100% | 100% | 100% |

Peak RSS was not measured. Gate-on overhead is compared with Sluis gate-off. Provider latency dominates live-model tails. See docs/benchmarks.md in the source repository for methodology.
