Skip to content

Budgets & caching

Workload rate limits and budgets, organisation caps, and the exact-match response cache.

Budgets & limits

Each workload owns rate limits (requests and tokens per minute) and a spend budget shared across its credentials. Choose a budget period: total, daily or monthly. An organisation-wide aggregate cap applies across all workloads.

Enforcement happens at the gateway, in real money: every request is priced from the live price book and debited before dispatch. Past the rate limit the call returns 429; past the budget, 402. The request never reaches a provider. If the counters are ever unreachable, the check fails closed and reconciles from the immutable audit ledger instead of guessing. Workspace Knowledge uploads are metered the same way: the embeddings a document needs are priced like a /v1/embeddings call and charged to the organisation, its prepaid credit and the uploading member's budget. An upload is refused with 402 when the organisation budget or credit is exhausted.

Usage: review consumption and spending.
Usage: review consumption and spending.English interface · illustrative demo data. Open the image for full size.
# past the key's rate limit — the request never reaches a provider
{
  "error": {
    "message": "rate limit exceeded",
    "type": "rate_limit_error",
    "param": null,
    "code": "rate_limit_exceeded"
  }
}

Caching

An opt-in exact response cache matches normalized requests. It is strictly isolated per organisation, encrypted at rest, and requires content retention. Enable it in the Console's Cache view.

Cache: configure exact response caching.
Cache: configure exact response caching.English interface · illustrative demo data. Open the image for full size.

New audit rows record how the call was served: none | exact.

# identical request twice — the second is served from the exact cache
curl https://api.sluis.ai/v1/chat/completions \
  -H "Authorization: Bearer $SLUIS_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "mistral/mistral-large-latest", "messages": [{ "role": "user", "content": "Define GDPR" }] }'

# audit row of call #2: cache_hit: exact · no provider dispatch, no token cost