Budgets & caching
Workload rate limits and budgets, organisation caps, and the exact-match response cache.
Budgets & limits
Each workload owns rate limits (requests and tokens per minute) and a spend budget shared across its credentials. Choose a budget period: total, daily or monthly. An organisation-wide aggregate cap applies across all workloads.
Enforcement happens at the gateway, in real money: every request is priced from the live price book and debited before dispatch. Past the rate limit the call returns 429; past the budget, 402. The request never reaches a provider. If the counters are ever unreachable, the check fails closed and reconciles from the immutable audit ledger instead of guessing. Workspace Knowledge uploads are metered the same way: the embeddings a document needs are priced like a /v1/embeddings call and charged to the organisation, its prepaid credit and the uploading member's budget. An upload is refused with 402 when the organisation budget or credit is exhausted.

# past the key's rate limit — the request never reaches a provider { "error": { "message": "rate limit exceeded", "type": "rate_limit_error", "param": null, "code": "rate_limit_exceeded" } }
# budget reached or no active plan — activate a plan or raise the budget { "error": { "message": "budget exceeded", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } }
Caching
An opt-in exact response cache matches normalized requests. It is strictly isolated per organisation, encrypted at rest, and requires content retention. Enable it in the Console's Cache view.

New audit rows record how the call was served: none | exact.
# identical request twice — the second is served from the exact cache curl https://api.sluis.ai/v1/chat/completions \ -H "Authorization: Bearer $SLUIS_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "mistral/mistral-large-latest", "messages": [{ "role": "user", "content": "Define GDPR" }] }' # audit row of call #2: cache_hit: exact · no provider dispatch, no token cost