Model call governance#
A sandbox never learns a real model URL or API key. Its model endpoint is MAQPNA_MODEL_ENDPOINT={gateway}/llm/<route>/v1, and its API key is its own session token. Every model call goes through the gateway's OpenAI-compatible proxy, which checks identity, the kill switch, scope, budgets and token limits, runs data loss prevention (DLP) on the prompt and the completion, picks a compliant endpoint, meters tokens and writes one audit record.
The proxy speaks the OpenAI wire format and forwards to OpenAI-compatible upstreams (vLLM, llm-d, KServe or a hosted API that offers the same format). There are no provider-specific adapters.
Routes#
| Path | Operation (the policy "tool") |
|---|---|
POST /llm/{route}/v1/chat/completions |
chat.completions |
POST /llm/{route}/v1/completions |
completions |
POST /llm/{route}/v1/embeddings |
embeddings |
GET /llm/{route}/v1/models |
models |
The operator renders one route per Agent with spec.model.endpoint into the ConfigMap maqpna-models (key models.json). The route name is <namespace>.<agent> (sanitised, at most 63 characters). Each route carries the upstream base URL, the forced model, the namespaces and agents allow-lists and, when set, jurisdiction, requestsPerMinute, fallbacks[] and preferJurisdiction. Model API keys from spec.model.apiKeySecretRef are copied into the gateway-only Secret maqpna-model-credentials; key material never enters models.json. Agents that break the sovereignty policy in enforce mode get no route. The session's token carries models:<route>, granted automatically when the agent has a model.
The pipeline#
sequenceDiagram
autonumber
participant Ag as Agent (OpenAI SDK)
participant GW as Gateway /llm
participant M as Model endpoint
participant L as Audit ledger
Ag->>GW: POST /llm/team-a.coder/v1/chat/completions<br/>Authorization Bearer session token
GW->>GW: verify token, kill switch, route exists
GW->>GW: rewrite body: force model,<br/>stream_options.include_usage = true
GW->>GW: scope models:route, namespace and agent allow-lists,<br/>policy (modelPolicy policy only), budgets
GW->>GW: clamp max_tokens, reserve tokensPerMinute
GW->>GW: DLP on the request, injection guards
opt auditFailurePolicy closed
GW->>L: intent record
end
GW->>M: pick endpoint: residency, data class, jurisdiction order,<br/>circuits, requestsPerMinute
M-->>GW: completion (JSON or SSE)
GW->>GW: DLP on the completion (streaming carry window)
GW-->>Ag: completion
GW->>GW: meter prompt and completion tokens, settle the reservation
GW->>L: record server llm/route, tool chat.completions, decision allow
Step by step (cmd/maqpna-gateway/llm.go, budgets.go, model_routing.go):
- Authenticate. The bearer token is verified like any session token. MCP-client tokens from a customer identity provider are refused here (
401 external_not_allowed). - Kill switch. A matching revocation answers
403, typemaqpna_revoked. - Route. An unknown route answers
404 route_not_found. - Rewrite. The gateway forces the route's
modelinto the body and setsstream_options.include_usage=trueon streaming requests, even if the caller set it to false, because metering and budgets depend on usage. It stripsAuthorization, cookies,Api-Key,X-Api-Key,OpenAI-Organization,OpenAI-Projectand client-suppliedX-Maqpna-*headers. - Authorise. The token must hold
models:<route>(ormodels:*); the route'snamespacesandagentsallow-lists must admit the caller. WithmodelPolicy: "policy"the policy engine also decides, with serverllm/<route>and the operation as tool. The defaultscope-onlymode does not evaluate policies for model calls. Arequire_approvaldecision is a denial: model calls never wait for approval. - Budgets. The legacy
sessionBudgetUSDand every hierarchical budget (agent, namespace, user; session, day, month) are checked. Spent budgets answer402 maqpna_budget_exceeded. - Token limits.
maxTokensPerRequestclamps (or sets)max_tokens.tokensPerMinutereserves an estimate (body bytes ÷ 4) per session and route, in shared Postgres windows or a local token bucket; when exhausted the call gets429 maqpna_rate_limitedwithRetry-After. - DLP and guards. The whole request body except the model name and numeric controls is scanned with the route's profile. A deny answers
403 maqpna_dlp_blocked. Injection guards may flag the prompt. - Intent record. With
auditFailurePolicy: closed, a durable intent record is written first; if it fails the call is refused (503,audit_unavailable). - Route and fail over. See below.
- Response DLP. JSON completions are scanned in full (message, tool-call arguments, refusal, reasoning;
logprobsis dropped from a redacted choice). Streamed completions pass each choice through a redactor that holds back the laststreamCarryBytes(256 by default) so matches split across chunks are caught. A streamed deny emits amaqpna_dlp_blockederror event and stops content, but still forwards the final usage chunk. - Meter and audit. Prompt and completion tokens are priced from the pricing table and added to the session's spend (so model spend also blocks tool calls once a budget is exhausted). One ledger record is written:
server=llm/<route>,tool=<operation>,argsSha256of the raw request body (prompts are never stored),decision,latencyMs,costUsd, and areasonwithusage:prompt=N,completion=M,model_override:a->band the upstream status when relevant.
Routing and failover#
flowchart TD
S["Endpoints: primary + fallbacks[]"] --> R1{"Passes the residency check?"}
R1 -- no --> X1["dropped"]
R1 -- yes --> R2{"Session touched a data class<br/>with a dataClassModels list?"}
R2 -- "yes, model not listed" --> X2["dropped"]
R2 -- "no, or listed" --> O["Order: tier jurisdiction first,<br/>then route preferJurisdiction,<br/>then sovereignty home, else declared order"]
O --> C{"Circuit open or<br/>requestsPerMinute spent?"}
C -- yes --> NEXT["skip to the next endpoint"]
C -- no --> TRY["send"]
TRY --> OK{"5xx or 429?"}
OK -- no --> DONE["serve, ext.modelEndpoint"]
OK -- yes --> NEXT
NEXT --> C
A circuit opens after failureThreshold consecutive transport errors, timeouts, 5xx or 429 answers, stays open for cooldownSeconds, then allows one half-open trial. Every skip and failover is recorded on the call's ledger record (ext.modelEndpoint, ext.modelFailover) and counted in maqpna_gateway_model_failovers_total. When the session's data class rules out every endpoint, the call is denied with model_not_allowed_for_data_class. All endpoint connections go through the residency-checking dialer.
Errors#
Model errors use the OpenAI shape, plus the MAQPNA reason and domain (keys are sorted; message format from cmd/maqpna-gateway/budgets.go, values illustrative):
{"error":{"code":"budget_exceeded","domain":"maqpna.com","message":"namespace budget (day) exhausted for team-a: spent 200.000000 of 200.000000 USD","reason":"budget_exceeded","type":"maqpna_budget_exceeded"}}
| HTTP | type |
When |
|---|---|---|
| 401 | maqpna_unauthorized |
Missing, expired or invalid token |
| 403 | maqpna_unauthorized |
Missing models: scope, namespace or agent not allowed |
| 403 | maqpna_revoked |
Kill switch |
| 403 | maqpna_policy_denied |
Policy deny (modelPolicy: policy) |
| 403 | maqpna_approval_required |
Policy asked for approval (not supported for models) |
| 403 | maqpna_dlp_blocked |
DLP deny on the prompt or completion |
| 403 | maqpna_egress_denied |
The residency dialer refused the endpoint |
| 402 | maqpna_budget_exceeded |
A budget is spent |
| 404 | maqpna_not_found |
Unknown route |
| 429 | maqpna_rate_limited |
tokensPerMinute exhausted (code tokens_per_minute) |
| 502 | maqpna_upstream_error |
No endpoint answered |
| 503 | maqpna_audit_unavailable |
Closed audit mode and the intent record failed |
What you see#
Under maqpna dev up --stub-llm, the model route stub is served by a scripted local model, and the agent's OpenAI SDK picks up OPENAI_BASE_URL and OPENAI_API_KEY (the session token). The timeline shows model calls next to tool calls (format from cmd/maqpna/dev.go; values illustrative):
TIME KIND SERVER/TOOL DECISION DETAIL
14:05:09 model_call llm/stub/chat.completions allow usage:prompt=212,completion=38
14:05:09 tool_call echo/echo allow matched rule read-only-tools
Metrics: maqpna_gateway_model_requests_total{route,status}, maqpna_gateway_model_tokens_total{route,kind}, gen_ai_client_token_usage, maqpna_gateway_model_circuits_open; the chart alert MaqpnaModelCircuitOpen fires while a circuit is open. See maqpna dev timeline and maqpna budgets list.
Failure modes#
| Failure | Effect |
|---|---|
| Every endpoint down or open | 502 maqpna_upstream_error, audited as a deny |
| Endpoint API key Secret missing | That endpoint fails closed and the gateway fails over; the agent condition ModelCredentialsReady=False names the Secret |
| Shared token counters unreachable | The gateway falls back to its local limiter; the request path never fails on a counter |
| Ledger unavailable in closed mode | 503 maqpna_audit_unavailable |