MAQPNADocs

Model call governance#

A sandbox never learns a real model URL or API key. Its model endpoint is MAQPNA_MODEL_ENDPOINT={gateway}/llm/<route>/v1, and its API key is its own session token. Every model call goes through the gateway's OpenAI-compatible proxy, which checks identity, the kill switch, scope, budgets and token limits, runs data loss prevention (DLP) on the prompt and the completion, picks a compliant endpoint, meters tokens and writes one audit record.

The proxy speaks the OpenAI wire format and forwards to OpenAI-compatible upstreams (vLLM, llm-d, KServe or a hosted API that offers the same format). There are no provider-specific adapters.

Routes#

Path Operation (the policy "tool")
POST /llm/{route}/v1/chat/completions chat.completions
POST /llm/{route}/v1/completions completions
POST /llm/{route}/v1/embeddings embeddings
GET /llm/{route}/v1/models models

The operator renders one route per Agent with spec.model.endpoint into the ConfigMap maqpna-models (key models.json). The route name is <namespace>.<agent> (sanitised, at most 63 characters). Each route carries the upstream base URL, the forced model, the namespaces and agents allow-lists and, when set, jurisdiction, requestsPerMinute, fallbacks[] and preferJurisdiction. Model API keys from spec.model.apiKeySecretRef are copied into the gateway-only Secret maqpna-model-credentials; key material never enters models.json. Agents that break the sovereignty policy in enforce mode get no route. The session's token carries models:<route>, granted automatically when the agent has a model.

The pipeline#

sequenceDiagram
    autonumber
    participant Ag as Agent (OpenAI SDK)
    participant GW as Gateway /llm
    participant M as Model endpoint
    participant L as Audit ledger
    Ag->>GW: POST /llm/team-a.coder/v1/chat/completions<br/>Authorization Bearer session token
    GW->>GW: verify token, kill switch, route exists
    GW->>GW: rewrite body: force model,<br/>stream_options.include_usage = true
    GW->>GW: scope models:route, namespace and agent allow-lists,<br/>policy (modelPolicy policy only), budgets
    GW->>GW: clamp max_tokens, reserve tokensPerMinute
    GW->>GW: DLP on the request, injection guards
    opt auditFailurePolicy closed
      GW->>L: intent record
    end
    GW->>M: pick endpoint: residency, data class, jurisdiction order,<br/>circuits, requestsPerMinute
    M-->>GW: completion (JSON or SSE)
    GW->>GW: DLP on the completion (streaming carry window)
    GW-->>Ag: completion
    GW->>GW: meter prompt and completion tokens, settle the reservation
    GW->>L: record server llm/route, tool chat.completions, decision allow

Step by step (cmd/maqpna-gateway/llm.go, budgets.go, model_routing.go):

  1. Authenticate. The bearer token is verified like any session token. MCP-client tokens from a customer identity provider are refused here (401 external_not_allowed).
  2. Kill switch. A matching revocation answers 403, type maqpna_revoked.
  3. Route. An unknown route answers 404 route_not_found.
  4. Rewrite. The gateway forces the route's model into the body and sets stream_options.include_usage=true on streaming requests, even if the caller set it to false, because metering and budgets depend on usage. It strips Authorization, cookies, Api-Key, X-Api-Key, OpenAI-Organization, OpenAI-Project and client-supplied X-Maqpna-* headers.
  5. Authorise. The token must hold models:<route> (or models:*); the route's namespaces and agents allow-lists must admit the caller. With modelPolicy: "policy" the policy engine also decides, with server llm/<route> and the operation as tool. The default scope-only mode does not evaluate policies for model calls. A require_approval decision is a denial: model calls never wait for approval.
  6. Budgets. The legacy sessionBudgetUSD and every hierarchical budget (agent, namespace, user; session, day, month) are checked. Spent budgets answer 402 maqpna_budget_exceeded.
  7. Token limits. maxTokensPerRequest clamps (or sets) max_tokens. tokensPerMinute reserves an estimate (body bytes ÷ 4) per session and route, in shared Postgres windows or a local token bucket; when exhausted the call gets 429 maqpna_rate_limited with Retry-After.
  8. DLP and guards. The whole request body except the model name and numeric controls is scanned with the route's profile. A deny answers 403 maqpna_dlp_blocked. Injection guards may flag the prompt.
  9. Intent record. With auditFailurePolicy: closed, a durable intent record is written first; if it fails the call is refused (503, audit_unavailable).
  10. Route and fail over. See below.
  11. Response DLP. JSON completions are scanned in full (message, tool-call arguments, refusal, reasoning; logprobs is dropped from a redacted choice). Streamed completions pass each choice through a redactor that holds back the last streamCarryBytes (256 by default) so matches split across chunks are caught. A streamed deny emits a maqpna_dlp_blocked error event and stops content, but still forwards the final usage chunk.
  12. Meter and audit. Prompt and completion tokens are priced from the pricing table and added to the session's spend (so model spend also blocks tool calls once a budget is exhausted). One ledger record is written: server=llm/<route>, tool=<operation>, argsSha256 of the raw request body (prompts are never stored), decision, latencyMs, costUsd, and a reason with usage:prompt=N,completion=M, model_override:a->b and the upstream status when relevant.

Routing and failover#

flowchart TD
    S["Endpoints: primary + fallbacks[]"] --> R1{"Passes the residency check?"}
    R1 -- no --> X1["dropped"]
    R1 -- yes --> R2{"Session touched a data class<br/>with a dataClassModels list?"}
    R2 -- "yes, model not listed" --> X2["dropped"]
    R2 -- "no, or listed" --> O["Order: tier jurisdiction first,<br/>then route preferJurisdiction,<br/>then sovereignty home, else declared order"]
    O --> C{"Circuit open or<br/>requestsPerMinute spent?"}
    C -- yes --> NEXT["skip to the next endpoint"]
    C -- no --> TRY["send"]
    TRY --> OK{"5xx or 429?"}
    OK -- no --> DONE["serve, ext.modelEndpoint"]
    OK -- yes --> NEXT
    NEXT --> C

A circuit opens after failureThreshold consecutive transport errors, timeouts, 5xx or 429 answers, stays open for cooldownSeconds, then allows one half-open trial. Every skip and failover is recorded on the call's ledger record (ext.modelEndpoint, ext.modelFailover) and counted in maqpna_gateway_model_failovers_total. When the session's data class rules out every endpoint, the call is denied with model_not_allowed_for_data_class. All endpoint connections go through the residency-checking dialer.

Errors#

Model errors use the OpenAI shape, plus the MAQPNA reason and domain (keys are sorted; message format from cmd/maqpna-gateway/budgets.go, values illustrative):

{"error":{"code":"budget_exceeded","domain":"maqpna.com","message":"namespace budget (day) exhausted for team-a: spent 200.000000 of 200.000000 USD","reason":"budget_exceeded","type":"maqpna_budget_exceeded"}}
HTTP type When
401 maqpna_unauthorized Missing, expired or invalid token
403 maqpna_unauthorized Missing models: scope, namespace or agent not allowed
403 maqpna_revoked Kill switch
403 maqpna_policy_denied Policy deny (modelPolicy: policy)
403 maqpna_approval_required Policy asked for approval (not supported for models)
403 maqpna_dlp_blocked DLP deny on the prompt or completion
403 maqpna_egress_denied The residency dialer refused the endpoint
402 maqpna_budget_exceeded A budget is spent
404 maqpna_not_found Unknown route
429 maqpna_rate_limited tokensPerMinute exhausted (code tokens_per_minute)
502 maqpna_upstream_error No endpoint answered
503 maqpna_audit_unavailable Closed audit mode and the intent record failed

What you see#

Under maqpna dev up --stub-llm, the model route stub is served by a scripted local model, and the agent's OpenAI SDK picks up OPENAI_BASE_URL and OPENAI_API_KEY (the session token). The timeline shows model calls next to tool calls (format from cmd/maqpna/dev.go; values illustrative):

TIME      KIND        SERVER/TOOL               DECISION  DETAIL
14:05:09  model_call  llm/stub/chat.completions  allow     usage:prompt=212,completion=38
14:05:09  tool_call   echo/echo                  allow     matched rule read-only-tools

Metrics: maqpna_gateway_model_requests_total{route,status}, maqpna_gateway_model_tokens_total{route,kind}, gen_ai_client_token_usage, maqpna_gateway_model_circuits_open; the chart alert MaqpnaModelCircuitOpen fires while a circuit is open. See maqpna dev timeline and maqpna budgets list.

Failure modes#

Failure Effect
Every endpoint down or open 502 maqpna_upstream_error, audited as a deny
Endpoint API key Secret missing That endpoint fails closed and the gateway fails over; the agent condition ModelCredentialsReady=False names the Secret
Shared token counters unreachable The gateway falls back to its local limiter; the request path never fails on a counter
Ledger unavailable in closed mode 503 maqpna_audit_unavailable