MAQPNADocs

HA and failure handling#

MAQPNA's rule for failures is simple to state: if a control cannot be checked, the call is denied — with a short, explicit list of exceptions where availability was chosen on purpose and the choice is configurable. This page lists every component's replica model and every dependency's failure behaviour, as implemented.

Replicas#

Component Chart default Production profile (values-production.yaml) How replicas coordinate
Gateway 1 3 state.backend: postgres shares approvals, revocations, taint, pins, outbox, meter, vault, memory, captures and the audit chain. With file and more than one replica the chart refuses to render unless guards.allowIndependentGatewayReplicas is set
Identity broker 1 2 Stateless; more than one replica needs a shared key (identity.key.mode: helm or existingSecret)
Attestation service 1 (when enabled) 2 Requires state.backend: postgres for more than one replica (releases and challenges claimed under a cluster-wide lock)
Operator 1 2 Kubernetes Lease maqpna-operator.maqpna.com; the chart adds --leader-elect only when operator.replicas > 1

PodDisruptionBudgets. The gateway PDB (minAvailable: 1) is rendered whenever gateway.pdb.enabled is true, which is the default, whatever the replica count. With the default single replica this blocks voluntary evictions such as node drains; run at least two replicas, or disable the PDB, before draining nodes. The attestation PDB is rendered only with more than one replica. The identity broker and operator have no PDB.

With more than one gateway replica, the chart adds a preferred pod anti-affinity so replicas spread across nodes.

Background jobs and leaders#

Gateway jobs that must run once per installation are guarded by PostgreSQL advisory-lock leases (see state backend schema): siem/<sink>, worm/<kind>, tenant-ledgers, memory-sweep, capture-sweep. Shipping positions are shared blobs, so a new leader resumes where the old one stopped. Delivery is at least once at the edges; receivers can de-duplicate on the record hash.

sequenceDiagram
    participant Agent
    participant LB as Gateway Service
    participant A as Replica A
    participant B as Replica B
    participant PG as PostgreSQL
    Agent->>LB: tools/call (held for approval, sync)
    LB->>A: route to A
    A->>PG: approvals journal: create apr_1
    Note over A: A is killed
    PG-->>PG: A's lease session ends, its locks are released
    B->>PG: takes leases siem/ocsf and worm/audit
    Agent->>LB: retry with X-Maqpna-Async: 1 (the TCP connection was lost)
    LB->>B: route to B
    B->>PG: create apr_2 and return -32002
    Note over B: an approver approves apr_2 on any replica
    Agent->>LB: retry with X-Maqpna-Approval-Id: apr_2
    LB->>B: consume apr_2 under the journal lock, then forward
    B->>PG: append to audit_chain (one chain, advisory lock)

Walkthrough:

  1. A replica dies. Kubernetes removes it from the Service; in-flight requests on it fail and agents retry them (tool calls are retried by the agent, and the SDKs surface the error).
  2. Its PostgreSQL sessions end, releasing its leases at once (or within about 8 seconds if the host vanished without a TCP reset).
  3. A standby takes each lease at its next retry (leader.retryMillis, 1 second) and resumes shipping from the saved cursor.
  4. Pending approvals, revocations and taint are in the shared journals, so any surviving replica serves them. A caller blocked on a synchronous approval on the dead replica lost its connection and must retry; the approval itself is not lost.
  5. The ledger is one chain: every replica appends under the same advisory lock and keeps a verified local mirror.

Failure modes#

What fails Behaviour Fails
PostgreSQL unreachable /readyz reports state backend unreachable and the pod leaves the Service. Creating, deciding and consuming approvals fail; break-glass changes return 500; audit appends fail (and count in maqpna_gateway_audit_errors_total); the attestation service refuses with reason store (renewals 503). Reads serve the last folded state Closed
Shared rate-limit counters unreachable The replica falls back to its local limiter; the request path never fails on a counter Open (limits per replica)
Revocation list missing or invalid Every /mcp, /llm, /a2a and egress call is denied with revocation_list_unavailable; /readyz 503; maqpna_gateway_revocation_list_ok 0 Closed
Audit ledger append fails, auditFailurePolicy: open (default) The call proceeds; the error is logged and counted Open (configurable)
Audit ledger append fails, auditFailurePolicy: closed No intent record means the call is refused (audit_unavailable); a failed completion record withholds a held result; /readyz 503 Closed
Policy file invalid on reload The previous policies stay in force. At start-up with no valid policies, /readyz reports policies not loaded and no call is allowed Closed
Unknown DLP profile dlp:profile_unknown deny Closed
Cedar text in a binary without -tags cedar, or a Cedar error cedar_unavailable / cedar_error deny Closed
OPA sidecar down or slow rego_unavailable deny Closed
Injection guard classifier down failOpen defaults to true: the call proceeds; set failOpen: false for guard_unavailable denials Open (configurable)
MCP server or model endpoint down upstream_unavailable (502); model routes fail over to the next endpoint, and open a circuit after repeated failures —
Upstream credential missing upstream_credential_unavailable; nothing is sent upstream Closed
Identity broker down New non-attested sessions cannot get a token: MintFailed on the session and the operator retries with backoff. Warm-pool bootstraps wait. Existing tokens keep working until they expire, and the gateway keeps its cached JWKS Closed for new sessions
Attestation service down The release returns to pending (verifier 503, broker 502) and the init container retries for about 60 seconds; the operator treats outages as "keep waiting", not as failure Closed (no token released)
Admin identity provider down OIDC tokens already issued keep verifying with cached keys until they expire; new logins fail. With adminAuth.mode: both the static break-glass token still works and every use is recorded —
Operator down Running sessions keep running and the gateway keeps enforcing the last rendered files. New sessions, policy changes and new revocations wait for the operator; break-glass revocations at the gateway still work —
Gateway down Sandboxes can reach nothing else (NetworkPolicy), so every agent call fails Closed
Licence missing, invalid or expired No effect on any call, session, approval or audit write Never blocks

What you see#

maqpna doctor and GET /v1/posture evaluate these settings as posture checks. A gateway running three replicas on the file backend fails the state-backend check (text from pkg/posture/gateway.go):

{
  "id": "state-backend",
  "component": "state",
  "title": "Shared state backend for HA replicas",
  "severity": "fail",
  "status": "fail",
  "detail": "stateBackend.type=file with 3 gateway replicas: approvals, revocations, taint, budgets and audit ledgers diverge per replica",
  "fix": "set state.backend: postgres (docs/state-backend.md) or run one gateway replica"
}

Each check also carries a source, and the report has a score from 0 to 100; maqpna doctor exits 3 when a check fails. See maqpna doctor and maqpna status.