MAQPNADocs

Scaling and high availability

Run several gateway, identity and attestation replicas on shared PostgreSQL state, size them from measured latency, and keep session start fast with warm pools.

A single gateway replica on the file state backend is fine for a lab. For production, run three gateway replicas on the PostgreSQL state backend with a PodDisruptionBudget, two identity broker and attestation replicas, and warm pools for the trust tiers that must start fast.

flowchart TB
  subgraph Agent namespaces
    S1[sandbox] & S2[sandbox] & S3[sandbox]
  end
  S1 & S2 & S3 --> SVC[maqpna-gateway Service]
  SVC --> G1[gateway 1] & G2[gateway 2] & G3[gateway 3]
  G1 & G2 & G3 --> PG[(PostgreSQL<br/>state + audit_chain)]
  G1 & G2 & G3 -.local verified mirror.-> L[(ledger PVC per replica)]
  OP[operator x2<br/>leader election] --> CM[ConfigMaps: policies,<br/>revocations, budgets]
  CM -.configSync 1 s.-> G1 & G2 & G3

Goal#

An installation that survives the loss of one node or one gateway replica with no lost approvals, revocations or audit records.

Prerequisites#

  • A highly available PostgreSQL in the same region and jurisdiction as the gateways (Patroni or CloudNativePG).
  • Enough nodes to spread replicas (set gateway.topologySpreadConstraints or affinity).

What scales how#

Component Replicas Requirement Value
Gateway 1 with file; 2 or more with postgres The chart refuses more than 1 with state.backend=file unless guards.allowIndependentGatewayReplicas=true (load tests only). gateway.replicas, gateway.pdb.minAvailable
Identity broker 1 with identity.key.mode=generate; 2 or more with helm or existingSecret Every replica must sign with the same key. identity.replicas
Attestation service 1 with file; 2 or more with postgres The release store must be shared. attestation.replicas, attestation.pdb
Operator 2 for failover Leader election; session metrics come from the leader. operator.replicas

The prod profile sets gateway 3 (PDB minAvailable: 2), identity 2, attestation 2 and operator 2.

Why the gateway needs PostgreSQL for more than one replica#

With the file backend, every replica keeps its own approvals, break-glass revocations, session taint, tool pins, budgets and audit ledger. maqpna kill would reach one replica and maqpna approvals list would show part of the queue. With state.backend: postgres:

  • an approval created on one replica is decided and consumed on any other, exactly once;
  • a revocation made on one replica is enforced on all within state.postgres.pollMillis (250 ms by default);
  • the audit ledger is one hash chain; every replica keeps a byte-identical, verified local mirror;
  • rate limits (tokensPerMinute, model-route requestsPerMinute) count cluster-wide (state.counters: shared);
  • SIEM and WORM shipping, the tenant ledger feed and retention sweeps run on one elected replica.

Steps#

1. Switch to the PostgreSQL state backend#

state:
  backend: postgres
  postgres:
    dsnSecret: maqpna-state
    maxConns: 10
    pollMillis: 250
gateway:
  replicas: 3
  pdb: {enabled: true, minAvailable: 2}
  auditStorage:
    persistent: true      # StatefulSet, one PVC per replica for the ledger mirror
    size: 50Gi
identity:
  replicas: 2
  key: {mode: existingSecret, existingSecret: maqpna-identity-key}
attestation:
  replicas: 2
operator:
  replicas: 2

The first replica that starts against an empty shared chain seeds it from its existing local ledger (verified, in one transaction), so moving from file to postgres keeps the evidence.

maqpna values validate --chart ./maqpna -f my-values.yaml --profile production
maqpna upgrade check -f my-values.yaml
maqpna upgrade -f my-values.yaml --wait --atomic

2. Keep revocations fast#

gateway.configSync.enabled (on by default) reads the operator's ConfigMaps (policies, revocations, budgets, …) from the Kubernetes API every gateway.configSync.pollMillis (1,000 ms) instead of waiting for the kubelet's 60–90 s volume refresh. Keep it on: it is what makes an AgentRevocation effective within seconds.

3. Add warm pools for fast session start#

A trust tier with a warm pool keeps pre-started sandboxes, so a session claims one instead of creating a sandbox.

apiVersion: maqpna.com/v1alpha1
kind: TrustTier
metadata:
  name: tier-0
spec:
  runtimeClassName: gvisor
  isolation: gvisor
  egress: gateway-only
  sandboxTemplateName: agent-default    # an agent-sandbox SandboxTemplate
  warmPoolSize: 5
  warmPoolNamespaces: [team-a]

Watch maqpna_warm_pool_ready{tier} and maqpna_session_ready_seconds{mode="claim"}. The start-latency targets are p90 under 1 s warm and under 5 s cold (MaqpnaSessionReadySLO). agent-sandbox treats a claim with environment variables as a cold start, and MAQPNA passes the session token as claim environment, so measure warm starts on your agent-sandbox version before you rely on the warm target.

4. Cap concurrency and spend#

  • Agent.spec.maxConcurrentSessions and budgets.namespaces["*"].maxConcurrentSessions (enforced by the operator).
  • maxCallsPerMinute on policy rules (per session).
  • Budgets per agent, namespace and user (Budgets and cost limits).

Sizing from measured results#

These figures come from a load test of the gateway on Fly.io Firecracker microVMs (October 2026), with auditFailurePolicy: closed, a default DLP profile, 64 sessions and a call mix of 80% allow, 5% DLP redaction, 10% deny and 5% approvals. They are measurements of one environment, not guarantees.

Configuration Load Gateway-added latency p50 / p95 / p99
File ledger, auditFailurePolicy: open, 4 shared vCPU 2,000 calls/s 2.37 / 3.65 / 2.15 ms
File ledger, auditFailurePolicy: closed, 4 shared vCPU 2,000 calls/s 4.01 / 8.64 / 9.80 ms
PostgreSQL state, closed, 2 dedicated vCPU (gateway, database, load) 2,000 calls/s 14.5 / 25.5 / 35.4 ms

What the results mean for sizing:

  • Closed audit mode costs one synchronous fsync per allowed call (the intent record): about 1.7 ms p50 at 2,000 calls/s on the file ledger.
  • PostgreSQL state adds two synchronous commits per allowed call in closed mode, so the lowest achievable in-region latency is about 7 ms p50. Choose it for availability, not for latency.
  • Keep active replicas in the database's region. A replica writing across regions measured a p50 of 196 ms; use remote replicas for failover only.
  • CPU per call is about 0.47 ms, so one vCPU serves roughly 2,000 calls/s before other work.
  • In a failover test (two replicas, one killed mid-run at 300 calls/s), 0.13% of calls failed at the moment of the kill, and no approvals and no audit records of answered calls were lost.
  • Revocations reached the other replica in a median of 195 ms (p99 302 ms), in line with pollMillis: 250.

Verification#

kubectl -n maqpna-system get pdb
maqpna status
maqpna audit verify --gateway "$GW"     # same head on every replica
maqpna doctor --component state,gateway

In Prometheus, maqpna_gateway_leader{job} must be 1 on exactly one replica per job.

Troubleshooting#

Symptom Cause Fix
doctor: gateway-replicas-state fails More than one replica on the file backend Switch to postgres, or scale to 1.
A replica refuses to start: its local ledger is not a prefix of the shared chain A stale or foreign volume (for example a second replica's own file ledger), or rewritten database rows Inspect first (Monitoring and alerts). Then set auditAcceptDivergedMirror: true in the gateway config once; the replica moves its mirror aside to <auditPath>.diverged-<unix> and rebuilds it.
Two replicas show different approvals list File backend with independent replicas Use postgres.

Next steps#