MAQPNADocs

Production install with Helm

Install MAQPNA on your own Kubernetes cluster with the Helm chart, from preflight checks to a smoke-tested, OIDC-protected installation.

This page takes you from an empty cluster to a production installation of MAQPNA: the operator, the gateway, the identity broker and (optionally) the attestation service, with shared PostgreSQL state, OIDC sign-in for the console and approvers, and a smoke test at the end.

flowchart LR
  A[Prerequisites<br/>agent-sandbox, runtimes, CNI] --> B[Secrets<br/>state DSN, keys, WORM]
  B --> C[Values<br/>profile + your overrides]
  C --> D[maqpna values validate]
  D --> E[maqpna preflight]
  E --> F[maqpna install --wait]
  F --> G[maqpna smoke<br/>helm test]
  G --> H[maqpna doctor]

Goal#

A running installation that passes maqpna doctor with no failed checks, built from the prod profile (values-production.yaml) or the sovereign-eu profile (values-sovereign-eu.yaml).

Prerequisites#

You need Why
Kubernetes 1.29 or later (1.30 or later for the sovereign profile's admission policy) The chart declares kubeVersion: ">=1.29.0-0".
Cluster-admin rights for the install The chart installs CRDs, ClusterRoles and RuntimeClasses.
The upstream agent-sandbox controller (v1.0.x) MAQPNA creates sandboxes through kubernetes-sigs/agent-sandbox; the chart does not install it.
The node runtimes for the trust tiers you enable tier-0 (gVisor, runsc), tier-1 (microVM, Kata Containers with Firecracker, needs /dev/kvm), tier-2 (confidential VM, Kata CoCo with AMD SEV-SNP or Intel TDX).
A CNI that enforces NetworkPolicy Sandboxes are default-deny; only DNS and the gateway are reachable. Calico, Cilium, OVN-Kubernetes and the EKS VPC CNI policy agent all work.
PostgreSQL in the same jurisdiction (for more than one gateway replica) The shared state backend. See the state backend.
The maqpna CLI and its maqpna-install plugin, same release preflight, install, upgrade, rollback and uninstall run the plugin.
An OIDC identity provider (Keycloak, Entra ID, Okta) Console sign-in and approver identity in production.

Install the agent-sandbox controller first:

kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v1.0.4/sandbox-with-extensions.yaml

Install the CLI and the plugin if you have not:

curl -fsSL https://maqpna.com/install.sh | sh                          # installs maqpna and maqpna-install
curl -fsSL https://maqpna.com/install.sh | sh -s -- --bin maqpna-install  # the plugin alone

Install profiles#

The chart ships three values files. --profile selects one; your -f files and --set flags override it.

--profile File in the chart Use it for Notable settings
dev values-dev.yaml kind, k3s, minikube Only the gVisor RuntimeClass, the sample attestation verifier (DEV ONLY), a generated identity key (DEV ONLY), sovereignty in audit mode, the echo test MCP server.
prod values-production.yaml Production adminAuth.mode: oidc, gateway.auditFailurePolicy: closed, state.backend: postgres, 3 gateway replicas, WORM audit shipping and signed checkpoints, a default DLP profile, adminListen: ":9090", Trustee attestation with signed responses, principalBinding.mode: requester, identity key from your own Secret, no echo server.
sovereign-eu values-sovereign-eu.yaml Disconnected EU installations Everything in prod, plus an in-country image mirror (airgap.enabled), an allow-list of registries and egress hosts, and PostgreSQL state.

Every value marked CHANGE-ME in the prod and sovereign-eu files must point at your systems before you install. Pull the chart to read them (REGISTRY is the release registry listed on maqpna.com/download):

helm pull "oci://$REGISTRY/charts/maqpna" --version 0.1.1 --untar
grep -n CHANGE-ME maqpna/values-production.yaml

Steps#

1. Create the Secrets the profile expects#

The prod profile reads keys and credentials from Secrets you create. The names below are the profile's defaults; change them in your values if you use others.

kubectl create namespace maqpna-system

# Shared state: the PostgreSQL DSN (never inline in values)
kubectl -n maqpna-system create secret generic maqpna-state \
  --from-literal=dsn='postgres://maqpna@pg-rw.db.svc:5432/maqpna?sslmode=verify-full'

# Identity signing key, synced from your HSM or KMS (PKCS#8 Ed25519 PEM)
kubectl -n maqpna-system create secret generic maqpna-identity-key --from-file=key.pem

# Ed25519 key that signs audit checkpoints
maqpna keygen -out ./audit
kubectl -n maqpna-system create secret generic maqpna-audit-signing \
  --from-file=audit-signing.pem=./audit/identity.key

# Attestation response signing key (tier-2 only)
maqpna keygen -out ./attest
kubectl -n maqpna-system create secret generic maqpna-attest-response-signing \
  --from-file=key.pem=./attest/identity.key --from-file=pub.pem=./attest/identity.pub

# WORM bucket credentials and the Trustee token keys
kubectl -n maqpna-system create secret generic maqpna-audit-s3 \
  --from-literal=access-key=AKIA... --from-literal=secret-key=...
kubectl -n maqpna-system create secret generic maqpna-trustee-jwks --from-file=jwks.json

2. Write your values file#

Keep your overrides in one file on top of the profile. The keys you set most often:

Key Default (values.yaml) What to set in production
image.digest "" Pin the digest you verified (Verify releases).
identity.key.mode generate (DEV ONLY) existingSecret with identity.key.existingSecret, or helm.
identity.trustDomain maqpna.local Your SPIFFE trust domain, for example eu.prod.example.org.
state.backend file postgres, with state.postgres.dsnSecret.
gateway.replicas 1 3 with postgres (the chart refuses more than 1 with file).
gateway.auditFailurePolicy open closed: a call is refused when its audit record cannot be written.
gateway.auditStorage.persistent false true: the gateway becomes a StatefulSet with one PVC per replica for the local ledger mirror.
gateway.config.adminListen unset ":9090": the admin API, /metrics and pprof leave the agent-facing port.
audit.sink / audit.worm.* stdout worm with an S3-compatible bucket that has Object Lock.
audit.checkpointSigning.enabled false true with secretName.
adminAuth.mode token oidc (or both to keep the static token as an audited break-glass credential).
adminAuth.roles / namespaceRoles groups maqpna-* Map your identity provider's groups to viewer, approver, auditor, admin, killswitch.
approvals.preview.mode redacted none, hash, redacted or full: what approvers see of the arguments.
approvals.quorum {} Four-eyes: {default: {minApprovers: 2}} or per "<policy>/<rule>".
approvals.notifiers [] Webhook, Slack, Teams or Matrix notifications of pending approvals.
dlp.profiles / dlp.defaultProfile none A default profile for all traffic (Protect data with DLP profiles).
budgets.enabled false true with namespace and user limits (Budgets and cost limits).
networkPolicy.agentNamespaces.names [maqpna-agents] The namespaces your agents run in.
serviceMonitor.enabled / monitoring.prometheusRule.enabled false true once the Prometheus Operator CRDs exist (Monitoring).
license.secretRef.name maqpna-license Keep; install the licence with maqpna license install (Licensing).
examples.mcpEcho.enabled true false, then register your MCP servers as MCPServer objects.
sovereignty.* enabled, jurisdiction EU Your jurisdiction, allowed registries and egress hosts (Sovereignty).

A minimal override for the prod profile:

# my-values.yaml
image:
  digest: sha256:...                        # from the release manifest
identity:
  trustDomain: eu.prod.example.org
sovereignty:
  jurisdiction: EU-DE
  allowedEgressHosts: ["*.svc", "*.svc.cluster.local", "idp.example.eu", "minio.audit.example.eu"]
adminAuth:
  oidc:
    issuer: https://idp.example.eu/realms/maqpna
    clientID: maqpna-console
  namespaceRoles:
    team-a: {approver: [team-a-leads]}
audit:
  worm:
    endpoint: https://minio.audit.example.eu:9000
networkPolicy:
  agentNamespaces:
    names: [team-a, team-b]

3. Validate the values offline#

maqpna values validate merges the chart's values.yaml with your files the way Helm does, checks the result against the chart's values.schema.json and the chart guards, and applies the production rules.

maqpna values validate --chart ./maqpna -f maqpna/values-production.yaml -f my-values.yaml --profile production

A values file with three gateway replicas and the default file state backend fails a chart guard. This is real output:

$ maqpna values validate --chart ./maqpna -f my-values.yaml
SEVERITY  RULE        PATH                      MESSAGE
warning   production  examples.mcpEcho.enabled  the mcp-echo demo upstream is enabled; set examples.mcpEcho.enabled=false and configure your real MCP servers
warning   production  identity.key.mode         identity.key.mode=generate makes a new signing key on every restart (tokens stop verifying); use helm or existingSecret
error     chart       gateway.replicas          gateway.replicas=3 with state.backend=file: each replica keeps its own approvals, break-glass revocations, taint and audit ledger; use state.backend=postgres or 1 replica (or guards.allowIndependentGatewayReplicas=true)

INVALID: 1 error(s), 2 warning(s) (./maqpna/values.yaml, my-values.yaml; schema maqpna/values.schema.json)

The exit status is 3. With the production profile file the same command passes:

$ maqpna values validate --chart ./maqpna -f prod.yaml --profile production
OK: values valid (./maqpna/values.yaml, prod.yaml; schema maqpna/values.schema.json)

4. Run the preflight checks#

maqpna preflight takes the same chart and value flags as install, evaluates the values install would use and checks the cluster. It creates nothing.

maqpna preflight --profile prod -f my-values.yaml --version 0.1.1 --strict

It checks, among others:

Check Fails when
kube-version The API server is older than 1.29 or unreachable.
install-permissions You cannot create CRDs, ClusterRoles, ClusterRoleBindings, namespaces, Deployments, Secrets or RuntimeClasses.
agent-sandbox sandboxes.agents.x-k8s.io is missing or its API version does not match operator.sandboxApiVersion.
runtimeclass-<name> An enabled RuntimeClass has no node carrying its handler.
cni-networkpolicy No NetworkPolicy-enforcing CNI is found.
image-registry-allowed An image registry is not in sovereignty.allowedRegistries.
gateway-replicas-state More than one gateway replica with state.backend=file.
chart-renders The chart does not render with your values.

The output is a table of PASS, WARN and FAIL lines with a fix: line under each problem, and a score. The exit status is 3 when a check fails (or a warning, with --strict).

5. Install#

maqpna install --profile prod -f my-values.yaml --version 0.1.1 --wait --timeout 15m

install applies the CRDs, installs the release maqpna into maqpna-system (--release and -n change them) and, with --wait, waits until every workload is ready. --dry-run renders on the API server, lookups included, and changes nothing. For a mirrored or local chart, pass --chart oci://registry.internal/maqpna/charts/maqpna or --chart ./maqpna.

With plain Helm (GitOps), the equivalent is:

helm upgrade --install maqpna ./maqpna -n maqpna-system --create-namespace \
  -f maqpna/values-production.yaml -f my-values.yaml

6. Smoke-test the installation#

helm test runs maqpna smoke in a pod (templates/tests/smoke.yaml):

helm test maqpna -n maqpna-system

You can also run it from your workstation against the gateway. Without an admin token the admin-API steps are skipped; without --agent-token-file the allow and deny tool calls are skipped; --sandbox NS/AGENT creates and deletes a real session. This output comes from a local MAQPNA (maqpna dev up); on a cluster the steps are the same:

$ maqpna smoke --gateway "$MAQPNA_GATEWAY_URL" --identity "$MAQPNA_BROKER_URL" --agent-token-file agent.tok \
    --deny-args '{"id":"x","namespace":"kube-system"}'
STATUS  CHECK                COMPONENT  DETAIL
PASS    gateway-ready        gateway    ready
PASS    gateway-version      gateway    gateway dev (0s)
PASS    identity-jwks        identity   1 key(s) (1ms)
PASS    mcp-unauthenticated  gateway    POST /mcp/echo -> HTTP 401 (0s)
PASS    mcp-bad-token        gateway    POST /mcp/echo -> HTTP 401 (0s)
PASS    admin-api            gateway    admin auth mode token (0s)
PASS    audit-verify         audit      chain verifies (5 records) (1ms)
PASS    mcp-allow            gateway    echo allowed (8ms)
PASS    mcp-deny             gateway    delete_resource: error -32001 denied by policy baseline-guardrails rule never-touch-system-namespaces: system namespaces are off-limits to agents (6ms)
SKIP    sandbox-lifecycle    sandbox    no --sandbox NS/AGENT

score 100/100: 0 fail, 0 warn, 9 pass, 0 info, 1 skip, 0 accepted

7. Check posture with doctor#

maqpna doctor -n maqpna-system --strict

doctor runs the gateway posture checks (through GET /v1/posture) and the cluster checks, and exits 3 when one fails. An installation made from the prod profile, with every CHANGE-ME replaced, has no failed check. The full list of checks is in Security hardening.

The PostgreSQL HA state backend#

With state.backend: file (the default), every store is a single-writer journal on the pod's volume: approvals, break-glass revocations, session taint, tool pins, budgets, the token vault, session memory and the audit ledger. That is fine for one gateway replica. With state.backend: postgres, all of them live in one PostgreSQL schema shared by every replica, and two or more replicas behave like one.

state:
  backend: postgres
  postgres:
    dsnSecret: maqpna-state      # Secret with key dsn (state.postgres.dsnSecretKey)
    schema: maqpna               # created by idempotent migrations at start
    maxConns: 10                 # pool size per pod
    pollMillis: 250              # how often a replica pulls other replicas' changes
  leader:
    retryMillis: 1000            # leader election of background jobs
  counters: ""                   # "" = shared rate-limit windows with postgres
gateway:
  replicas: 3
  pdb: {enabled: true, minAvailable: 2}
attestation:
  replicas: 2                    # refused with state.backend=file
  • Run an in-jurisdiction, highly available PostgreSQL (Patroni or CloudNativePG) with point-in-time recovery.
  • The gateway migrates the schema at start under an advisory lock. Migrations are additive and never roll back.
  • Every gateway pod keeps a verified local mirror of the whole ledger at its audit path. A replica whose local ledger is not a prefix of the shared chain refuses to start; see Troubleshooting.
  • SIEM and WORM shipping, the tenant ledger feed and retention sweeps run on one elected replica (maqpna_gateway_leader{job} shows which).
  • Approvals and revocations made on one replica are effective on every replica within pollMillis.

What happens when PostgreSQL is down: /readyz reports state backend unreachable and the pod leaves the Service; approvals cannot be created or decided; with auditFailurePolicy: closed every allowed call is refused until audit appends succeed. maqpna kill … --emit-yaml | kubectl apply -f - still works, because it goes through Kubernetes. Decide between open and closed before an incident.

Ingress and TLS#

The chart creates no Ingress. Agents reach the gateway inside the cluster (maqpna-gateway Service, port 8080). You need an external route only for:

  • people: the console and maqpna approvals, maqpna desk or maqpna login from outside the cluster;
  • external MCP clients (IDE and desktop agents) that use your identity provider's tokens;
  • OAuth callbacks of connected accounts (delegation.resourceBaseURL).

Choose one way to terminate TLS:

Option Set
TLS at your ingress controller or Gateway API route Create an Ingress or HTTPRoute to the maqpna-gateway Service, port 8080, with a certificate from cert-manager.
TLS at the gateway gateway.tls.enabled: true and gateway.tls.certSecret (a kubernetes.io/tls Secret). Required if you bind agent tokens to workload certificates (svidMode: required, with spire.enabled).

The chart's NetworkPolicy for the gateway admits only agent namespaces and pods in the release namespace on the gateway port. Add a NetworkPolicy that admits your ingress controller's namespace. With adminListen: ":9090", the admin port is a separate Service port that no NetworkPolicy opens: reach it with kubectl port-forward, or add a policy for the console's ingress and Prometheus only.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: maqpna-gateway-from-ingress
  namespace: maqpna-system
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: maqpna-gateway
  policyTypes: ["Ingress"]
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: ingress-nginx
      ports:
        - {protocol: TCP, port: 8080}

OIDC for the console and approvers#

With adminAuth.mode: oidc, the admin API accepts bearer tokens from your identity provider (RS256, ES256 or EdDSA), and the approver recorded in the audit ledger comes from the verified token, never from the request body.

adminAuth:
  mode: oidc                 # both = OIDC plus the static token as audited break-glass
  breakGlassToken:
    enabled: false
  oidc:
    issuer: https://idp.example.eu/realms/maqpna   # exact iss
    audience: maqpna-admin
    clientID: maqpna-console                        # public client for console and CLI login (PKCE)
    groupsClaim: groups
    usernameClaim: ""                               # default preferred_username, then email, then sub
  roles:
    viewer: ["maqpna-viewers"]
    approver: ["maqpna-approvers"]
    auditor: ["maqpna-auditors"]
    admin: ["maqpna-admins"]
    killswitch: ["maqpna-sre"]
  namespaceRoles:
    team-a: {approver: [team-a-leads]}

In the identity provider:

  1. Create a public client (here maqpna-console) for the authorization code flow with PKCE, and register http://127.0.0.1/callback (any port) as a redirect URI for maqpna login and maqpna console.
  2. Add an audience mapper so access tokens carry aud: maqpna-admin.
  3. Emit group names (not full paths) in the groups claim.
  4. Pin the issuer URL. Keycloak derives iss from the request host; set --hostname so every token has the same iss.
  5. Add the identity provider host to sovereignty.allowedEgressHosts: discovery and JWKS requests go through the gateway's residency-checked dialer.

Then sign in and check your roles:

maqpna context set prod --gateway https://gateway.example.eu --kube-context prod-eu --use
maqpna login            # browser, authorization code + PKCE; --device for headless hosts
maqpna whoami
maqpna console --open   # the console on 127.0.0.1, admin API proxied to the gateway

Approvers cannot approve calls made on their own behalf (409 self_approval), and the same person cannot vote twice on a four-eyes approval (409 duplicate_approver). User approval on the user's own device (CIBA) is configured under delegation.ciba; see Human approvals.

Cloud-specific notes#

Platform tier-0 tier-1 (microVM) tier-2 (confidential VM) Notes
AKS Pod Sandboxing (kata-mshv-vm-isolation); gVisor is not a supported AKS runtime Pod Sandboxing Confidential Containers (kata-cc-isolation, DCasv5/ECasv5) Disable the chart's RuntimeClasses and point the TrustTiers at AKS's classes. Use Azure CNI with Cilium for NetworkPolicy.
EKS gVisor installed by a launch template or DaemonSet Bare-metal (*.metal) instances only Not available; disable it Turn on the VPC CNI policy agent (enableNetworkPolicy: "true") or use Cilium or Calico.
GKE GKE Sandbox creates RuntimeClass gvisor; disable the chart's copy (runtimeClasses.gvisor.enabled=false) Nested virtualisation plus kata-deploy Confidential GKE Nodes protect the node only; per-pod attestation needs Kata CoCo, which GKE does not offer managed Dataplane V2 enforces NetworkPolicy.
OpenShift No gVisor from Red Hat; use Kata OpenShift sandboxed containers (kata, or kata-remote for peer pods) Sandboxed containers with confidential containers Grant the nonroot-v2 SCC to the MAQPNA service accounts (pods run as UID 65532). OVN-Kubernetes enforces NetworkPolicy.
k3s, kind gVisor through the containerd config template Kata via kata-deploy on hosts with /dev/kvm (k3s) — kind runs tier-0 only. k3s ships kube-router, which enforces NetworkPolicy. Use --profile dev.

A TrustTier overlay for a platform whose RuntimeClasses already exist (AKS shown):

runtimeClasses:
  gvisor: {enabled: false}
  kataFc: {enabled: false}
  kataQemuSnp: {enabled: false}
trustTiers:
  create: true
  items:
    - name: tier-0
      spec: {runtimeClassName: kata-mshv-vm-isolation, isolation: microvm, egress: gateway-only}
    - name: tier-1
      spec: {runtimeClassName: kata-mshv-vm-isolation, isolation: microvm, egress: gateway-only}
    - name: tier-2
      spec:
        runtimeClassName: kata-cc-isolation
        isolation: confidential
        attestationRequired: true
        egress: gateway-only

Verification#

kubectl -n maqpna-system get pods
kubectl get trusttiers
maqpna status
maqpna version --check
maqpna doctor --strict

maqpna status shows the release, workloads, sessions by phase, pending approvals, the audit ledger head and the licence state on one screen. maqpna version --check exits 3 when the components are not on the same minor version.

Troubleshooting#

Symptom Cause Fix
values validate reports gateway.replicas=3 with state.backend=file The chart guard against independent per-replica state Set state.backend: postgres with state.postgres.dsnSecret, or run 1 replica.
preflight: agent-sandbox sandboxes.agents.x-k8s.io not found The upstream controller is not installed kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v1.0.4/sandbox-with-extensions.yaml
preflight: install-permissions missing: create apiextensions.k8s.io/customresourcedefinitions, … Your Kubernetes user is not cluster admin Run the install as a cluster administrator.
an oci:// chart needs --version (the CLI is a development build) A CLI built from source has no default chart version Pass --version 0.1.1 or --chart ./maqpna.
Gateway pods not ready, /readyz lists state backend unreachable PostgreSQL unreachable or the DSN Secret is wrong Check the Secret maqpna-state and the network path to the database.
Console login fails with token issuer mismatch Tokens fetched through another host name have another iss Pin the identity provider's issuer URL.
More on errors and diagnostics — Troubleshooting

Next steps#