MAQPNADocs

Monitoring, alerts and runbooks

Scrape MAQPNA's metrics, turn on the shipped Prometheus alerts and Grafana dashboards, and follow the runbook for each alert and incident.

The chart ships a ServiceMonitor, a PrometheusRule with 21 alerts and three Grafana dashboards. All three are off by default. Each alert links to its runbook.

flowchart LR
  subgraph MAQPNA
    O[operator :8080/metrics]
    G[gateway /metrics<br/>admin port with adminListen]
    T[attestation /metrics]
  end
  O & G & T --> SM[ServiceMonitor] --> P[Prometheus]
  P --> R[PrometheusRule<br/>21 alerts] --> AM[Alertmanager] --> RB[Runbook]
  P --> D[Grafana dashboards<br/>Sessions, Gateway, Evidence]

Goal#

Every MAQPNA failure that affects evidence, approvals, the kill switch or session start pages someone, with a runbook that says what to do.

Prerequisites#

  • Prometheus Operator CRDs (kube-prometheus-stack, for example). helm template with serviceMonitor.enabled=true fails on purpose when they are missing.
  • kube-state-metrics, for MaqpnaGatewayNotReady (or set monitoring.prometheusRule.kubeStateMetrics=false).
  • Grafana with the dashboard sidecar, for the dashboards.

Steps#

1. Turn it on#

# monitoring-values.yaml
serviceMonitor:
  enabled: true
  interval: 30s
monitoring:
  prometheusRule:
    enabled: true
    labels:
      release: kube-prometheus-stack      # must match your Prometheus ruleSelector
    additionalAlertLabels:
      team: platform
    disabled: []                          # e.g. [MaqpnaBudgetExceeded]
  dashboards:
    enabled: true
    annotations:
      grafana_folder: MAQPNA
maqpna upgrade -f monitoring-values.yaml --wait
Value Renders
serviceMonitor.enabled A ServiceMonitor for the operator (:8080/metrics), the gateway and the attestation service. With gateway.config.adminListen, the gateway serves /metrics only on its admin port and the ServiceMonitor scrapes that port.
monitoring.prometheusRule.enabled The PrometheusRule. Tune every threshold under monitoring.prometheusRule.thresholds. runbookBaseURL sets the link in each alert.
monitoring.dashboards.enabled One ConfigMap per dashboard, labelled grafana_dashboard: "1". Each dashboard has a datasource variable.

2. Check that targets are up#

In Prometheus, query:

up{job=~".*maqpna.*"}
maqpna_build_info

maqpna_build_info{component,version,commit,goversion,tags} is exported by every server, so you can see versions per pod. maqpna version --check reads the same information from GET /version.

3. Validate the rules in CI (optional)#

From a source checkout, hack/check-alerts.sh renders the PrometheusRule, runs promtool check rules and the alert unit tests in deploy/helm/maqpna/tests/alerts_test.yaml. It needs helm and promtool.

Metrics#

Operator#

Served on --metrics-bind-address (default :8080, port metrics). Session metrics come from the leader replica.

Metric Type Meaning
maqpna_session_ready_seconds{tier,mode} histogram Time from session creation to its first Running phase. mode is claim (warm pool) or direct.
maqpna_sessions{namespace,tier,phase} gauge Sessions by phase.
maqpna_session_phase_transitions_total{from,to,reason} counter Phase changes.
maqpna_session_outcomes_total{tier,result} counter Ended sessions by result.
maqpna_token_mint_seconds histogram Identity broker token mint latency.
maqpna_warm_pool_ready{tier} gauge Ready warm sandboxes per tier.
maqpna_reconcile_errors_total{controller,reason} counter Reconcile errors by Kubernetes API status reason.
maqpna_rendered_configmap_bytes{name} gauge Size of each gateway ConfigMap the operator renders (limit 1 MiB).
maqpna_license_state, maqpna_license_days_left, maqpna_license_node_limit gauge Licence state; never affects agent traffic.
maqpna_billable_nodes{period} gauge Distinct nodes that ran sandbox pods (current, day, month).

Gateway and attestation service#

Metric Used by
maqpna_gateway_decisions_total{action,server} MaqpnaGatewayDenyRateHigh, Gateway dashboard
maqpna_gateway_audit_commit_seconds MaqpnaGatewayAuditCommitSlow
maqpna_gateway_audit_errors_total MaqpnaAuditLedgerUnavailable
maqpna_gateway_approvals_pending MaqpnaApprovalsBacklog
maqpna_gateway_revocation_list_ok MaqpnaRevocationListUnavailable
maqpna_gateway_configsync_last_success_timestamp_seconds{configmap} MaqpnaRevocationLag
maqpna_gateway_configsync_errors_total{configmap} Failed configSync reads
maqpna_gateway_leader{job} Which replica runs each background job (PostgreSQL state)
maqpna_gateway_policy_reloads_total{result} MaqpnaConfigSyncFailures
maqpna_gateway_upstream_errors_total{server}, maqpna_gateway_upstream_latency_seconds MaqpnaGatewayUpstreamErrors
maqpna_gateway_budget_denied_total{scope} MaqpnaBudgetExceeded
maqpna_gateway_model_circuits_open MaqpnaModelCircuitOpen
maqpna_gateway_audit_sink_lag{sink} MaqpnaAuditSinkLag
maqpna_attest_failure_total{reason}, maqpna_attest_success_total MaqpnaAttestationFailureRate

Alerts#

Alert Severity Fires when First step
MaqpnaGatewayNotReady critical A gateway pod is unready for 5 minutes GET /readyz on the pod lists the problems.
MaqpnaAuditLedgerUnavailable critical Audit appends fail Check the ledger volume or the state backend; then maqpna audit verify.
MaqpnaRevocationListUnavailable critical The gateway cannot read the revocation list and denies every call Check the maqpna-revocations ConfigMap and the operator's agentrevocation controller.
MaqpnaRevocationLag critical The gateway's copy of the revocations ConfigMap is older than 10 s for 1 minute New revocations are not enforced by that replica; check configsync_errors_total and the gateway Role.
MaqpnaLicenseExpiring warning, then critical Licence expires in under 30, then 7 days maqpna license status; install the renewal. Agent traffic is never affected.
MaqpnaSessionReadySLO warning p90 session start above 1 s (warm) or 5 s (cold) for 15 minutes Is the warm pool full? Read the slow session's events.
MaqpnaSessionFailureRate warning More than 25% of ended sessions Failed Agent failures (ExitCode) or platform refusals (ScopeEscalation, AttestationFailed).
MaqpnaWarmPoolEmpty warning A tier's warm pool has no ready sandbox for 10 minutes Check the SandboxWarmPool and its pods.
MaqpnaTokenMintSlow warning p99 token mint above 1 s Identity broker logs and its KMS or HSM latency.
MaqpnaOperatorReconcileErrors warning A controller keeps failing Forbidden: re-apply the chart. Other: the broker or attestation service is unreachable.
MaqpnaConfigMapNearLimit warning A rendered ConfigMap is above 800 KiB Split large policies; remove unused MCP servers or model routes.
MaqpnaGatewayAuditCommitSlow warning p99 durable audit commit above 50 ms Ledger volume IOPS and fsync latency, or PostgreSQL.
MaqpnaApprovalsBacklog warning More than 50 approvals pending for 10 minutes Add approvers, or review require_approval rules.
MaqpnaConfigSyncFailures warning The gateway fails to reload its policy file The gateway log names the parse error; it keeps the last good policy set.
MaqpnaGatewayDenyRateHigh warning More than half of tool calls denied for 15 minutes maqpna audit tail --decision deny: a strict policy, a revoked agent or an agent probing for tools.
MaqpnaGatewayUpstreamErrors warning More than 5% of MCP round trips fail MCP server health; kubectl get mcpservers -A.
MaqpnaModelCircuitOpen warning A model endpoint's circuit breaker is open for 10 minutes Check the model endpoint.
MaqpnaAuditSinkLag warning A SIEM or WORM sink is over 1,000 records behind maqpna audit sinks; check reachability and credentials.
MaqpnaAttestationFailureRate warning More than 10% of attestations fail maqpna_attest_failure_total{reason}.
MaqpnaBudgetExceeded info Calls are denied because a budget is exhausted Raise the budget or wait for the window to reset.
MaqpnaLicenseNodeLimitExceeded info Billable nodes this month above the licence limit Nothing is blocked; extend the licence or use fewer nodes.

Runbooks#

Incident Trigger Steps
Kill switch Misuse, a runaway loop, a leaked token maqpna kill --gateway "$GW" -n NS --agent A --reason "INC-123" --terminate; make it durable with --emit-yaml \| kubectl apply -f -; verify with maqpna revocations list. See Kill switch and revocations.
Compromised agent Prompt injection, a poisoned tool, a stolen token Contain with maqpna kill … --terminate; preserve evidence with maqpna audit verify --gateway, audit export, audit fetch ledger; find the entry point with maqpna taint list, maqpna mcp tools NS/SERVER --drift-only and maqpna replay; rotate what could have leaked (maqpna keys rotate identity).
Ledger divergence or tamper audit verify reports TAMPERED, replicas report different heads maqpna audit verify --gateway, --postgres dsn.txt, and offline with --jwks; compare with the last backup, WORM and SIEM copies and signed checkpoints; restore into a new ledger, never rewrite the chain.
PostgreSQL outage /readyz: state backend unreachable; audit errors rising Fail over the database; then maqpna audit verify --postgres dsn.txt, maqpna audit sinks --max-lag 1000, maqpna approvals list. While it is down, maqpna kill … --emit-yaml \| kubectl apply -f - still works.
Attestation failure MaqpnaAttestationFailureRate, sessions in AttestationFailed Read maqpna_attest_failure_total{reason}: measurement means update reference values after review; nonce/replay means check node clocks; verifier means check the Trustee. Never switch to the sample verifier in production.
Key rotation Schedule, or a suspected leak maqpna keys rotate identity\|audit-checkpoint\|vault\|attest; see Backup, restore and DR.
Backup and restore Data loss, a monthly drill maqpna backup create, backup verify, restore, dr drill; see Backup, restore and DR.

Troubleshooting#

Symptom Cause Fix
No gateway targets with adminListen set Your Prometheus scrapes the agent port The ServiceMonitor scrapes the admin port; with your own scrape config, use it too.
maqpna_sessions shows exported_namespace Label clash on the operator endpoint The shipped ServiceMonitor sets honorLabels: true; set it in your own scrape config.
MaqpnaRevocationLag never fires gateway.configSync.enabled=false The metric is absent without configSync; revocations then propagate with the kubelet's 60–90 s volume refresh.
Alerts render but never load monitoring.prometheusRule.labels do not match the Prometheus ruleSelector Set the label your Prometheus selects (release: <its release name>).

Next steps#