Monitoring, alerts and runbooks
Scrape MAQPNA's metrics, turn on the shipped Prometheus alerts and Grafana dashboards, and follow the runbook for each alert and incident.
The chart ships a ServiceMonitor, a PrometheusRule with 21 alerts and three Grafana dashboards. All three are off by default. Each alert links to its runbook.
flowchart LR
subgraph MAQPNA
O[operator :8080/metrics]
G[gateway /metrics<br/>admin port with adminListen]
T[attestation /metrics]
end
O & G & T --> SM[ServiceMonitor] --> P[Prometheus]
P --> R[PrometheusRule<br/>21 alerts] --> AM[Alertmanager] --> RB[Runbook]
P --> D[Grafana dashboards<br/>Sessions, Gateway, Evidence]
Goal#
Every MAQPNA failure that affects evidence, approvals, the kill switch or session start pages someone, with a runbook that says what to do.
Prerequisites#
- Prometheus Operator CRDs (kube-prometheus-stack, for example).
helm templatewithserviceMonitor.enabled=truefails on purpose when they are missing. - kube-state-metrics, for
MaqpnaGatewayNotReady(or setmonitoring.prometheusRule.kubeStateMetrics=false). - Grafana with the dashboard sidecar, for the dashboards.
Steps#
1. Turn it on#
# monitoring-values.yaml
serviceMonitor:
enabled: true
interval: 30s
monitoring:
prometheusRule:
enabled: true
labels:
release: kube-prometheus-stack # must match your Prometheus ruleSelector
additionalAlertLabels:
team: platform
disabled: [] # e.g. [MaqpnaBudgetExceeded]
dashboards:
enabled: true
annotations:
grafana_folder: MAQPNA
maqpna upgrade -f monitoring-values.yaml --wait
| Value | Renders |
|---|---|
serviceMonitor.enabled |
A ServiceMonitor for the operator (:8080/metrics), the gateway and the attestation service. With gateway.config.adminListen, the gateway serves /metrics only on its admin port and the ServiceMonitor scrapes that port. |
monitoring.prometheusRule.enabled |
The PrometheusRule. Tune every threshold under monitoring.prometheusRule.thresholds. runbookBaseURL sets the link in each alert. |
monitoring.dashboards.enabled |
One ConfigMap per dashboard, labelled grafana_dashboard: "1". Each dashboard has a datasource variable. |
2. Check that targets are up#
In Prometheus, query:
up{job=~".*maqpna.*"}
maqpna_build_info
maqpna_build_info{component,version,commit,goversion,tags} is exported by every server, so you can see versions
per pod. maqpna version --check reads the same information from GET /version.
3. Validate the rules in CI (optional)#
From a source checkout, hack/check-alerts.sh renders the PrometheusRule, runs promtool check rules and the alert
unit tests in deploy/helm/maqpna/tests/alerts_test.yaml. It needs helm and promtool.
Metrics#
Operator#
Served on --metrics-bind-address (default :8080, port metrics). Session metrics come from the leader replica.
| Metric | Type | Meaning |
|---|---|---|
maqpna_session_ready_seconds{tier,mode} |
histogram | Time from session creation to its first Running phase. mode is claim (warm pool) or direct. |
maqpna_sessions{namespace,tier,phase} |
gauge | Sessions by phase. |
maqpna_session_phase_transitions_total{from,to,reason} |
counter | Phase changes. |
maqpna_session_outcomes_total{tier,result} |
counter | Ended sessions by result. |
maqpna_token_mint_seconds |
histogram | Identity broker token mint latency. |
maqpna_warm_pool_ready{tier} |
gauge | Ready warm sandboxes per tier. |
maqpna_reconcile_errors_total{controller,reason} |
counter | Reconcile errors by Kubernetes API status reason. |
maqpna_rendered_configmap_bytes{name} |
gauge | Size of each gateway ConfigMap the operator renders (limit 1 MiB). |
maqpna_license_state, maqpna_license_days_left, maqpna_license_node_limit |
gauge | Licence state; never affects agent traffic. |
maqpna_billable_nodes{period} |
gauge | Distinct nodes that ran sandbox pods (current, day, month). |
Gateway and attestation service#
| Metric | Used by |
|---|---|
maqpna_gateway_decisions_total{action,server} |
MaqpnaGatewayDenyRateHigh, Gateway dashboard |
maqpna_gateway_audit_commit_seconds |
MaqpnaGatewayAuditCommitSlow |
maqpna_gateway_audit_errors_total |
MaqpnaAuditLedgerUnavailable |
maqpna_gateway_approvals_pending |
MaqpnaApprovalsBacklog |
maqpna_gateway_revocation_list_ok |
MaqpnaRevocationListUnavailable |
maqpna_gateway_configsync_last_success_timestamp_seconds{configmap} |
MaqpnaRevocationLag |
maqpna_gateway_configsync_errors_total{configmap} |
Failed configSync reads |
maqpna_gateway_leader{job} |
Which replica runs each background job (PostgreSQL state) |
maqpna_gateway_policy_reloads_total{result} |
MaqpnaConfigSyncFailures |
maqpna_gateway_upstream_errors_total{server}, maqpna_gateway_upstream_latency_seconds |
MaqpnaGatewayUpstreamErrors |
maqpna_gateway_budget_denied_total{scope} |
MaqpnaBudgetExceeded |
maqpna_gateway_model_circuits_open |
MaqpnaModelCircuitOpen |
maqpna_gateway_audit_sink_lag{sink} |
MaqpnaAuditSinkLag |
maqpna_attest_failure_total{reason}, maqpna_attest_success_total |
MaqpnaAttestationFailureRate |
Alerts#
| Alert | Severity | Fires when | First step |
|---|---|---|---|
MaqpnaGatewayNotReady |
critical | A gateway pod is unready for 5 minutes | GET /readyz on the pod lists the problems. |
MaqpnaAuditLedgerUnavailable |
critical | Audit appends fail | Check the ledger volume or the state backend; then maqpna audit verify. |
MaqpnaRevocationListUnavailable |
critical | The gateway cannot read the revocation list and denies every call | Check the maqpna-revocations ConfigMap and the operator's agentrevocation controller. |
MaqpnaRevocationLag |
critical | The gateway's copy of the revocations ConfigMap is older than 10 s for 1 minute | New revocations are not enforced by that replica; check configsync_errors_total and the gateway Role. |
MaqpnaLicenseExpiring |
warning, then critical | Licence expires in under 30, then 7 days | maqpna license status; install the renewal. Agent traffic is never affected. |
MaqpnaSessionReadySLO |
warning | p90 session start above 1 s (warm) or 5 s (cold) for 15 minutes | Is the warm pool full? Read the slow session's events. |
MaqpnaSessionFailureRate |
warning | More than 25% of ended sessions Failed |
Agent failures (ExitCode) or platform refusals (ScopeEscalation, AttestationFailed). |
MaqpnaWarmPoolEmpty |
warning | A tier's warm pool has no ready sandbox for 10 minutes | Check the SandboxWarmPool and its pods. |
MaqpnaTokenMintSlow |
warning | p99 token mint above 1 s | Identity broker logs and its KMS or HSM latency. |
MaqpnaOperatorReconcileErrors |
warning | A controller keeps failing | Forbidden: re-apply the chart. Other: the broker or attestation service is unreachable. |
MaqpnaConfigMapNearLimit |
warning | A rendered ConfigMap is above 800 KiB | Split large policies; remove unused MCP servers or model routes. |
MaqpnaGatewayAuditCommitSlow |
warning | p99 durable audit commit above 50 ms | Ledger volume IOPS and fsync latency, or PostgreSQL. |
MaqpnaApprovalsBacklog |
warning | More than 50 approvals pending for 10 minutes | Add approvers, or review require_approval rules. |
MaqpnaConfigSyncFailures |
warning | The gateway fails to reload its policy file | The gateway log names the parse error; it keeps the last good policy set. |
MaqpnaGatewayDenyRateHigh |
warning | More than half of tool calls denied for 15 minutes | maqpna audit tail --decision deny: a strict policy, a revoked agent or an agent probing for tools. |
MaqpnaGatewayUpstreamErrors |
warning | More than 5% of MCP round trips fail | MCP server health; kubectl get mcpservers -A. |
MaqpnaModelCircuitOpen |
warning | A model endpoint's circuit breaker is open for 10 minutes | Check the model endpoint. |
MaqpnaAuditSinkLag |
warning | A SIEM or WORM sink is over 1,000 records behind | maqpna audit sinks; check reachability and credentials. |
MaqpnaAttestationFailureRate |
warning | More than 10% of attestations fail | maqpna_attest_failure_total{reason}. |
MaqpnaBudgetExceeded |
info | Calls are denied because a budget is exhausted | Raise the budget or wait for the window to reset. |
MaqpnaLicenseNodeLimitExceeded |
info | Billable nodes this month above the licence limit | Nothing is blocked; extend the licence or use fewer nodes. |
Runbooks#
| Incident | Trigger | Steps |
|---|---|---|
| Kill switch | Misuse, a runaway loop, a leaked token | maqpna kill --gateway "$GW" -n NS --agent A --reason "INC-123" --terminate; make it durable with --emit-yaml \| kubectl apply -f -; verify with maqpna revocations list. See Kill switch and revocations. |
| Compromised agent | Prompt injection, a poisoned tool, a stolen token | Contain with maqpna kill … --terminate; preserve evidence with maqpna audit verify --gateway, audit export, audit fetch ledger; find the entry point with maqpna taint list, maqpna mcp tools NS/SERVER --drift-only and maqpna replay; rotate what could have leaked (maqpna keys rotate identity). |
| Ledger divergence or tamper | audit verify reports TAMPERED, replicas report different heads |
maqpna audit verify --gateway, --postgres dsn.txt, and offline with --jwks; compare with the last backup, WORM and SIEM copies and signed checkpoints; restore into a new ledger, never rewrite the chain. |
| PostgreSQL outage | /readyz: state backend unreachable; audit errors rising |
Fail over the database; then maqpna audit verify --postgres dsn.txt, maqpna audit sinks --max-lag 1000, maqpna approvals list. While it is down, maqpna kill … --emit-yaml \| kubectl apply -f - still works. |
| Attestation failure | MaqpnaAttestationFailureRate, sessions in AttestationFailed |
Read maqpna_attest_failure_total{reason}: measurement means update reference values after review; nonce/replay means check node clocks; verifier means check the Trustee. Never switch to the sample verifier in production. |
| Key rotation | Schedule, or a suspected leak | maqpna keys rotate identity\|audit-checkpoint\|vault\|attest; see Backup, restore and DR. |
| Backup and restore | Data loss, a monthly drill | maqpna backup create, backup verify, restore, dr drill; see Backup, restore and DR. |
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
No gateway targets with adminListen set |
Your Prometheus scrapes the agent port | The ServiceMonitor scrapes the admin port; with your own scrape config, use it too. |
maqpna_sessions shows exported_namespace |
Label clash on the operator endpoint | The shipped ServiceMonitor sets honorLabels: true; set it in your own scrape config. |
MaqpnaRevocationLag never fires |
gateway.configSync.enabled=false |
The metric is absent without configSync; revocations then propagate with the kubelet's 60–90 s volume refresh. |
| Alerts render but never load | monitoring.prometheusRule.labels do not match the Prometheus ruleSelector |
Set the label your Prometheus selects (release: <its release name>). |
Next steps#
- Scaling and HA
- Troubleshooting
- Command reference:
maqpna status,maqpna doctor,maqpna audit sinks.