State backend#
The gateway and the attestation service keep durable state: approvals, break-glass revocations, session taint, tool-pin observations, the notification outbox, FinOps meter state, the token vault and consent flows, session memory, argument captures, the audit ledger and per-tenant ledgers, SIEM and WORM shipping positions, rate-limit windows, and attestation releases and renewals. pkg/state puts all of them behind one abstraction with two backends.
| Backend | Where state lives | Replicas | Use it for |
|---|---|---|---|
file (default) |
fsynced JSONL journals on each pod's volume | Gateway: any, but state is per replica; attestation: 1 | Development, single node, air-gapped edge |
postgres |
One PostgreSQL schema shared by every replica | Gateway 2 or more with a PDB; attestation 2 or more | Production |
Code: pkg/state, pkg/state/postgres (the only place the pgx driver is linked; not in the operator, broker, CLI or SDKs).
Neighbours#
flowchart TB
subgraph G["Gateway replicas"]
GA["replica A<br/>memory view + local ledger mirror"]
GB["replica B<br/>memory view + local ledger mirror"]
end
subgraph A["maqpna-attest replicas"]
AA["replica 1"]
AB["replica 2"]
end
subgraph PG["PostgreSQL, schema maqpna"]
E[("state_events<br/>journals per store")]
C[("audit_chain<br/>append-only trigger")]
B[("state_blobs<br/>FinOps, memory, captures, cursors")]
N[("counters<br/>rate-limit windows")]
L["advisory locks<br/>writers and leader leases"]
end
GA & GB -- "append under pg_advisory_xact_lock,<br/>poll every pollMillis" --> E
GA & GB -- "seal and insert records" --> C
GA & GB --> B & N
GA & GB -. "pg_try_advisory_lock per job" .-> L
AA & AB -- "journal attest-releases" --> E
Responsibilities#
| Abstraction | Backed by | Used for |
|---|---|---|
Journal |
Table state_events, one ordered log per store |
approvals, revocations, taint, toolpins, notify-outbox, attest-releases, vault, vault-refresh, vault-flows |
Exclusive (check-then-act under the store lock) |
Same | Deciding and consuming approvals, registering attestation nonces, claiming releases and renewal challenges: each happens exactly once across replicas |
Chain |
Table audit_chain |
The platform ledger (audit) and one chain per tenant (tenant/<name>) |
Blobs |
Table state_blobs |
FinOps meter snapshot (finops), memory documents, captures (capture/<ns>/<sha>), shipping cursors (siem/<sink>/cursor, worm/<kind>/offset) |
Counters |
Table counters |
tokensPerMinute and model requestsPerMinute fixed one-minute windows |
Leader |
Session-level advisory lock on a dedicated connection | One replica per background job |
How it works#
- Writes. A replica writes a store's event in a transaction holding
pg_advisory_xact_lockfor that store. Inside it, it first folds every event other replicas committed since its last one, then appends. Per-store order is commit order, and every replica folds the same events in the same order. - Reads. Reads are served from memory. Each replica pulls new events every
pollMillis(250 ms) with a cheapmax(seq)probe. Lookups where staleness matters refresh on demand: taint (a session tainted on replica A is tainted on replica B for the very next call), unknown or pending approvals, attestation release status. - Ledger appends. An append takes the chain's advisory lock, reads the head, seals the record (
seq = head + 1,prevHash = head hash) and inserts it. Each replica keeps a byte-identical, verified local mirror atauditPath; a mirror that does not match the database stops the replica (ErrMirrorDiverged) until an operator setsauditAcceptDivergedMirrorafter inspection. - Leader leases. Each background job campaigns for a lease every
leader.retryMillis(1 s). Gateway jobs:siem/<sink>,worm/<kind>,tenant-ledgers,memory-sweep,capture-sweep. Jobs check the lease before each unit of work, and a new leader resumes from the shared cursor. - Compaction. At start-up a replica replaces each store's log with snapshot rows; a replica that missed events (it was partitioned) replays the store from scratch.
The schema and locking are detailed in state backend schema and leader election.
Configuration#
Gateway (gateway.json):
| Key | Default | Meaning |
|---|---|---|
stateBackend.type |
file |
file or postgres |
stateBackend.dsnFile |
— (required for postgres) | File holding the DSN; never inline, so credentials stay out of ConfigMaps |
stateBackend.schema |
maqpna |
Schema, created if missing |
stateBackend.maxConns |
10 | Pool size per process |
stateBackend.pollMillis |
250 | Pull interval for other replicas' changes |
stateBackend.leader.retryMillis |
1000 | Lease campaign and verification interval |
stateBackend.counters |
shared with postgres, local with file |
Where rate-limit windows live |
Attestation service: -state-backend postgres -state-dsn-file FILE [-state-schema maqpna] [-state-max-conns 10].
Helm: state.backend, state.postgres.{dsnSecret, schema, maxConns, pollMillis}, state.leader.retryMillis, state.counters. The chart mounts the DSN Secret at /var/run/maqpna-state/dsn.
Failure modes#
| Situation | Behaviour |
|---|---|
| Database unreachable | Gateway /readyz 503 state backend unreachable (pod leaves the Service); creating, deciding and consuming approvals fail closed; break-glass changes return 500; audit appends fail and are counted; attestation refused (reason store); reads serve the last folded state; rate limits fall back to local limiters |
| Replica partitioned, then back | The poller catches up; after a compaction it replays from scratch |
| PostgreSQL failover (Patroni, CloudNativePG) | The pool reconnects; in-flight transactions fail and callers retry |
| Leader crashes | The server releases its lock at once on a TCP reset (takeover in about 1 s), or after about 8 s of keepalive if the host vanishes |
Uncertain commit (connection lost during COMMIT) |
The journal is marked for a full replay so memory never keeps a change that might not be durable |
Forged or truncated audit_chain rows |
The verified mirror stops instead of propagating them |
Scaling and limits#
- Approval creates are grouped into one transaction per batch (
GroupAppender); a laptop benchmark went from 684 to 7,393 creates per second. - Still per replica: ToolPolicy
maxCallsPerMinuteandmaxConcurrentSessionscounters (the operator enforces concurrency). - Budgets see other replicas' spend with about 1 second of lag (the ledger mirror sync).
- Planned:
LISTEN/NOTIFYto cut the polling interval, batching of the tenant ledger feed, Redis or etcd backends.
Metrics#
maqpna_gateway_state_backend_info{type}, maqpna_gateway_leader{job} (1 on the leader), maqpna_gateway_audit_errors_total, maqpna_gateway_audit_commit_seconds.
What you see#
Verify the shared ledger straight from the database, independently of any gateway:
$ maqpna audit verify --postgres /var/run/maqpna-state/dsn --chain audit
OK: postgres maqpna.audit_chain chain=audit: 48211 records, hash chain intact (head seq 48211, hash 9f2c4e1a...)
--chain tenant/<name> verifies a tenant ledger; the command exits 3 on a break. The line format comes from cmd/maqpna/admin_audit.go; the values are illustrative. maqpna doctor reports state-backend and gateway-replicas-state. See maqpna audit verify.