State backend schema and leader election#
The gateway and the attestation service keep durable state behind one abstraction, pkg/state. A state.Backend offers five things: journals (ordered event logs that a store replays into memory), chains (the audit hash chain), blobs (opaque documents), leader leases and counters. Two backends implement it:
| Backend | Where state lives | Replicas | Use it for |
|---|---|---|---|
file (default) |
fsynced JSONL files on each pod's volume, one per store | One writer per file; gateway replicas keep separate state; attest runs one replica | Development, single node, air-gapped edge |
postgres |
One PostgreSQL schema shared by every replica (pkg/state/postgres, driver pgx through database/sql) |
Gateway 2 or more with a PodDisruptionBudget; attest 2 or more | Production |
The operator, identity broker and CLI do not link the PostgreSQL driver. See state backend for the component view and configuration.
Tables#
erDiagram
state_events {
text store PK
bigint seq PK "GENERATED ALWAYS AS IDENTITY"
bigint through "0, or snapshot through seq"
text data "the JSON line a file journal would write"
timestamptz created_at
}
audit_chain {
text chain PK "audit, or tenant/name"
bigint seq PK
text hash
text prev_hash
text line "exact JSON of the record"
timestamptz appended_at
}
state_blobs {
text name PK "finops, memory/..., capture/..., cursors"
bytea data
timestamptz updated_at
}
counters {
text key PK "tpm/..., rpm/..."
bigint window_ms PK
bigint bucket PK
bigint value
timestamptz expires_at
}
schema_migrations {
int version PK
text name
timestamptz applied_at
}
The DDL lives in pkg/state/postgres/migrations/0001_init.sql and 0002_leader_counters_blobs.sql; {{schema}} is replaced with the configured schema (default maqpna, pattern [a-z_][a-z0-9_]*). Migrations run under a session advisory lock, are idempotent (IF NOT EXISTS) and are recorded in schema_migrations, so two replicas can start at once.
state_events: one ordered log per store. Rows withthrough > 0are compaction snapshots of the store as ofseq <= through.audit_chain: the ledger. A trigger,audit_chain_append_only, raises an exception on everyUPDATEorDELETE(defence in depth; the hash chain and signed checkpoints remain the tamper evidence, and a database superuser can still drop the trigger).state_blobs: snapshots and documents, with a byte-order index (state_blobs_name_c) for prefix listing and age sweeps.counters: fixed-window rate-limit counters; expired rows are deleted every minute.
What is stored#
| Kind | Name | Owner |
|---|---|---|
| Journal | approvals |
pkg/approval |
| Journal | revocations (break-glass entries) |
gateway |
| Journal | taint |
pkg/taint |
| Journal | toolpins |
pkg/toolpin |
| Journal | notify-outbox |
pkg/notify |
| Journal | vault, vault-refresh, vault-flows |
token vault and consent flows |
| Journal | attest-releases |
maqpna-attest |
| Chain | audit and tenant/<name> |
pkg/audit |
| Blob | finops |
FinOps meter snapshot |
| Blob | memory/<sha256(store)>/<sha256(user)> |
memory stores |
| Blob | capture/<ns>/<argsSha256> |
argument captures |
| Blob | siem/<sink>/cursor, worm/<kind>/offset, tenant-ledgers/cursor |
shipping positions |
| Counter | tokensPerMinute, requestsPerMinute windows |
gateway |
How a journal write works#
- The writer opens a transaction and takes
pg_advisory_xact_lock(key)for the store. The key is the FNV-64a hash ofmaqpna\0<schema>\0<kind>\0<name>, so writers of one store are serialised and per-storeseqorder is commit order. - Inside the lock it catches up: it folds every event other replicas committed since its last one.
- It appends its event and commits.
Exclusiveruns a check-then-act step in the same lock after catching up, so the check sees every committed decision of every replica. It is used to decide and consume approvals (quorum, "already used"), register attestation nonces, claim a release and claim a renewal challenge. This makes them happen exactly once across replicas.- Reads are served from memory. Each replica polls for new events every
pollMillis(default 250 ms; amax(seq)probe, then a range read). Reads where staleness matters refresh first: taint lookups, unknown or pending approvals, attestation release status, break-glass deletion. - Compaction replaces a store's log with snapshot rows under the lock. A replica that missed events while partitioned replays the store from scratch.
- An uncertain commit (connection lost during
COMMIT) marks the journal for a full replay, so memory never keeps a change that might not be durable.
Concurrent approval creates of one replica share one transaction (state.GroupAppender), which keeps asynchronous approvals fast under load.
Leader election#
Background jobs that must run once per installation (SIEM shipping per sink, write-once storage (WORM) shipping, the tenant-ledger feed, memory and capture retention sweeps) run on the holder of a lease. On PostgreSQL a lease is a session-level pg_try_advisory_lock held on one dedicated connection per replica. The file backend's lease is always held.
sequenceDiagram
participant A as Gateway replica A
participant PG as PostgreSQL
participant B as Gateway replica B
A->>PG: pg_try_advisory_lock(worm/audit) on its lease connection
PG-->>A: true (A leads)
B->>PG: pg_try_advisory_lock(worm/audit)
PG-->>B: false (B is standby)
loop every leader.retryMillis (1 s)
A->>PG: ping lease session (2 s timeout)
B->>PG: try the free leases again
end
Note over A: A crashes (TCP reset)
PG->>PG: session ends, lock released
B->>PG: pg_try_advisory_lock(worm/audit)
PG-->>B: true (B leads)
B->>PG: read blob worm/audit/offset
B->>B: resume shipping from the saved offset
| Event | Takeover |
|---|---|
| Leader shuts down cleanly | Lock released at once; a standby takes over at its next retry (within leader.retryMillis, 1 s) |
| Leader process crashes (TCP reset) | The server ends the session at once; same as above |
| Leader's host vanishes (no reset) | TCP keepalive (5 s idle, 3 probes of 1 s) ends the session after about 8 s |
| Leader loses the database | Its ping fails within one retry: it reports not leading and stops the jobs |
Jobs check Held() before each unit of work. Lease names: siem/<sink>, worm/<kind> (worm/audit, and one per tenant prefix), tenant-ledgers, memory-sweep, capture-sweep.
The operator does not use pkg/state; it uses the Kubernetes Lease maqpna-operator.maqpna.com (--leader-elect).
Counters#
Counters().Add(key, delta, window) adds to a fixed window with one INSERT … ON CONFLICT … DO UPDATE … RETURNING value. With stateBackend.counters: shared (the default with Postgres) the gateway uses it for tokensPerMinute and model-route requestsPerMinute, so these limits hold for the whole cluster. If the database cannot be reached, the gateway falls back to its local limiter: the request path never fails on a counter. maxCallsPerMinute of policy rules is still counted per replica.
What you see#
$ maqpna audit verify --postgres /var/run/maqpna-state/dsn --chain audit
OK: postgres maqpna.audit_chain chain=audit: 48211 records, hash chain intact (head seq 48211, hash 9c1e5d0b…)
The gateway exports maqpna_gateway_state_backend_info{type} and maqpna_gateway_leader{job} (1 on the leader). See maqpna audit verify and maqpna backup create.