MAQPNADocs

State backend schema and leader election#

The gateway and the attestation service keep durable state behind one abstraction, pkg/state. A state.Backend offers five things: journals (ordered event logs that a store replays into memory), chains (the audit hash chain), blobs (opaque documents), leader leases and counters. Two backends implement it:

Backend Where state lives Replicas Use it for
file (default) fsynced JSONL files on each pod's volume, one per store One writer per file; gateway replicas keep separate state; attest runs one replica Development, single node, air-gapped edge
postgres One PostgreSQL schema shared by every replica (pkg/state/postgres, driver pgx through database/sql) Gateway 2 or more with a PodDisruptionBudget; attest 2 or more Production

The operator, identity broker and CLI do not link the PostgreSQL driver. See state backend for the component view and configuration.

Tables#

erDiagram
    state_events {
        text store PK
        bigint seq PK "GENERATED ALWAYS AS IDENTITY"
        bigint through "0, or snapshot through seq"
        text data "the JSON line a file journal would write"
        timestamptz created_at
    }
    audit_chain {
        text chain PK "audit, or tenant/name"
        bigint seq PK
        text hash
        text prev_hash
        text line "exact JSON of the record"
        timestamptz appended_at
    }
    state_blobs {
        text name PK "finops, memory/..., capture/..., cursors"
        bytea data
        timestamptz updated_at
    }
    counters {
        text key PK "tpm/..., rpm/..."
        bigint window_ms PK
        bigint bucket PK
        bigint value
        timestamptz expires_at
    }
    schema_migrations {
        int version PK
        text name
        timestamptz applied_at
    }

The DDL lives in pkg/state/postgres/migrations/0001_init.sql and 0002_leader_counters_blobs.sql; {{schema}} is replaced with the configured schema (default maqpna, pattern [a-z_][a-z0-9_]*). Migrations run under a session advisory lock, are idempotent (IF NOT EXISTS) and are recorded in schema_migrations, so two replicas can start at once.

  • state_events: one ordered log per store. Rows with through > 0 are compaction snapshots of the store as of seq <= through.
  • audit_chain: the ledger. A trigger, audit_chain_append_only, raises an exception on every UPDATE or DELETE (defence in depth; the hash chain and signed checkpoints remain the tamper evidence, and a database superuser can still drop the trigger).
  • state_blobs: snapshots and documents, with a byte-order index (state_blobs_name_c) for prefix listing and age sweeps.
  • counters: fixed-window rate-limit counters; expired rows are deleted every minute.

What is stored#

Kind Name Owner
Journal approvals pkg/approval
Journal revocations (break-glass entries) gateway
Journal taint pkg/taint
Journal toolpins pkg/toolpin
Journal notify-outbox pkg/notify
Journal vault, vault-refresh, vault-flows token vault and consent flows
Journal attest-releases maqpna-attest
Chain audit and tenant/<name> pkg/audit
Blob finops FinOps meter snapshot
Blob memory/<sha256(store)>/<sha256(user)> memory stores
Blob capture/<ns>/<argsSha256> argument captures
Blob siem/<sink>/cursor, worm/<kind>/offset, tenant-ledgers/cursor shipping positions
Counter tokensPerMinute, requestsPerMinute windows gateway

How a journal write works#

  1. The writer opens a transaction and takes pg_advisory_xact_lock(key) for the store. The key is the FNV-64a hash of maqpna\0<schema>\0<kind>\0<name>, so writers of one store are serialised and per-store seq order is commit order.
  2. Inside the lock it catches up: it folds every event other replicas committed since its last one.
  3. It appends its event and commits.
  4. Exclusive runs a check-then-act step in the same lock after catching up, so the check sees every committed decision of every replica. It is used to decide and consume approvals (quorum, "already used"), register attestation nonces, claim a release and claim a renewal challenge. This makes them happen exactly once across replicas.
  5. Reads are served from memory. Each replica polls for new events every pollMillis (default 250 ms; a max(seq) probe, then a range read). Reads where staleness matters refresh first: taint lookups, unknown or pending approvals, attestation release status, break-glass deletion.
  6. Compaction replaces a store's log with snapshot rows under the lock. A replica that missed events while partitioned replays the store from scratch.
  7. An uncertain commit (connection lost during COMMIT) marks the journal for a full replay, so memory never keeps a change that might not be durable.

Concurrent approval creates of one replica share one transaction (state.GroupAppender), which keeps asynchronous approvals fast under load.

Leader election#

Background jobs that must run once per installation (SIEM shipping per sink, write-once storage (WORM) shipping, the tenant-ledger feed, memory and capture retention sweeps) run on the holder of a lease. On PostgreSQL a lease is a session-level pg_try_advisory_lock held on one dedicated connection per replica. The file backend's lease is always held.

sequenceDiagram
    participant A as Gateway replica A
    participant PG as PostgreSQL
    participant B as Gateway replica B
    A->>PG: pg_try_advisory_lock(worm/audit) on its lease connection
    PG-->>A: true (A leads)
    B->>PG: pg_try_advisory_lock(worm/audit)
    PG-->>B: false (B is standby)
    loop every leader.retryMillis (1 s)
        A->>PG: ping lease session (2 s timeout)
        B->>PG: try the free leases again
    end
    Note over A: A crashes (TCP reset)
    PG->>PG: session ends, lock released
    B->>PG: pg_try_advisory_lock(worm/audit)
    PG-->>B: true (B leads)
    B->>PG: read blob worm/audit/offset
    B->>B: resume shipping from the saved offset
Event Takeover
Leader shuts down cleanly Lock released at once; a standby takes over at its next retry (within leader.retryMillis, 1 s)
Leader process crashes (TCP reset) The server ends the session at once; same as above
Leader's host vanishes (no reset) TCP keepalive (5 s idle, 3 probes of 1 s) ends the session after about 8 s
Leader loses the database Its ping fails within one retry: it reports not leading and stops the jobs

Jobs check Held() before each unit of work. Lease names: siem/<sink>, worm/<kind> (worm/audit, and one per tenant prefix), tenant-ledgers, memory-sweep, capture-sweep.

The operator does not use pkg/state; it uses the Kubernetes Lease maqpna-operator.maqpna.com (--leader-elect).

Counters#

Counters().Add(key, delta, window) adds to a fixed window with one INSERT … ON CONFLICT … DO UPDATE … RETURNING value. With stateBackend.counters: shared (the default with Postgres) the gateway uses it for tokensPerMinute and model-route requestsPerMinute, so these limits hold for the whole cluster. If the database cannot be reached, the gateway falls back to its local limiter: the request path never fails on a counter. maxCallsPerMinute of policy rules is still counted per replica.

What you see#

$ maqpna audit verify --postgres /var/run/maqpna-state/dsn --chain audit
OK: postgres maqpna.audit_chain chain=audit: 48211 records, hash chain intact (head seq 48211, hash 9c1e5d0b…)

The gateway exports maqpna_gateway_state_backend_info{type} and maqpna_gateway_leader{job} (1 on the leader). See maqpna audit verify and maqpna backup create.