Skip to content
v0.3.8GitHub

Operations runbook

Verified against cerbix f5240f5Report a problem ↗

These are the operator tasks you will meet most often: upgrades, backups, key rotation, the two recovery commands, and the reliability states that need a person. Each section is a short procedure. The full runbook on GitHub has the background and the rarer cases.

Before every upgrade, take a database backup and read the release notes. Migrations are forward-only.

Single node. Replace the binary or image and start it. Migrations apply on startup.

Terminal window
# Docker Compose: bump CERBIX_IMAGE in docker/.env, then
docker compose --env-file docker/.env -f docker/docker-compose.prod.yml pull cerbix
docker compose --env-file docker/.env -f docker/docker-compose.prod.yml up -d

Distributed. Make the schema change a deliberate step:

  1. Stop every outbox owner: all all, api and scheduler processes. Workers and agents do not deliver notifications and can keep probing.

  2. Run the migration once with the new binary:

    Terminal window
    cerbix migrate --config /etc/cerbix/config.yaml
  3. Start the new scheduler, then the workers, then api.

Step 1 is required when an upgrade crosses migration 00088. Without it, an old process can deliver one already-claimed batch of an incident’s events out of order. Nothing is lost.

Each refusal below stops before it changes the schema: the process exits with db_migrate_failed (or cerbix migrate exits non-zero), and the database stays on the migration before it. Nothing is lost.

Refusal Cause Action
The server is older than PostgreSQL 15 cerbix migrate checks the server version before applying anything. Upgrade PostgreSQL to 15 or newer (16 matches the repository images). There is no compatibility mode.
Migration 00104: rows at protocol_version 4 in pull_tests No cerbix build writes such rows, so the migration will not delete them for you. Inspect, then delete explicitly (below).
Migration 00106: a cross-project alert-routing reference A monitor, escalation step or on-call participant points at a channel or schedule in another project. Keep the old version running, find the rows with the read-only checks, correct them through the project API, then retry.
SELECT id, region, protocol_version, created_at FROM pull_tests WHERE protocol_version = 4;
-- once you know what wrote them:
DELETE FROM pull_tests WHERE protocol_version = 4;
  • PostgreSQL holds all cerbix state. The binary writes nothing to local disk. Back up with your usual PostgreSQL tooling, for example pg_dump -U cerbix cerbix > cerbix.sql.
  • Back up the keys separately. security.encryption_key (and any security.previous_keys) decrypts stored channel credentials, webhook secrets, TOTP secrets and project secrets. Without it a restored database cannot use them. Keep the regional security.dispatch keys with it.
  • RabbitMQ holds only jobs and results in flight. The scheduler re-issues expired jobs on its next tick. For broker upgrades, follow the staged RabbitMQ procedure: RabbitMQ supports neither downgrades nor a direct 3.12 to 4.3 jump.

The at-rest key lives on core roles (all, api, scheduler) only.

  1. On every core role, set the new key as security.encryption_key and move the old one into security.previous_keys.

  2. Roll the core roles. Readers decrypt old rows; writers use the new key.

  3. Re-encrypt stored secrets:

    Terminal window
    cerbix reencrypt --config /etc/cerbix/config.yaml

    Exit 0 means no project secret remains under an old key; webhook and channel secrets are rewritten under the new key too. A non-zero exit logs the failure. If it reports that re-encryption did not converge (a concurrent secret rotation), run it again.

  4. Remove the old key from security.previous_keys only after the command succeeds and every core replica runs the new config.

A key is openssl rand -base64 32: a base64-encoded 32-byte value.

reencrypt does not touch job payloads in RabbitMQ or the pull queue. For one region:

  1. Deploy {primary: <new>, previous: [<old>]} for that region to the core roles and to the region’s workers or agents. Executors need the overlap before core starts sealing with the new key.
  2. Check executor /readyz and confirm that cerbix_executor_probe_error_total{reason=~"unknown_key_id|decrypt_auth_failed"} is not rising.
  3. Keep the old key for at least the longest job or test TTL and until pull leases have drained. Purge or inspect checks.dead, which can hold payloads sealed with the old key.
  4. Remove the old entry from previous everywhere.

cerbix ships exactly two recovery commands. Both change persisted reliability evidence.

Service-reliability facts are partitioned by month; rows written before a month’s partition existed land in a DEFAULT partition. The scheduler adopts such months automatically. Use the command when cerbix_service_fact_maintenance_failing stays at 1 and the repeating ensure_service_fact_partitions_failed warning names a month that is over the automatic bound (100,000 rows) or still fails after you stopped the writers keeping it busy.

Terminal window
cerbix adopt-fact-month --config /etc/cerbix/config.yaml --month 2026-05 --timeout 10m

The copy phase takes no lock on the parent table. The final cutover holds the parent lock until commit, within --timeout (default 10m). An error rolls back and the next attempt resumes. Running it for an already attached month is a no-op success. Fact values do not change; only their physical partition does.

Terminal window
cerbix enqueue-service-repair --config /etc/cerbix/config.yaml \
--project <project-id> --service <service-id> \
--from 2026-06-01T00:00:00Z --to 2026-07-01T00:00:00Z

All flags are required; --to must be later than --from. The command only queues an audited admin repair, widened to whole buckets. The scheduler leader recomputes it and records before-and-after evidence for every sealed bucket that changes. Choose the range within raw heartbeat retention: where the raw evidence is gone, the repair stops in the terminal error state, the sealed facts stay as they were, and the subsystem reports wedged (below).

UNKNOWN means a clause the policy assigns block or warn could not be answered from the facts. The token, the network and the gate itself are fine. reasons[] names each unanswered clause:

Reason Meaning Action
seal_stale sealed_through lags more than the policy’s max_seal_lag_seconds. Check cerbix_service_watermark_lag_seconds and the stalled states below. Do not loosen the policy to hide it.
facts_stale A burn rule’s lease expired or it was never evaluated. Check the burn evaluator: cerbix_service_alert_lag_seconds{signal="burn"}. A new target reads facts_stale until its first evaluation.
never_sealed No fact sealed for the service yet. Wait for the first seal; check that the service has SLI members and a governing definition.
window_target_missing The policy’s window has no SLO target any more. Recreate the target or change the policy’s window.
no_objective The target has no objective. Set the objective.
budget_withheld The service page withholds the number for this window. Resolve the withholding reason shown on the service page.

never_evaluated and no_governing_revision can also appear in reasons[] as evidence; they do not decide the state. The policy’s unknown_behavior turns UNKNOWN into action WARN (CLI exit 0) or BLOCK (exit 2). NOT_CONFIGURED (exit 4) and a 503 (exit 1, nothing recorded) are not UNKNOWN. See Release gate.

Wedged. cerbix_service_wedged is 1 and the scheduler fails /readyz. The cause is a repair range parked in state error, usually because raw evidence for a sealed bucket is gone. It will not retry itself:

SELECT id, service_id, reason, last_error FROM service_repair_ranges WHERE state = 'error';

Resolve it deliberately: delete the row, or enqueue a narrower range. Readiness recovers on the next stats pass (within about 15 s).

Alerting evaluator stalled. When a pass of the live signal (every 30 s) or the burn signal (every 60 s) lags more than three cadences, the scheduler reports not ready. Coverage then dis-arms and member monitors page for themselves: noisier, not silent. The API stays ready. Restart the scheduler leader; a standby takes over and re-arms by evaluating.

Watermark lag is not a wedge. A newly declared service backfilling history lags heavily while progressing. Alert on a lag that grows without a backfill in flight.

Expression Severity Meaning
cerbix_service_wedged == 1 for 5m page Service reliability needs an operator.
cerbix_service_alert_lag_seconds{signal="health"} > 90 (or {signal="burn"} > 180) for 10m page The evaluator is stalled; members page for themselves.
time() - cerbix_service_alert_last_success_seconds > 300 ticket An evaluator arm keeps failing.
cerbix_service_repair_ranges{state="error"} > 0 for 15m ticket A repair range is parked.
cerbix_service_fact_maintenance_failing == 1 and last success older than 30 min ticket A fact month is stuck; see the recovery commands.
cerbix_gate_evaluate_errors_total{kind="ledger_unwritable"} > 0 page The gate ledger ran out of partitions; pipelines get no decision.
cerbix_gate_decisions_writable_horizon_seconds < 172800, paired with absent() ticket Gate ledger maintenance stopped.
cerbix_pull_agent_lag_seconds > 120 for 2m warning A pull region is not draining its queue.
cerbix_database_up == 0 page The process lost PostgreSQL; /readyz fails.

The repository also ships rules for audit retention, credential dispatch and Monitoring as Code. The metric catalogue is in Metrics and alerts.