Operations runbook
f5240f5Report a problem ↗These are the operator tasks you will meet most often: upgrades, backups, key rotation, the two recovery commands, and the reliability states that need a person. Each section is a short procedure. The full runbook on GitHub has the background and the rarer cases.
Upgrade cerbix
Section titled “Upgrade cerbix”Before every upgrade, take a database backup and read the release notes. Migrations are forward-only.
Single node. Replace the binary or image and start it. Migrations apply on startup.
# Docker Compose: bump CERBIX_IMAGE in docker/.env, thendocker compose --env-file docker/.env -f docker/docker-compose.prod.yml pull cerbixdocker compose --env-file docker/.env -f docker/docker-compose.prod.yml up -dDistributed. Make the schema change a deliberate step:
-
Stop every outbox owner: all
all,apiandschedulerprocesses. Workers and agents do not deliver notifications and can keep probing. -
Run the migration once with the new binary:
Terminal window cerbix migrate --config /etc/cerbix/config.yaml -
Start the new
scheduler, then the workers, thenapi.
Step 1 is required when an upgrade crosses migration 00088. Without it, an old process can deliver one already-claimed batch of an incident’s events out of order. Nothing is lost.
When a migration refuses an upgrade
Section titled “When a migration refuses an upgrade”Each refusal below stops before it changes the schema: the process exits with db_migrate_failed (or cerbix migrate exits non-zero), and the database stays on the migration before it. Nothing is lost.
| Refusal | Cause | Action |
|---|---|---|
| The server is older than PostgreSQL 15 | cerbix migrate checks the server version before applying anything. |
Upgrade PostgreSQL to 15 or newer (16 matches the repository images). There is no compatibility mode. |
Migration 00104: rows at protocol_version 4 in pull_tests |
No cerbix build writes such rows, so the migration will not delete them for you. | Inspect, then delete explicitly (below). |
Migration 00106: a cross-project alert-routing reference |
A monitor, escalation step or on-call participant points at a channel or schedule in another project. | Keep the old version running, find the rows with the read-only checks, correct them through the project API, then retry. |
SELECT id, region, protocol_version, created_at FROM pull_tests WHERE protocol_version = 4;-- once you know what wrote them:DELETE FROM pull_tests WHERE protocol_version = 4;Back up and restore
Section titled “Back up and restore”- PostgreSQL holds all cerbix state. The binary writes nothing to local disk. Back up with your usual PostgreSQL tooling, for example
pg_dump -U cerbix cerbix > cerbix.sql. - Back up the keys separately.
security.encryption_key(and anysecurity.previous_keys) decrypts stored channel credentials, webhook secrets, TOTP secrets and project secrets. Without it a restored database cannot use them. Keep the regionalsecurity.dispatchkeys with it. - RabbitMQ holds only jobs and results in flight. The scheduler re-issues expired jobs on its next tick. For broker upgrades, follow the staged RabbitMQ procedure: RabbitMQ supports neither downgrades nor a direct 3.12 to 4.3 jump.
Rotate the secrets-at-rest key
Section titled “Rotate the secrets-at-rest key”The at-rest key lives on core roles (all, api, scheduler) only.
-
On every core role, set the new key as
security.encryption_keyand move the old one intosecurity.previous_keys. -
Roll the core roles. Readers decrypt old rows; writers use the new key.
-
Re-encrypt stored secrets:
Terminal window cerbix reencrypt --config /etc/cerbix/config.yamlExit
0means no project secret remains under an old key; webhook and channel secrets are rewritten under the new key too. A non-zero exit logs the failure. If it reports that re-encryption did not converge (a concurrent secret rotation), run it again. -
Remove the old key from
security.previous_keysonly after the command succeeds and every core replica runs the new config.
A key is openssl rand -base64 32: a base64-encoded 32-byte value.
Rotate a dispatch key
Section titled “Rotate a dispatch key”reencrypt does not touch job payloads in RabbitMQ or the pull queue. For one region:
- Deploy
{primary: <new>, previous: [<old>]}for that region to the core roles and to the region’s workers or agents. Executors need the overlap before core starts sealing with the new key. - Check executor
/readyzand confirm thatcerbix_executor_probe_error_total{reason=~"unknown_key_id|decrypt_auth_failed"}is not rising. - Keep the old key for at least the longest job or test TTL and until pull leases have drained. Purge or inspect
checks.dead, which can hold payloads sealed with the old key. - Remove the old entry from
previouseverywhere.
Recovery commands
Section titled “Recovery commands”cerbix ships exactly two recovery commands. Both change persisted reliability evidence.
Adopt a stranded fact month
Section titled “Adopt a stranded fact month”Service-reliability facts are partitioned by month; rows written before a month’s partition existed land in a DEFAULT partition. The scheduler adopts such months automatically. Use the command when cerbix_service_fact_maintenance_failing stays at 1 and the repeating ensure_service_fact_partitions_failed warning names a month that is over the automatic bound (100,000 rows) or still fails after you stopped the writers keeping it busy.
cerbix adopt-fact-month --config /etc/cerbix/config.yaml --month 2026-05 --timeout 10mThe copy phase takes no lock on the parent table. The final cutover holds the parent lock until commit, within --timeout (default 10m). An error rolls back and the next attempt resumes. Running it for an already attached month is a no-op success. Fact values do not change; only their physical partition does.
Restate a range of service facts
Section titled “Restate a range of service facts”cerbix enqueue-service-repair --config /etc/cerbix/config.yaml \ --project <project-id> --service <service-id> \ --from 2026-06-01T00:00:00Z --to 2026-07-01T00:00:00ZAll flags are required; --to must be later than --from. The command only queues an audited admin repair, widened to whole buckets. The scheduler leader recomputes it and records before-and-after evidence for every sealed bucket that changes. Choose the range within raw heartbeat retention: where the raw evidence is gone, the repair stops in the terminal error state, the sealed facts stay as they were, and the subsystem reports wedged (below).
The gate says UNKNOWN
Section titled “The gate says UNKNOWN”UNKNOWN means a clause the policy assigns block or warn could not be answered from the facts. The token, the network and the gate itself are fine. reasons[] names each unanswered clause:
| Reason | Meaning | Action |
|---|---|---|
seal_stale |
sealed_through lags more than the policy’s max_seal_lag_seconds. |
Check cerbix_service_watermark_lag_seconds and the stalled states below. Do not loosen the policy to hide it. |
facts_stale |
A burn rule’s lease expired or it was never evaluated. | Check the burn evaluator: cerbix_service_alert_lag_seconds{signal="burn"}. A new target reads facts_stale until its first evaluation. |
never_sealed |
No fact sealed for the service yet. | Wait for the first seal; check that the service has SLI members and a governing definition. |
window_target_missing |
The policy’s window has no SLO target any more. | Recreate the target or change the policy’s window. |
no_objective |
The target has no objective. | Set the objective. |
budget_withheld |
The service page withholds the number for this window. | Resolve the withholding reason shown on the service page. |
never_evaluated and no_governing_revision can also appear in reasons[] as evidence; they do not decide the state. The policy’s unknown_behavior turns UNKNOWN into action WARN (CLI exit 0) or BLOCK (exit 2). NOT_CONFIGURED (exit 4) and a 503 (exit 1, nothing recorded) are not UNKNOWN. See Release gate.
Service reliability is wedged or stalled
Section titled “Service reliability is wedged or stalled”Wedged. cerbix_service_wedged is 1 and the scheduler fails /readyz. The cause is a repair range parked in state error, usually because raw evidence for a sealed bucket is gone. It will not retry itself:
SELECT id, service_id, reason, last_error FROM service_repair_ranges WHERE state = 'error';Resolve it deliberately: delete the row, or enqueue a narrower range. Readiness recovers on the next stats pass (within about 15 s).
Alerting evaluator stalled. When a pass of the live signal (every 30 s) or the burn signal (every 60 s) lags more than three cadences, the scheduler reports not ready. Coverage then dis-arms and member monitors page for themselves: noisier, not silent. The API stays ready. Restart the scheduler leader; a standby takes over and re-arms by evaluating.
Watermark lag is not a wedge. A newly declared service backfilling history lags heavily while progressing. Alert on a lag that grows without a backfill in flight.
Suggested alerts
Section titled “Suggested alerts”| Expression | Severity | Meaning |
|---|---|---|
cerbix_service_wedged == 1 for 5m |
page | Service reliability needs an operator. |
cerbix_service_alert_lag_seconds{signal="health"} > 90 (or {signal="burn"} > 180) for 10m |
page | The evaluator is stalled; members page for themselves. |
time() - cerbix_service_alert_last_success_seconds > 300 |
ticket | An evaluator arm keeps failing. |
cerbix_service_repair_ranges{state="error"} > 0 for 15m |
ticket | A repair range is parked. |
cerbix_service_fact_maintenance_failing == 1 and last success older than 30 min |
ticket | A fact month is stuck; see the recovery commands. |
cerbix_gate_evaluate_errors_total{kind="ledger_unwritable"} > 0 |
page | The gate ledger ran out of partitions; pipelines get no decision. |
cerbix_gate_decisions_writable_horizon_seconds < 172800, paired with absent() |
ticket | Gate ledger maintenance stopped. |
cerbix_pull_agent_lag_seconds > 120 for 2m |
warning | A pull region is not draining its queue. |
cerbix_database_up == 0 |
page | The process lost PostgreSQL; /readyz fails. |
The repository also ships rules for audit retention, credential dispatch and Monitoring as Code. The metric catalogue is in Metrics and alerts.