Skip to content
v0.3.8GitHub

On-call and notifications

Verified against cerbix f5240f5Report a problem ↗

This page explains who gets told when something breaks: the channels cerbix delivers to, how escalation ladders and rotations pick a recipient, how a Service pages in place of its members, and how delivery survives restarts and failing endpoints.

Channels belong to a project. Create one with POST /api/v1/projects/{projectID}/notification-channels and link it to a monitor with POST /api/v1/monitors/{monitorID}/notifications and {"channel_id": …}.

type Required config keys Optional
webhook url (http or https) —
slack url (incoming webhook) —
telegram bot_token, chat_id —
email smtp_host, from, to (comma-separated) smtp_port (default 587), smtp_username, smtp_password

url, bot_token and smtp_password are returned blank on every read. When you edit a channel, a blank secret keeps the stored value.

Outbound delivery is guarded against private destinations. By default webhooks, Slack, Telegram and SMTP cannot reach loopback, private or cloud-metadata addresses. Set notification_egress.allow_private_ips or notification_egress.allow_metadata_ips only when an internal endpoint is required. See Configuration.

A monitor notifies its linked channels on any move to down, on a recovery from down to up, and on reminders every renotify_seconds while still down (0 turns reminders off).

down requires failure_threshold consecutive failed checks (minimum 1; push monitors always use 1). With confirm_interval_seconds set, cerbix probes faster between the first failure and the verdict, then returns to the normal interval. The value is clamped to 5 s up to the monitor interval, applies only when failure_threshold is above 1, and is ignored for push and composite monitors. 0 turns it off.

An escalation policy is an ordered list of steps in a project, created with POST /api/v1/projects/{projectID}/escalation-policies. Example data:

{
"name": "checkout-primary",
"repeat_last": true,
"steps": [
{ "after_seconds": 0, "targets": [{ "type": "schedule", "id": "<schedule-id>" }] },
{ "after_seconds": 900, "targets": [{ "type": "channel", "id": "<channel-id>" }] }
]
}

Every policy needs at least one step. after_seconds must be 0 or more and never decrease, and each step needs a target. Targets must belong to the same project. A schedule target resolves to whoever is on call when the step fires.

For a monitor, set escalation_policy_id and keep auto_incident on. The ladder then runs over the monitor’s open auto-incident: each step fires once its offset from the incident start has passed. The monitor’s flat down notification is dropped; recoveries still go to every linked channel. With repeat_last, the last step repeats every renotify_seconds until acknowledged; with renotify_seconds: 0, it does not repeat.

The ladder stops when the incident is acknowledged (POST /api/v1/incidents/{incidentID}/acknowledge) or resolved, or when the monitor is no longer down or is disabled. It pauses while a parent monitor in depends_on is down.

An on-call schedule rotates over notification channels. Create one with POST /api/v1/projects/{projectID}/oncall-schedules. Example data:

{
"name": "checkout-weekly",
"shift_seconds": 604800,
"anchor_at": "2026-10-05T09:00:00Z",
"participants": ["<channel-a>", "<channel-b>", "<channel-c>"]
}

At instant t, the participant at index floor((t − anchor_at) / shift_seconds) mod n is on call. The schedule has no time-zone field: express shift boundaries through anchor_at.

Cover a vacation with POST /api/v1/oncall-schedules/{scheduleID}/overrides and {"channel_id", "starts_at", "ends_at"}. While starts_at ≤ t < ends_at, the override channel is on call regardless of the rotation. GET /api/v1/oncall-schedules/{scheduleID}/current returns the channel on call now.

A Service can page in place of its member monitors. Declare it with PATCH /api/v1/projects/{projectID}/services/{serviceID}/alerting; omitted fields stay unchanged.

Field Default Meaning
owns_paging false The Service pages; its members stop delivering alerts while coverage is armed.
page_on ["down"] Subset of down, degraded. [] pages for no state.
page_on_unknown false Page when cerbix cannot see the Service.
confirm_evaluations 2 1–10 consecutive evaluations before an onset or close.
renotify_seconds 0 0 or 60–86,400; repeats the last escalation step.

The Service is evaluated every 30 s, so the nominal confirmation delay is confirm_evaluations × 30 s. An evaluator stall extends it.

Routing. An announcement goes to the channel on call in the Service’s on-call schedule (set as oncall_schedule_id when the Service is created) if that channel exists and is enabled. Otherwise it goes to every enabled channel of the project. An onset with nobody to tell is withheld and retried on the next evaluation (cerbix_service_alert_withheld_total). The close goes to the same recipients as the onset and states its reason, for example recovered, visibility_lost or entered_maintenance; losing sight of a Service is never announced as a recovery.

Escalation. PUT /api/v1/projects/{projectID}/services/{serviceID}/escalation-policy with {"escalation_policy_id": "<id>"} attaches a ladder; "" clears it. The ladder runs over the Service’s auto-incident and is frozen when that incident opens, so a policy attached later applies from the next incident. A step advances only while the Service still owns paging and its fresh live verdict is still firing; if that cannot be established, the ladder waits. Burn-rate alerts notify but never escalate.

How member alerts are muted, and why ambiguity always pages, is described in Paging suppression.

A global admin silences alerts with PUT /api/v1/settings/alerting:

{ "global_silence": { "enabled": true, "until": "2026-10-05T18:00:00Z" } }

until is optional; the silence ends on its own after it. Silence mutes monitor transitions, burn alerts, Service alerts, escalation steps and region-worker alerts. Incident webhooks, status-page subscriber emails and SLA reports are not muted. Incidents, statuses and escalation progress are recorded as usual.

Delivery: outbox, retries and dead letters

Section titled “Delivery: outbox, retries and dead letters”

Every notification is written to an outbox in the same database transaction as the state change, so a restart loses nothing. Workers claim due events every 2 s with row locks, so several replicas can run them safely.

  • At-least-once. One event can arrive twice. If any channel of an event fails, the whole event is retried, and channels that already succeeded receive it again.
  • Ordered per incident. An incident event is not released while an earlier event of the same incident is undelivered, including one in the dead-letter queue.
  • Stale transitions are dropped. A retried down that a newer transition has superseded is not delivered.
  • Retry backoff starts at 10 s, doubles per attempt and is capped at 1 h. HTTP requests time out after 10 s.
  • Dead letters. After 10 failed attempts an event is parked as dead and counted in cerbix_outbox_dead_total.

A global admin inspects and replays dead events:

Terminal window
curl -H "Authorization: Bearer $CERBIX_TOKEN" \
"https://cerbix.example.com/api/v1/admin/outbox/dead?limit=100"
curl -X POST -H "Authorization: Bearer $CERBIX_TOKEN" \
"https://cerbix.example.com/api/v1/admin/outbox/dead/<event-id>/replay"

POST /api/v1/admin/outbox/dead/replay-all requeues every dead event. Replay resets the attempt count. See the runbook and metrics and alerts.