Skip to content
v0.3.8GitHub

Incidents

Verified against cerbix f5240f5Report a problem ↗

An incident is a tracked disruption inside a project, with a status, an impact level and a timeline of updates. This page explains where incidents come from, what they are anchored to, how their lifecycle works, and how cerbix avoids paging twice for one outage.

Source Opened by Defaults
auto A monitor going down, or a Service that owns paging announcing an outage investigating, impact major
manual A person in the UI investigating, impact minor unless set
api An API token, including the Alertmanager receiver investigating, impact minor unless set

Monitor auto-incidents. When a monitor with auto_incident enabled transitions to down, cerbix opens <monitor> is down. No incident opens while the monitor is inside an active maintenance window. Any later transition into up resolves it with the update “Monitor recovered — automatically resolved.” At most one auto-incident is open per monitor.

Service auto-incidents. A Service with owns_paging enabled opens <service> — service <state> when its live signal announces an onset, confirmed over confirm_evaluations consecutive evaluations. A burn-rate breach never opens an incident. The incident resolves in the same transaction as the announcement that ends the outage. If paging ownership is turned off, the paging policy stops covering the state, or the Service is deleted, the incident resolves with a note that states this is not a recovery. An incident a person already resolved is left alone. At most one auto-incident is open per Service.

Manual and API incidents. POST /api/v1/projects/{projectID}/incidents (editor or above) takes title (required), status, impact, an opening Markdown body and an optional monitor_id from the same project. Only a Service’s own paging sets a Service anchor.

Every incident has at most one anchor, enforced by the database:

  • Monitor — monitor_id is set.
  • Service — service_id is set.
  • None — a project-level incident, for example a manual one or one opened by Alertmanager.

Deleting a Service clears the anchor. The incident stays in the project as a project-level record with its timeline intact. A Service incident does not change the incidents of its member monitors; both timelines remain true.

Status Meaning
investigating The issue is acknowledged; the cause is unknown.
identified The cause is understood; a fix is in progress.
monitoring A fix is applied and being watched.
resolved Over. Terminal.

The lifecycle only moves forward. An update with an earlier status is refused with 400, and an update against a resolved incident is refused. An update that omits status keeps whatever status the incident has when the update lands.

Impact is one of none, minor, major, critical.

Acknowledge with POST /api/v1/incidents/{incidentID}/acknowledge. It records who took ownership and stops the escalation ladder. A repeated acknowledgement changes nothing. A resolved incident answers 409. See On-call.

Post updates with POST /api/v1/incidents/{incidentID}/updates and a body of {"status": …, "body": …}. The system adds its own notes, each under a fixed prefix that a redelivered event never duplicates.

Prefix Written on Content
⚡ Context: Monitor auto-incidents Other monitors of the project that failed within ±5 min, the dominant error class, and the region when all failures share one.
🕸 Impact: Anchored incidents with impact links Probable-root and affected Services. One note per batch of new links.
⏸ Suppressed: A monitor’s open auto-incident Which parent monitor or owning Service muted its alerts.
🚀 Changes: Service auto-incidents Changes that preceded the incident. See Change intelligence.

Every write by a person or token adds one row to the organization audit log: incident.create, incident.status, incident.note, incident.acknowledge or incident.postmortem. Bodies never enter the audit log. Machine writes are recorded only in the timeline.

A postmortem attaches to a resolved incident with PUT /api/v1/incidents/{incidentID}/postmortem and {"body": …}. A second PUT replaces it. Public status pages show it under past incidents.

When a Service incident opens, cerbix snapshots the Service’s member monitors with their names and roles. The authenticated incident detail returns them as members and keeps naming a monitor after it has been deleted. members_unavailable: true means the snapshot could not be read; it is not an empty list.

Dependency impact: candidates, never a culprit

Section titled “Dependency impact: candidates, never a culprit”

When an anchored incident opens, cerbix walks the Service dependency graph from the incident’s own Services:

  • probable_root — every upstream Service on a path that has an open incident.
  • affected — every downstream Service that has an open incident.

Each link carries the canonical path, root first. The relation records candidates. It never elects a single cause, never links an incident to its own subject, and never suppresses or hides anything. Links appear as impacts on the authenticated detail (GET /api/v1/incidents/{incidentID}) only. impacts_unavailable: true means the read failed. Public status pages never carry impact links.

Suppression mutes delivery, never facts. Heartbeats, statuses, incidents and SLO history are recorded unchanged. Recoveries and burn-alert clears are never suppressed.

Mechanism What is muted Condition
Service ownership Member down alerts, reminders and escalation steps (live signal); member firing burn alerts (burn signal) The owning Service’s coverage for that signal is armed
Monitor depends_on The child’s down alerts and firing burn alerts A transitive parent is down or has an open auto-incident; the child’s escalation ladder also pauses while a parent is down
Maintenance window down alerts The monitor is inside an active window; a reminder after the window delivers if still down

Service coverage is armed per signal, never by owns_paging alone. The Service’s own evaluation must be fresh for its current configuration, its route must resolve to at least one recipient, and any outage it is in must already be announced and delivered. Read the state at GET /api/v1/projects/{projectID}/services/{serviceID}/alerting/state; reason names the first unmet condition.

POST /api/v1/projects/{projectID}/alerts/alertmanager accepts the Alertmanager webhook payload. Authenticate with an API token that has project write access (editor or above).

receivers:
- name: cerbix
webhook_configs:
- url: https://cerbix.example.com/api/v1/projects/<project-id>/alerts/alertmanager
send_resolved: true
http_config:
authorization:
credentials_file: /etc/alertmanager/cerbix-token
Alert field Effect
fingerprint Correlation key; falls back to labels.alertname. With neither, the alert is ignored.
status: firing Opens a project-level incident (source api) unless one is already open for the key.
status: resolved Resolves the open incident for the key with “Resolved by Alertmanager.” Unknown keys are ignored.
annotations.summary Title; falls back to labels.alertname, then “External alert”.
annotations.description Opening update.
labels.severity critical or page → critical; warning or major → major; anything else → minor.

The response counts opened, resolved and ignored alerts. Duplicate deliveries, including concurrent ones, are counted as ignored.

Lifecycle changes are sent to the organization’s webhooks as incident.opened, incident.updated or incident.resolved, with X-Cerbix-Event and X-Cerbix-Signature: sha256=<hex> (HMAC-SHA256 of the body with the webhook secret). Delivery is at-least-once: deduplicate and order on the incident id plus seq. Status page subscribers receive the same events by email.