Incidents
f5240f5Report a problem ↗An incident is a tracked disruption inside a project, with a status, an impact level and a timeline of updates. This page explains where incidents come from, what they are anchored to, how their lifecycle works, and how cerbix avoids paging twice for one outage.
How an incident opens
Section titled “How an incident opens”| Source | Opened by | Defaults |
|---|---|---|
auto |
A monitor going down, or a Service that owns paging announcing an outage | investigating, impact major |
manual |
A person in the UI | investigating, impact minor unless set |
api |
An API token, including the Alertmanager receiver | investigating, impact minor unless set |
Monitor auto-incidents. When a monitor with auto_incident enabled transitions to down, cerbix opens <monitor> is down. No incident opens while the monitor is inside an active maintenance window. Any later transition into up resolves it with the update “Monitor recovered — automatically resolved.” At most one auto-incident is open per monitor.
Service auto-incidents. A Service with owns_paging enabled opens <service> — service <state> when its live signal announces an onset, confirmed over confirm_evaluations consecutive evaluations. A burn-rate breach never opens an incident. The incident resolves in the same transaction as the announcement that ends the outage. If paging ownership is turned off, the paging policy stops covering the state, or the Service is deleted, the incident resolves with a note that states this is not a recovery. An incident a person already resolved is left alone. At most one auto-incident is open per Service.
Manual and API incidents. POST /api/v1/projects/{projectID}/incidents (editor or above) takes title (required), status, impact, an opening Markdown body and an optional monitor_id from the same project. Only a Service’s own paging sets a Service anchor.
Anchors
Section titled “Anchors”Every incident has at most one anchor, enforced by the database:
- Monitor —
monitor_idis set. - Service —
service_idis set. - None — a project-level incident, for example a manual one or one opened by Alertmanager.
Deleting a Service clears the anchor. The incident stays in the project as a project-level record with its timeline intact. A Service incident does not change the incidents of its member monitors; both timelines remain true.
Lifecycle and impact
Section titled “Lifecycle and impact”| Status | Meaning |
|---|---|
investigating |
The issue is acknowledged; the cause is unknown. |
identified |
The cause is understood; a fix is in progress. |
monitoring |
A fix is applied and being watched. |
resolved |
Over. Terminal. |
The lifecycle only moves forward. An update with an earlier status is refused with 400, and an update against a resolved incident is refused. An update that omits status keeps whatever status the incident has when the update lands.
Impact is one of none, minor, major, critical.
Acknowledge with POST /api/v1/incidents/{incidentID}/acknowledge. It records who took ownership and stops the escalation ladder. A repeated acknowledgement changes nothing. A resolved incident answers 409. See On-call.
Timeline and system notes
Section titled “Timeline and system notes”Post updates with POST /api/v1/incidents/{incidentID}/updates and a body of {"status": …, "body": …}. The system adds its own notes, each under a fixed prefix that a redelivered event never duplicates.
| Prefix | Written on | Content |
|---|---|---|
⚡ Context: |
Monitor auto-incidents | Other monitors of the project that failed within ±5 min, the dominant error class, and the region when all failures share one. |
🕸 Impact: |
Anchored incidents with impact links | Probable-root and affected Services. One note per batch of new links. |
⏸ Suppressed: |
A monitor’s open auto-incident | Which parent monitor or owning Service muted its alerts. |
🚀 Changes: |
Service auto-incidents | Changes that preceded the incident. See Change intelligence. |
Every write by a person or token adds one row to the organization audit log: incident.create, incident.status, incident.note, incident.acknowledge or incident.postmortem. Bodies never enter the audit log. Machine writes are recorded only in the timeline.
Postmortems and member snapshots
Section titled “Postmortems and member snapshots”A postmortem attaches to a resolved incident with PUT /api/v1/incidents/{incidentID}/postmortem and {"body": …}. A second PUT replaces it. Public status pages show it under past incidents.
When a Service incident opens, cerbix snapshots the Service’s member monitors with their names and roles. The authenticated incident detail returns them as members and keeps naming a monitor after it has been deleted. members_unavailable: true means the snapshot could not be read; it is not an empty list.
Dependency impact: candidates, never a culprit
Section titled “Dependency impact: candidates, never a culprit”When an anchored incident opens, cerbix walks the Service dependency graph from the incident’s own Services:
probable_root— every upstream Service on a path that has an open incident.affected— every downstream Service that has an open incident.
Each link carries the canonical path, root first. The relation records candidates. It never elects a single cause, never links an incident to its own subject, and never suppresses or hides anything. Links appear as impacts on the authenticated detail (GET /api/v1/incidents/{incidentID}) only. impacts_unavailable: true means the read failed. Public status pages never carry impact links.
Paging suppression
Section titled “Paging suppression”Suppression mutes delivery, never facts. Heartbeats, statuses, incidents and SLO history are recorded unchanged. Recoveries and burn-alert clears are never suppressed.
| Mechanism | What is muted | Condition |
|---|---|---|
| Service ownership | Member down alerts, reminders and escalation steps (live signal); member firing burn alerts (burn signal) |
The owning Service’s coverage for that signal is armed |
Monitor depends_on |
The child’s down alerts and firing burn alerts |
A transitive parent is down or has an open auto-incident; the child’s escalation ladder also pauses while a parent is down |
| Maintenance window | down alerts |
The monitor is inside an active window; a reminder after the window delivers if still down |
Service coverage is armed per signal, never by owns_paging alone. The Service’s own evaluation must be fresh for its current configuration, its route must resolve to at least one recipient, and any outage it is in must already be announced and delivered. Read the state at GET /api/v1/projects/{projectID}/services/{serviceID}/alerting/state; reason names the first unmet condition.
Alertmanager webhook
Section titled “Alertmanager webhook”POST /api/v1/projects/{projectID}/alerts/alertmanager accepts the Alertmanager webhook payload. Authenticate with an API token that has project write access (editor or above).
receivers: - name: cerbix webhook_configs: - url: https://cerbix.example.com/api/v1/projects/<project-id>/alerts/alertmanager send_resolved: true http_config: authorization: credentials_file: /etc/alertmanager/cerbix-token| Alert field | Effect |
|---|---|
fingerprint |
Correlation key; falls back to labels.alertname. With neither, the alert is ignored. |
status: firing |
Opens a project-level incident (source api) unless one is already open for the key. |
status: resolved |
Resolves the open incident for the key with “Resolved by Alertmanager.” Unknown keys are ignored. |
annotations.summary |
Title; falls back to labels.alertname, then “External alert”. |
annotations.description |
Opening update. |
labels.severity |
critical or page → critical; warning or major → major; anything else → minor. |
The response counts opened, resolved and ignored alerts. Duplicate deliveries, including concurrent ones, are counted as ignored.
Events that leave cerbix
Section titled “Events that leave cerbix”Lifecycle changes are sent to the organization’s webhooks as incident.opened, incident.updated or incident.resolved, with X-Cerbix-Event and X-Cerbix-Signature: sha256=<hex> (HMAC-SHA256 of the body with the webhook secret). Delivery is at-least-once: deduplicate and order on the incident id plus seq. Status page subscribers receive the same events by email.