Skip to content
v0.3.8GitHub

SLOs, error budgets and burn rate

Verified against cerbix f5240f5Report a problem ↗

An objective turns a measured availability into a budget you can spend. This page explains how objectives are declared, which formulas produce the error budget and burn rate, and how burn-rate alert rules decide when to notify.

An objective is a target availability in percent for one standard rolling window: 24h, 7d, 30d or 90d. Each scope holds at most one objective per window.

Scope Endpoint Can alert on burn
Monitor PUT /api/v1/monitors/{monitorID}/sla-target yes
Project PUT /api/v1/projects/{projectID}/sla-target no
Service PUT /api/v1/projects/{projectID}/services/{serviceID}/sla-target yes, through a separate endpoint

window defaults to 30d. The objective must lie strictly between 0 and 100. cerbix rounds it half-up to four decimal places and checks the rounded value again, so the largest accepted objective is 99.9999. A value of 100 is rejected: a zero error budget is not a supported configuration. The response echoes the stored value.

A monitor or project window longer than the raw heartbeat retention (heartbeats.retention_days, default 30) still covers its whole length: days whose raw heartbeats were purged are read from the daily rollup. When no data older than the window’s start exists, the numbers cover only what exists and the response adds data_from, the UTC day the data starts; the UI marks the window since DD.MM.YYYY UTC. Latency comes from raw heartbeats only, and latency_from says where it starts when that is later.

A Service objective is not part of its definition revision. Changing it re-evaluates the budget over the same facts immediately. Every report states the objective it used and objective_updated_at. A Service report never borrows an objective from a monitor or project.

  • Monitors measure the share of checks that were up over the window ending now. Checks inside a maintenance window are excluded from both sides.
  • Services measure decidable time from sealed facts over [sealed_through − window, sealed_through). When coverage or storage is insufficient, the number is withheld with a reason. See How the SLI is computed.

With objective as a percentage and availability as the measured percentage:

Field Formula Notes
allowed_downtime_ratio 1 − objective / 100 the permitted bad fraction
actual_downtime_ratio 1 − availability / 100 the measured bad fraction
remaining_ratio allowed − actual negative once the budget is exceeded
burned_percent actual / allowed × 100 can exceed 100
met availability ≥ objective false when nothing was measured

Example data: an objective of 99.9 over 30d allows a bad fraction of 0.001. If the whole window was measured, that is 2,592 s, or 43.2 min, of bad time.

A Service budget appears only when the window’s availability can be quoted and an objective exists for that window. The Service page shows the remaining budget as 100% − burned_percent.

burn_rate = actual bad fraction in a window / allowed_downtime_ratio

At 1× the budget is used up exactly at the end of the objective’s window. At 14.4× a 30d budget is gone in about two days. A window with nothing measured never yields 0×; it yields no rate.

The Service page shows burn rate over two fixed windows, 1h and 6h, each ending at sealed_through, against the objective of the selected window. The card never notifies anyone. Each window carries its own status:

Reason Status Rate shown
— ok yes
decidable_coverage_below_min partial yes, with coverage
storage_gap partial no
spans_definition_revisions unavailable no
zero_decidable_time unavailable no
window_precedes_materialization_era insufficient_history no
sealed_through_behind_window insufficient_sealed_coverage no

A rule pairs a long window, which filters noise, with a short window, which confirms the burn is still happening. It fires when the burn rate is at or above threshold in both windows.

Field Rule
long_window_seconds greater than the short window; at most 7 days (604,800 s)
short_window_seconds at least 60 s for a monitor; at least 300 s for a Service
threshold a finite number greater than 0
severity page or ticket

A Service window ends at the seal watermark, which trails real time by the 120 s late-arrival grace plus up to a bucket. A shorter short window could never, or only intermittently, contain sealed time, so the API refuses a Service rule below 300 s with 400.

A target holds at most 4 rules. Two rules with the same severity, windows and threshold are rejected. Each rule has its own server-owned latch: one notification when it starts firing and one when it clears, not one per evaluation. severity is carried in the notification and does not change routing.

Set burn_alert: true on the monitor’s SLA target. If you send no burn_rules, cerbix keeps the existing rules or, on a target without rules, seeds the default pair:

Severity Threshold Long window Short window
page 14.4× 1 h 5 min
ticket 6× 6 h 30 min

Sending burn_rules replaces the set. A latch survives an edit for any rule whose configuration did not change. Windows are computed from checks ending now, with maintenance excluded. A window with no checks makes no decision and the latch stays as it was. Notifications go to the monitor’s channels.

Service rules are declared on their own endpoint:

Terminal window
curl -X PUT "$CERBIX_URL/api/v1/projects/<project-id>/services/<service-id>/sla-target/burn-alerting" \
-H "Authorization: Bearer $CERBIX_TOKEN" -H "Content-Type: application/json" \
-d '{"window":"30d","burn_alert_enabled":true,
"burn_rules":[{"long_window_seconds":3600,"short_window_seconds":300,"threshold":14.4,"severity":"page"}]}'
  • The body is a full replacement. Omitting burn_rules declares none.
  • No rules are seeded. A target with no rules never fires.
  • The window must already have an objective, otherwise the call returns 404.
  • A file-managed Service refuses this write with 409 managed_by_file.
  • Rules are evaluated only while the Service owns paging (owns_paging). See On-call.

Both windows are computed from sealed facts ending at sealed_through, so this signal trails the watermark. Each evaluation produces one verdict per rule:

Verdict When Effect
fire both windows quotable and at or above the threshold the rule fires
clear both windows quotable and at least one below the threshold the rule clears
hold at least one window cannot be quoted (any reason in the table above, including low coverage) the previous level is kept and the reason recorded

A Service burn alert notifies; it does not open an incident. Removing a firing rule or disabling burn alerting closes the open alert with reason rule_removed or burn_disabled.

PUT /api/v1/projects/{projectID}/sla-report with {"enabled": true} turns on a weekly report for a project. The scheduler leader checks every hour and sends a report when 7 days have passed since the last one. The report goes to the project’s notification channels as text: the share of up checks across all of the project’s monitors for 7d and 30d, with maintenance excluded, or no data when there were no checks. It is computed from monitor checks, not from Service facts.