SLOs, error budgets and burn rate
f5240f5Report a problem ↗An objective turns a measured availability into a budget you can spend. This page explains how objectives are declared, which formulas produce the error budget and burn rate, and how burn-rate alert rules decide when to notify.
Objectives per window
Section titled “Objectives per window”An objective is a target availability in percent for one standard rolling window:
24h, 7d, 30d or 90d. Each scope holds at most one objective per window.
| Scope | Endpoint | Can alert on burn |
|---|---|---|
| Monitor | PUT /api/v1/monitors/{monitorID}/sla-target |
yes |
| Project | PUT /api/v1/projects/{projectID}/sla-target |
no |
| Service | PUT /api/v1/projects/{projectID}/services/{serviceID}/sla-target |
yes, through a separate endpoint |
window defaults to 30d. The objective must lie strictly between 0 and 100. cerbix rounds it
half-up to four decimal places and checks the rounded value again, so the largest accepted
objective is 99.9999. A value of 100 is rejected: a zero error budget is not a supported
configuration. The response echoes the stored value.
A monitor or project window longer than the raw heartbeat retention (heartbeats.retention_days, default 30)
still covers its whole length: days whose raw heartbeats were purged are read from the daily rollup. When no
data older than the window’s start exists, the numbers cover only what exists and the response adds
data_from, the UTC day the data starts; the UI marks the window since DD.MM.YYYY UTC. Latency comes from
raw heartbeats only, and latency_from says where it starts when that is later.
A Service objective is not part of its definition revision. Changing it re-evaluates the budget
over the same facts immediately. Every report states the objective it used and
objective_updated_at. A Service report never borrows an objective from a monitor or project.
What a window measures
Section titled “What a window measures”- Monitors measure the share of checks that were
upover the window ending now. Checks inside a maintenance window are excluded from both sides. - Services measure decidable time from sealed facts over
[sealed_through − window, sealed_through). When coverage or storage is insufficient, the number is withheld with a reason. See How the SLI is computed.
Error budget
Section titled “Error budget”With objective as a percentage and availability as the measured percentage:
| Field | Formula | Notes |
|---|---|---|
allowed_downtime_ratio |
1 − objective / 100 |
the permitted bad fraction |
actual_downtime_ratio |
1 − availability / 100 |
the measured bad fraction |
remaining_ratio |
allowed − actual |
negative once the budget is exceeded |
burned_percent |
actual / allowed × 100 |
can exceed 100 |
met |
availability ≥ objective |
false when nothing was measured |
Example data: an objective of 99.9 over 30d allows a bad fraction of 0.001. If the whole
window was measured, that is 2,592 s, or 43.2 min, of bad time.
A Service budget appears only when the window’s availability can be quoted and an objective
exists for that window. The Service page shows the remaining budget as 100% − burned_percent.
Burn rate
Section titled “Burn rate”burn_rate = actual bad fraction in a window / allowed_downtime_ratioAt 1× the budget is used up exactly at the end of the objective’s window. At 14.4× a 30d
budget is gone in about two days. A window with nothing measured never yields 0×; it yields no
rate.
The informational 1h/6h card
Section titled “The informational 1h/6h card”The Service page shows burn rate over two fixed windows, 1h and 6h, each ending at
sealed_through, against the objective of the selected window. The card never notifies anyone.
Each window carries its own status:
| Reason | Status | Rate shown |
|---|---|---|
| — | ok |
yes |
decidable_coverage_below_min |
partial |
yes, with coverage |
storage_gap |
partial |
no |
spans_definition_revisions |
unavailable |
no |
zero_decidable_time |
unavailable |
no |
window_precedes_materialization_era |
insufficient_history |
no |
sealed_through_behind_window |
insufficient_sealed_coverage |
no |
Burn-rate alert rules
Section titled “Burn-rate alert rules”A rule pairs a long window, which filters noise, with a short window, which confirms the burn is
still happening. It fires when the burn rate is at or above threshold in both windows.
| Field | Rule |
|---|---|
long_window_seconds |
greater than the short window; at most 7 days (604,800 s) |
short_window_seconds |
at least 60 s for a monitor; at least 300 s for a Service |
threshold |
a finite number greater than 0 |
severity |
page or ticket |
A Service window ends at the seal watermark, which trails real time by the 120 s late-arrival grace plus up to
a bucket. A shorter short window could never, or only intermittently, contain sealed time, so the API refuses a
Service rule below 300 s with 400.
A target holds at most 4 rules. Two rules with the same severity, windows and threshold are
rejected. Each rule has its own server-owned latch: one notification when it starts firing and
one when it clears, not one per evaluation. severity is carried in the notification and does
not change routing.
Monitor targets
Section titled “Monitor targets”Set burn_alert: true on the monitor’s SLA target. If you send no burn_rules, cerbix keeps the
existing rules or, on a target without rules, seeds the default pair:
| Severity | Threshold | Long window | Short window |
|---|---|---|---|
page |
14.4× | 1 h | 5 min |
ticket |
6× | 6 h | 30 min |
Sending burn_rules replaces the set. A latch survives an edit for any rule whose configuration
did not change. Windows are computed from checks ending now, with maintenance excluded. A window
with no checks makes no decision and the latch stays as it was. Notifications go to the
monitor’s channels.
Service targets
Section titled “Service targets”Service rules are declared on their own endpoint:
curl -X PUT "$CERBIX_URL/api/v1/projects/<project-id>/services/<service-id>/sla-target/burn-alerting" \ -H "Authorization: Bearer $CERBIX_TOKEN" -H "Content-Type: application/json" \ -d '{"window":"30d","burn_alert_enabled":true, "burn_rules":[{"long_window_seconds":3600,"short_window_seconds":300,"threshold":14.4,"severity":"page"}]}'- The body is a full replacement. Omitting
burn_rulesdeclares none. - No rules are seeded. A target with no rules never fires.
- The window must already have an objective, otherwise the call returns
404. - A file-managed Service refuses this write with
409 managed_by_file. - Rules are evaluated only while the Service owns paging (
owns_paging). See On-call.
Both windows are computed from sealed facts ending at sealed_through, so this signal trails the
watermark. Each evaluation produces one verdict per rule:
| Verdict | When | Effect |
|---|---|---|
fire |
both windows quotable and at or above the threshold | the rule fires |
clear |
both windows quotable and at least one below the threshold | the rule clears |
hold |
at least one window cannot be quoted (any reason in the table above, including low coverage) | the previous level is kept and the reason recorded |
A Service burn alert notifies; it does not open an incident. Removing a firing rule or disabling
burn alerting closes the open alert with reason rule_removed or burn_disabled.
Weekly SLA reports
Section titled “Weekly SLA reports”PUT /api/v1/projects/{projectID}/sla-report with {"enabled": true} turns on a weekly report
for a project. The scheduler leader checks every hour and sends a report when 7 days have passed
since the last one. The report goes to the project’s notification channels as text: the share of
up checks across all of the project’s monitors for 7d and 30d, with maintenance excluded,
or no data when there were no checks. It is computed from monitor checks, not from Service
facts.