How the SLI is computed
src/content/docs/docs/concepts/sli.md · verified against cerbix f5240f5Edit on GitHub ↗A Service’s SLI is the share of decidable time during which it was available. cerbix measures it from its own checks, in one-minute buckets that are sealed before they are reported. This page follows one heartbeat to the number on the service page, and lists every case in which cerbix withholds that number instead.
At a glance
Section titled “At a glance”availability = good / (good + bad) × 100 # absent when good + bad = 0coverage = (good + bad) / (good + bad + unknown)budget burned = (bad / (good + bad)) / (1 − objective / 100) × 100Every term is a duration summed over sealed one-minute buckets in the window. excluded time —
maintenance and disabled members — appears in no denominator. Below 95% coverage the result is
partial; with nothing decidable it is absent, never 100%.
The path is always the same:
- Heartbeat — the raw
up/downresult of a check. - Member state —
GOOD,BAD,UNKNOWNorEXCLUDEDat every instant. - Region, then Service — members are combined with the Service’s aggregation policy.
- Sealed bucket — one minute, duration-weighted, sealed two minutes after it closes.
- Window report — a window that ends at the seal watermark, never at now.
1. The state of one member at an instant
Section titled “1. The state of one member at an instant”For every SLI member and every instant t, cerbix applies these rules in order. The first match wins.
| # | Condition | Active probe | Push |
|---|---|---|---|
| 1 | Member disabled in this evaluation epoch | EXCLUDED |
EXCLUDED |
| 2 | Inside a maintenance window, with maintenance: exclude (default) |
EXCLUDED |
EXCLUDED |
| 3 | No observation yet | UNKNOWN (no_observation) |
BAD once armed + stale_after has passed, UNKNOWN before |
| 4 | Last observation is at least stale_after old |
UNKNOWN (stale) |
BAD |
| 5 | Last observation is up |
GOOD |
GOOD |
| 6 | Last observation is down |
BAD |
BAD |
Silence means different things by check type. For an active probe a missing result is uncertainty. For a push monitor it is the failure the dead-man’s switch exists to catch.
How long an observation holds
Section titled “How long an observation holds”An observation stays in force until the next one arrives or until stale_after runs out, whichever
comes first. The bound is fixed per evaluation epoch:
stale_after(active) = max(3 × interval, 90 s)stale_after(push) = interval + grace| Check interval | Result holds for |
|---|---|
| 10 s | 90 s |
| 30 s | 90 s |
| 60 s | 180 s |
| 5 min | 15 min |
The Service’s missing_data policy then decides what an UNKNOWN member means:
unknown(default) keeps itUNKNOWN;badcounts it asBAD;ignoreremoves it from the count. Ignoring never turns a member intoEXCLUDED: if nothing is left to count, the result isUNKNOWN.
2. From members to one service state
Section titled “2. From members to one service state”Only members declared in sli[] count. Context members (monitors[]) are shown for diagnosis and never
change the number. The order is fixed: normalize → exclude → apply missing_data → combine within each
region → combine across regions.
Within a region, with n eligible members of which g are GOOD, b are BAD and u are UNKNOWN:
| Mode | GOOD when |
UNKNOWN when |
BAD when |
|---|---|---|---|
all (default) |
b = 0 and u = 0 |
b = 0 and u > 0 |
b > 0 |
any |
g > 0 |
g = 0 and u > 0 |
b = n |
quorum with threshold k |
g ≥ k |
g < k ≤ g + u |
g + u < k |
A quorum is BAD only when it would fail even if every UNKNOWN member turned out GOOD. For a quorum
of 2 out of 3 members:
| Members | Result |
|---|---|
GOOD GOOD UNKNOWN |
GOOD |
GOOD UNKNOWN BAD |
UNKNOWN |
UNKNOWN UNKNOWN BAD |
UNKNOWN |
GOOD BAD BAD |
BAD |
If exclusions leave fewer than k eligible members, the threshold is clamped to n and the bucket is
marked weakened, so the relaxed rule stays visible.
Regions. Each region is decided first, then the regions are combined with the same quorum logic.
A region that reports nothing is UNKNOWN, never dropped. With the default per_region policy, one
region is enough to be available and all regions are needed to be healthy, so a single dark region makes
the Service DEGRADED. DEGRADED still counts as available: health is shown beside the SLI, and only
availability feeds the SLO. With all_regions, every region must be decided GOOD; one dark region makes
the Service UNKNOWN for that time.
3. Duration weighting inside a one-minute bucket
Section titled “3. Duration weighting inside a one-minute bucket”Time is cut into one-minute buckets, half-open [start, end) and aligned to UTC. Inside a bucket the
Service state can change only at a breakpoint: a bucket edge, an observation, an observation going
stale, or a maintenance edge. Each piece between breakpoints goes whole into one bin. Durations are
stored in integer microseconds, and every bucket satisfies:
good + bad + unknown + excluded = 60 sFor example, a bucket that starts GOOD, turns BAD when a check fails, becomes UNKNOWN when a
required member’s last result goes stale, and enters a maintenance window before it closes might record
22 s good, 9 s bad, 10 s unknown and 19 s excluded (example data).
4. Sealing and the watermark
Section titled “4. Sealing and the watermark”- A bucket
[start, end)is sealed oncenow ≥ end + 2 min. The two minutes are the late-arrival grace. sealed_throughis the latest instant up to which every bucket exists and is sealed. A missing bucket holds the watermark back; it is never skipped.- Every window is
[sealed_through − W, sealed_through). It never ends at now, so the same window gives the same number every time it is read. Normal lag is two to three minutes. - A result that arrives for an already sealed bucket triggers an audited recompute of that bucket.
- The live state on the service page is a separate point-in-time evaluation. It is never mixed into the SLI.
5. Windows and formulas
Section titled “5. Windows and formulas”Reports are available for 24h, 7d, 30d (default) and 90d. Over the sealed buckets in a window:
availability = good / (good + bad) × 100 # absent when good + bad = 0coverage = (good + bad) / (good + bad + unknown)# excluded time is in neither denominatorUNKNOWN time does not count for or against availability. It lowers coverage instead, which decides
whether the number may be quoted.
6. When the number is withheld
Section titled “6. When the number is withheld”Each window gets one status. The first matching row wins.
| Condition | Status · reason | Number shown |
|---|---|---|
| The window starts before the Service was first measured | insufficient_history |
No |
| A bucket in the window is missing | partial · storage_gap |
No |
| No decidable time at all | unavailable · zero_decidable_time |
No |
| Coverage below 95% | partial · decidable_coverage_below_min |
Yes, labelled partial |
| Otherwise | ok |
Yes |
Independently of the status, a window whose facts span more than one definition revision has its
aggregate withheld (spans_definition_revisions). It is reported as segments, one per revision and
evaluation epoch, each with its own numbers. A Service with no SLI members reports no_sli; one with
nothing sealed yet reports nothing_sealed. The 95% threshold is fixed, not a setting.
7. Error budget and burn rate
Section titled “7. Error budget and burn rate”An objective is set per Service and per window, between 0 and 100 exclusive, up to 99.9999. A budget
is computed only when the availability can be quoted.
allowed = 1 − objective / 100actual = bad / (good + bad)burned % = actual / allowed × 100budget left = 100 − burned %burn rate = actual(w) / allowed # over a burn window wThe budget is a share of measured time, not of wall-clock time: maintenance and UNKNOWN minutes
neither spend it nor earn it. Burn-rate alert rules and their FIRE / CLEAR / HOLD semantics are
described in SLOs, error budgets and burn rate.
8. Worked example
Section titled “8. Worked example”A 30d window (2,592,000 s, all buckets sealed, one definition revision) with an objective of 99.9
(example data):
| Input | Duration |
|---|---|
good |
2,581,800 s |
bad |
1,200 s (20 min) |
unknown |
1,800 s (30 min) |
excluded |
7,200 s (2 h maintenance) |
| Result | Value |
|---|---|
| availability | 2,581,800 / 2,583,000 = 99.9535% |
| coverage | 2,583,000 / 2,584,800 = 99.93% → ok |
| budget | 0.001 × 2,583,000 s = 2,583 s ≈ 43 min of bad time |
| burned | 46.46% · 53.54% left |
If the last five minutes before sealed_through were BAD, the burn rate is 83.3× over 1 h and 13.9×
over 6 h.
| What changes | What cerbix reports |
|---|---|
Two days UNKNOWN instead of 30 minutes |
Coverage 93.3% → partial. Shows 99.9502% and 49.75% burned, labelled partial. Burn rules over these windows hold. |
| Definition edited mid-window | No aggregate and no budget. Each segment shows its own numbers. |
| Service created 10 days ago | 30d: insufficient_history. 7d is reported normally. |
9. Monitor uptime is a different number
Section titled “9. Monitor uptime is a different number”The Uptime on a monitor card is computed per monitor from raw results. It is useful at a glance, but it is not the Service SLI.
| Monitor uptime | Service SLI | |
|---|---|---|
| Unit | results: up ÷ total | time: good ÷ (good + bad) |
| Silence | not counted | UNKNOWN, lowers coverage |
| No data | 0% | withheld |
| Window ends | now | seal watermark |
| Scope | one monitor | declared members × regions, versioned |
| Also reports | average and p95 latency of successful checks | coverage, segments, burn |
Both paths exclude maintenance windows.