Skip to content
v0.3.8GitHub

How the SLI is computed

Source: src/content/docs/docs/concepts/sli.md · verified against cerbix f5240f5Edit on GitHub ↗

A Service’s SLI is the share of decidable time during which it was available. cerbix measures it from its own checks, in one-minute buckets that are sealed before they are reported. This page follows one heartbeat to the number on the service page, and lists every case in which cerbix withholds that number instead.

availability = good / (good + bad) × 100 # absent when good + bad = 0
coverage = (good + bad) / (good + bad + unknown)
budget burned = (bad / (good + bad)) / (1 − objective / 100) × 100

Every term is a duration summed over sealed one-minute buckets in the window. excluded time — maintenance and disabled members — appears in no denominator. Below 95% coverage the result is partial; with nothing decidable it is absent, never 100%.

The path is always the same:

  1. Heartbeat — the raw up / down result of a check.
  2. Member state — GOOD, BAD, UNKNOWN or EXCLUDED at every instant.
  3. Region, then Service — members are combined with the Service’s aggregation policy.
  4. Sealed bucket — one minute, duration-weighted, sealed two minutes after it closes.
  5. Window report — a window that ends at the seal watermark, never at now.

For every SLI member and every instant t, cerbix applies these rules in order. The first match wins.

# Condition Active probe Push
1 Member disabled in this evaluation epoch EXCLUDED EXCLUDED
2 Inside a maintenance window, with maintenance: exclude (default) EXCLUDED EXCLUDED
3 No observation yet UNKNOWN (no_observation) BAD once armed + stale_after has passed, UNKNOWN before
4 Last observation is at least stale_after old UNKNOWN (stale) BAD
5 Last observation is up GOOD GOOD
6 Last observation is down BAD BAD

Silence means different things by check type. For an active probe a missing result is uncertainty. For a push monitor it is the failure the dead-man’s switch exists to catch.

An observation stays in force until the next one arrives or until stale_after runs out, whichever comes first. The bound is fixed per evaluation epoch:

stale_after(active) = max(3 × interval, 90 s)
stale_after(push) = interval + grace
Check interval Result holds for
10 s 90 s
30 s 90 s
60 s 180 s
5 min 15 min

The Service’s missing_data policy then decides what an UNKNOWN member means:

  • unknown (default) keeps it UNKNOWN;
  • bad counts it as BAD;
  • ignore removes it from the count. Ignoring never turns a member into EXCLUDED: if nothing is left to count, the result is UNKNOWN.

Only members declared in sli[] count. Context members (monitors[]) are shown for diagnosis and never change the number. The order is fixed: normalize → exclude → apply missing_data → combine within each region → combine across regions.

Within a region, with n eligible members of which g are GOOD, b are BAD and u are UNKNOWN:

Mode GOOD when UNKNOWN when BAD when
all (default) b = 0 and u = 0 b = 0 and u > 0 b > 0
any g > 0 g = 0 and u > 0 b = n
quorum with threshold k g ≥ k g < k ≤ g + u g + u < k

A quorum is BAD only when it would fail even if every UNKNOWN member turned out GOOD. For a quorum of 2 out of 3 members:

Members Result
GOOD GOOD UNKNOWN GOOD
GOOD UNKNOWN BAD UNKNOWN
UNKNOWN UNKNOWN BAD UNKNOWN
GOOD BAD BAD BAD

If exclusions leave fewer than k eligible members, the threshold is clamped to n and the bucket is marked weakened, so the relaxed rule stays visible.

Regions. Each region is decided first, then the regions are combined with the same quorum logic. A region that reports nothing is UNKNOWN, never dropped. With the default per_region policy, one region is enough to be available and all regions are needed to be healthy, so a single dark region makes the Service DEGRADED. DEGRADED still counts as available: health is shown beside the SLI, and only availability feeds the SLO. With all_regions, every region must be decided GOOD; one dark region makes the Service UNKNOWN for that time.

3. Duration weighting inside a one-minute bucket

Section titled “3. Duration weighting inside a one-minute bucket”

Time is cut into one-minute buckets, half-open [start, end) and aligned to UTC. Inside a bucket the Service state can change only at a breakpoint: a bucket edge, an observation, an observation going stale, or a maintenance edge. Each piece between breakpoints goes whole into one bin. Durations are stored in integer microseconds, and every bucket satisfies:

good + bad + unknown + excluded = 60 s

For example, a bucket that starts GOOD, turns BAD when a check fails, becomes UNKNOWN when a required member’s last result goes stale, and enters a maintenance window before it closes might record 22 s good, 9 s bad, 10 s unknown and 19 s excluded (example data).

  • A bucket [start, end) is sealed once now ≥ end + 2 min. The two minutes are the late-arrival grace.
  • sealed_through is the latest instant up to which every bucket exists and is sealed. A missing bucket holds the watermark back; it is never skipped.
  • Every window is [sealed_through − W, sealed_through). It never ends at now, so the same window gives the same number every time it is read. Normal lag is two to three minutes.
  • A result that arrives for an already sealed bucket triggers an audited recompute of that bucket.
  • The live state on the service page is a separate point-in-time evaluation. It is never mixed into the SLI.

Reports are available for 24h, 7d, 30d (default) and 90d. Over the sealed buckets in a window:

availability = good / (good + bad) × 100 # absent when good + bad = 0
coverage = (good + bad) / (good + bad + unknown)
# excluded time is in neither denominator

UNKNOWN time does not count for or against availability. It lowers coverage instead, which decides whether the number may be quoted.

Each window gets one status. The first matching row wins.

Condition Status · reason Number shown
The window starts before the Service was first measured insufficient_history No
A bucket in the window is missing partial · storage_gap No
No decidable time at all unavailable · zero_decidable_time No
Coverage below 95% partial · decidable_coverage_below_min Yes, labelled partial
Otherwise ok Yes

Independently of the status, a window whose facts span more than one definition revision has its aggregate withheld (spans_definition_revisions). It is reported as segments, one per revision and evaluation epoch, each with its own numbers. A Service with no SLI members reports no_sli; one with nothing sealed yet reports nothing_sealed. The 95% threshold is fixed, not a setting.

An objective is set per Service and per window, between 0 and 100 exclusive, up to 99.9999. A budget is computed only when the availability can be quoted.

allowed = 1 − objective / 100
actual = bad / (good + bad)
burned % = actual / allowed × 100
budget left = 100 − burned %
burn rate = actual(w) / allowed # over a burn window w

The budget is a share of measured time, not of wall-clock time: maintenance and UNKNOWN minutes neither spend it nor earn it. Burn-rate alert rules and their FIRE / CLEAR / HOLD semantics are described in SLOs, error budgets and burn rate.

A 30d window (2,592,000 s, all buckets sealed, one definition revision) with an objective of 99.9 (example data):

Input Duration
good 2,581,800 s
bad 1,200 s (20 min)
unknown 1,800 s (30 min)
excluded 7,200 s (2 h maintenance)
Result Value
availability 2,581,800 / 2,583,000 = 99.9535%
coverage 2,583,000 / 2,584,800 = 99.93% → ok
budget 0.001 × 2,583,000 s = 2,583 s ≈ 43 min of bad time
burned 46.46% · 53.54% left

If the last five minutes before sealed_through were BAD, the burn rate is 83.3× over 1 h and 13.9× over 6 h.

What changes What cerbix reports
Two days UNKNOWN instead of 30 minutes Coverage 93.3% → partial. Shows 99.9502% and 49.75% burned, labelled partial. Burn rules over these windows hold.
Definition edited mid-window No aggregate and no budget. Each segment shows its own numbers.
Service created 10 days ago 30d: insufficient_history. 7d is reported normally.

The Uptime on a monitor card is computed per monitor from raw results. It is useful at a glance, but it is not the Service SLI.

Monitor uptime Service SLI
Unit results: up ÷ total time: good ÷ (good + bad)
Silence not counted UNKNOWN, lowers coverage
No data 0% withheld
Window ends now seal watermark
Scope one monitor declared members × regions, versioned
Also reports average and p95 latency of successful checks coverage, segments, burn

Both paths exclude maintenance windows.