Skip to content
v0.3.8GitHub

Reading reliability views

Verified against cerbix f5240f5Report a problem ↗

cerbix shows a number only when it can defend it, and applies the same rule to pictures. This page explains what each state, colour and mark means on reliability views, so you can tell a measured failure from missing evidence.

Every state has its own encoding, and every encoding comes with a text label. Colour is never the only signal: legends name each state, status pills carry a dot and a word, and each timeline cell has a readout.

State Meaning Encoding
good measured and serving status green
down measured and failing (BAD) status red
unknown cerbix should have measured and could not solid neutral grey, never a status hue
excluded declared out of scope, for example maintenance maintenance blue
no stored bucket no reliability fact exists for that time diagonal hatch with a cell outline
provisional not yet sealed, counted in no number the state’s colour at reduced opacity

UNKNOWN is a verdict, not a softer failure. It means cerbix decided that it does not know, and that time counts against coverage. A missing bucket is different: there is no verdict at all, so it is never painted as unknown. Opacity means provisional and nothing else.

A Service’s health card answers “what is happening now”. It is categorical, unstable by definition and never a percentage.

  • SLI status: healthy, degraded, down or unknown, from the SLI members only.
  • Diagnostics: ok, failing or unknown, from every monitor on the Service.

While the request is pending, the card says it is checking. If the request fails, the card reports a transport problem, not unknown. A failed read is never presented as a product state.

The reliability strip covers the requested UTC range at a fixed grain: one cell per hour for 24h, one per day for longer windows.

  • A cell’s width is its real duration. Time with no stored bucket still takes its width, drawn in the hatch encoding. A month with three days of facts shows 27 days of hatch.
  • A cell is a proportional stack of the states inside it. A day at 99.5% good is mostly green with a thin red slice.
  • A down, unknown or excluded slice thinner than 2 px is raised to 2 px when the cell can spare the height. When it cannot, the cell is marked, and the readout names the state and its exact duration. A problem is never hidden.
  • Hover or focus a cell to see its local and UTC extent, the good/down/unknown/excluded split with exact durations, the stored bucket count, and any provisional or repairing flags.

Definition revision boundaries are accent-coloured vertical marks; evaluation epoch boundaries are thinner grey marks. Marks that would overlap are merged into one mark anchored at the earliest boundary, so its position stays true.

Reported numbers cover [sealed_through − window, sealed_through). Buckets after sealed_through are provisional: they appear on the strip at reduced opacity and in no number. A line marks sealed_through, and a note states how many buckets lie beyond it.

A healthy Service seals about three minutes behind real time: a 60 s bucket plus 120 s of late-arrival grace. The Services list shows sealed_through for each Service and flags it as behind once the lag reaches 5 min. Ranges being recomputed are listed as repairing and masked on the strip as work in progress, never as data.

A number that cannot be stated is replaced by a dash with its reason. cerbix never shows 100% or 0× for an empty denominator.

Status What you see
ok the number
partial, reason decidable_coverage_below_min the number, labelled partial, with coverage below 95%
partial, reason storage_gap a dash: stored buckets are missing inside the window
unavailable a dash: nothing measured, or no SLI members declared
insufficient_history a dash: the window starts before the Service’s history
insufficient_sealed_coverage a dash: nothing is sealed for this window yet

When the definition changed inside the window, the card shows a ∅ banner instead of one availability figure and lists each segment with its own numbers. A segment states its storage as contiguous or <stored> of <extent> · incomplete. An incomplete segment quotes no availability, but its coverage stays printed, because coverage answers a different question. A segment built by a first-adoption backfill is labelled declared reconstruction.

For the formulas behind these statuses, see How the SLI is computed.

The monitor’s response time panel plots the checks it fetched at their real timestamps.

  • Every check is drawn. A failure with no latency, such as a composite evaluation, sits on the baseline instead of disappearing.
  • Points are not joined by a line or a fill unless the expected-run ledger proves the span was covered. The ledger records only when ledger.carrier_enabled is on; it is off by default.
  • An observation ruler under the plot has one tick per recorded check. An empty span means only that no check was recorded between two points. It is not a verdict and uses no threshold.
  • The header always states the monitor’s timeout. It is drawn as a line only when it falls inside the plotted range.

Bucket identity and all arithmetic are UTC. Times are shown in your browser’s zone with the offset named, and readouts add the UTC instant on its own line. A UTC day is never labelled as your calendar day: a viewer at UTC+05 sees 01.09 05:00 → 02.09 05:00 (UTC+05).

A public component with no measurement shows No data in a neutral colour. It is never ranked against the outage states. A 90-day figure that cannot be quoted is replaced by its reason in words. See Status pages.

The Dashboard distinguishes loading, loaded, no-data and failed reads. Placeholders such as 0 / 0 are not shown while data loads, missing availability buckets stay neutral and are left out of uptime arithmetic, and a failed read shows an error instead of an empty dashboard.