Skip to content
v0.3.8GitHub

Services and definitions

Verified against cerbix f5240f5Report a problem ↗

A Service is where you state what reliability means for one operational unit. This page covers its inputs and policies, and how cerbix versions the declaration so every number matches the definition it was measured under.

A monitor runs one check in one region and records up or not up. A Service runs no checks. It names which monitors matter, how their results combine across regions, and who owns the response: a reference to an escalation policy and/or an on-call schedule, not a free-text team name. cerbix turns that declaration into per-minute reliability facts.

A monitor can feed several Services, or none. Each Service has an immutable, project-unique slug. A Service is not a security boundary: access is decided at the project level.

A declaration carries two lists of monitors, and you declare them separately:

Field Meaning
monitors Operational context: what is shown on the Service and available for diagnosis.
sli Reliability inputs: what counts toward availability, error budget and burn rate. Must be a subset of monitors.

The split is deliberate. Adding a Redis check for diagnosis does not change what availability means, because the check enters monitors only. An SLI member outside monitors is rejected (sli_not_in_monitors), so every input to the number is visible beside it.

# Example data — a format-2 bundle entry
services:
checkout:
name: Checkout
monitors: [checkout-http, checkout-db, checkout-redis, checkout-synthetic]
sli: [checkout-http, checkout-synthetic]

A Service with an empty sli is valid. It shows operational context and reports availability as unavailable with reason no_sli. It never reports 100%.

On the live health card, a failing diagnostic never degrades the SLI status.

Limit Default Hard maximum
services.max_services_per_project 50 200
services.max_members_per_revision (size of monitors) 50 200
services.max_services_per_monitor (Services using one monitor as an SLI member) 10 25

Each SLI member has a state at every instant: GOOD, BAD, UNKNOWN or EXCLUDED. The policy combines them in two stages, first within each region, then across regions. The exact evaluation is on How the SLI is computed.

Within a region, aggregation.mode:

Mode GOOD when BAD when Otherwise
all (default) every eligible member is GOOD any member is BAD UNKNOWN
any at least one member is GOOD every eligible member is BAD UNKNOWN
quorum at least degraded_min members are GOOD GOOD plus UNKNOWN members cannot reach degraded_min UNKNOWN

For quorum, healthy_min splits GOOD into HEALTHY and DEGRADED. It affects health only, never availability. Defaults: degraded_min: 1, and healthy_min equals the largest number of SLI members declared in one region.

Across regions, region.mode:

Mode Effect
per_region (default) GOOD when at least degraded_min_regions regions are GOOD; HEALTHY when at least healthy_min_regions are HEALTHY. Defaults: degraded_min_regions: 1, healthy_min_regions = every expected region.
any_region Shorthand for degraded_min_regions: 1, healthy_min_regions: 1.
all_regions Shorthand for both thresholds equal to the number of expected regions.

The expected regions are the distinct regions of the declared SLI members. A region with no observations is UNKNOWN, not ignored. With the defaults, one dark region makes a multi-region Service DEGRADED, not DOWN.

Thresholds are validated at write time against the declared member counts. While members are excluded, they are clamped to the eligible count and the clamp is recorded, so maintenance is never a definition error.

Field Values Effect
missing_data unknown (default), bad, ignore What an UNKNOWN member contributes. ignore drops it only while other known members keep the interval decidable. If it removes the last source of information, the interval is UNKNOWN, never EXCLUDED.
maintenance exclude (default, only value) A member inside a maintenance window leaves both the numerator and the denominator. Exclusion wins over the member’s own state.
freshness active_multiplier (default 3), active_floor (default 90s) How long an active check’s last result stays valid: the larger of multiplier × interval and the floor. A push member uses its own interval plus grace.

A disabled member is also excluded. An active check whose result went stale is UNKNOWN; a stale push member is BAD, because for a dead-man’s switch silence is the failure.

cerbix keeps two immutable, versioned records per Service:

  • A definition revision is what a person declared: monitors, sli and policies. Each write creates a new revision with the next revision number. The write must send expected_revision; a stale value returns 409 revision_conflict instead of merging.
  • An evaluation epoch is what the system measured: a snapshot of each SLI member’s evaluation settings and resolved freshness deadline. Every stored fact references one epoch, and each epoch resolves to exactly one revision.

A new epoch is created by:

  • every declaration write, together with its revision;
  • a monitor change that alters how an SLI member is evaluated: type, target, method, interval, timeout, retries, grace, conditions, region, enabled, type-specific settings, the name of a referenced secret, or a rotation of that secret.

No epoch is created by a change to a monitor’s name, description, tags, notification or escalation settings, dependencies, failure_threshold or confirmation interval, or by status updates. Changes to monitors that are only context members create no epoch either.

A window that spans a revision boundary is reported as separate segments, never one number.

Declaration writes are prospective. effective_at is the next minute boundary on the database clock: a write at 12:00:30 governs from 12:01:00, and a write at exactly 12:01:00 governs from 12:01:00. Sealed history is never restated under a new definition. If two writes land before the same boundary, the later one wins and the earlier is kept for audit as superseded_before_effect.

The first declaration may set backfill_from to adopt existing history. Those facts are evaluated with today’s members and settings and are labelled declared reconstruction. They are not evidence of how the monitors were configured at the time.

Maintenance windows are declarations about time too. Creating or annulling one that overlaps already-sealed time requires a preview and runs an audited repair of the affected range.

A Service is owned by the UI/API or by one Monitoring as Code file provider, never both.

  • UI/API. POST /api/v1/projects/{projectID}/services creates the Service with no declaration. PUT …/services/{serviceID}/declaration writes each revision.
  • File. A format: 2 bundle declares a services map keyed by slug, referencing monitor slugs. The schema is strict. An unchanged canonical hash creates no revision. A bundle that claims a slug held by a UI-owned Service is rejected with service_slug_owned_by_ui.

A file-managed Service carries managed_by. UI writes to its declaration, alerting, burn alerting, dependencies and deletion return 409 managed_by_file. Removing it from a bundle marks it orphaned; it is not deleted and stays file-owned.

References may cross owners in either direction. Removing a monitor that a Service declares is a declaration change:

  • Deleting a UI-owned monitor used by a UI-owned Service writes a system revision that removes it.
  • If the Service is file-owned, the delete is refused until the bundle drops the reference.

Epochs are exempt from ownership. You can edit the interval of a monitor that a file-owned Service uses; cerbix records a new epoch and leaves the declaration unchanged.