Reliability
you can prove.
Define what reliable means for a service. cerbix measures it with checks from inside your own perimeter — and withholds any number the evidence cannot defend.
Most tools tell you a service is down. cerbix makes you say what reliable means — then holds you to it.
Every observation becomes GOOD, BAD or UNKNOWN with a reason. The window ends at the seal watermark, never at now. Below 95% decidable coverage the number is marked partial; with nothing to decide it is absent — never rounded to 100%.
1398 of 1,440 minutes produced a result, all up. The silence is simply not counted.
42 min UNKNOWN — excluded from availability, counted against coverage.
Illustrative: one check, no failures — it simply stopped reporting.
Do not hide uncertainty
GOOD, BAD, UNKNOWN, no data, provisional and withheld are distinct states. Two independent coverage axes govern every aggregate.
GOOD · BAD · UNKNOWN(reason)Measure inside your perimeter
Self-hosted. Probe private segments and remote regions from the inside — AMQP worker pools or an HTTPS-pull agent with no broker in the geo.
--role worker · --role agentTurn evidence into action
The same Service drives SLO and error budget, incidents, status pages, on-call escalation and the release gate. One operating loop.
define → measure → respondEvery number is traceable to the rule that produced it.
Definitions are immutable revisions. You can always answer: under which definition was this month measured?
- 01 · Definerev 7
Declare the Service
Which checks are its SLI, how they aggregate, what is pageable. Click a member:
- all
- BAD
- any
- GOOD
- quorum 2 of 3
- UNKNOWN
- 02 · Observe
Probe from the inside
17 check types run from your perimeter, private segments and geo regions. cerbix produces its own observations.
eu-centralus-eastdc-private - 03 · Seal
Reduce to one timeline
GOOD / BAD / UNKNOWN, duration-weighted into one-minute buckets and sealed two minutes after each one closes.
See the math → - 04 · Act
Drive the response
SLO, error budget and burn rate; suppression, incidents, status pages, escalation and a release gate — from the same facts.
incidentpagegate
One binary. The whole operating loop.
- SLO, budget & burn rate24h, 7d, 30d, 90d — on sealed facts only.
- Release gateALLOW / WARN / BLOCK / UNKNOWN for CI, with reasons.
- Incidents & status pagesTimelines, postmortems, public service-first pages.
- On-call that does not doubleRotations, ladders, suppression that fails open.
- Geo probers, five rolesFrom one process to region pools and pull agents.
- OIDC, RBAC, secrets at restMulti-tenant, 2FA, AES-256-GCM with rotation.
See how evidence becomes a reliability decision.
Deploy your own in minutes: one binary and PostgreSQL.