Skip to content
v0.3.8GitHub

Introduction

Source: src/content/docs/docs/getting-started/introduction.md · verified against cerbix f5240f5Edit on GitHub ↗

cerbix is a self-hosted, multi-tenant service reliability platform. A Service declares what its reliability is — which checks are its SLI, how regions aggregate, what counts as pageable — in versioned definitions, so every number is traceable to the rule that produced it.

cerbix measures that from its own checks, reports SLO, error budget and burn rate, and withholds any number it cannot defend. The same Service then drives the response: suppression, incidents, dependency impact, status pages, on-call escalation and a release gate.

Capability What it gives you
Service as the reliability object Explicit SLI membership, separately declared context members, aggregation policies (all, any, quorum, region-aware), immutable definition revisions and evaluation epochs.
Reliability that refuses to guess GOOD, BAD and UNKNOWN with reasons, two independent coverage axes, sealed facts, and a withheld value with a reason instead of a confident 100%. See How the SLI is computed.
SLO, error budget and burn rate Per window, quoted against the objective in force, with burn-rate alerting evaluated on sealed facts.
Release gate A pipeline asks cerbix gate check whether the error budget allows a release and gets ALLOW, WARN, BLOCK or UNKNOWN with every reason. cerbix decides; the pipeline acts.
Change intelligence The pipeline records deploys, rollbacks and flag flips; the Service shows which changes preceded an incident — never “caused” — and the SLI before and after.
Paging that does not double A Service can own paging for its members. Suppression is per signal and only while coverage is armed. Anything ambiguous fails open.
Incidents and status pages Automatic, manual and API incidents anchored to a monitor or a Service; timelines, postmortems and public status pages.
Dependency impact A same-project service graph correlates an incident with its upstreams and downstreams. It records candidates; it never elects a culprit.
On-call and escalation Escalation ladders, rotations with vacation overrides, acknowledge-to-stop.
Reliable delivery A transactional outbox with retry, backoff and dead-letter; confirmations, re-notify and instance-wide silence.
Authentication and tenancy OIDC with any provider, local argon2id login, TOTP two-factor authentication; organization → project multi-tenancy with role-based access control.
Security AES-256-GCM secrets at rest with key rotation; a distroless, non-root image.

cerbix produces its own observations. Every check runs from the region that owns it.

Group Types
Network http, tcp, icmp, dns, tls, grpc, websocket, ssh
Data and queues postgres, mysql, redis, rabbitmq
Metrics promql
Flows synthetic — scripted multi-step HTTP flows; async_canary — a typed submit → poll → verify transaction
Composition composite — a group of monitors; push — a dead-man’s switch

Details and settings for each type: Check types.

The whole application ships as one static Go binary. The web UI, the REST API and the database migrations are embedded, so there is no separate web server, frontend bundle or migration script to deploy. The binary runs in one of five roles — all, api, scheduler, worker, agent — so the same artifact scales from a single process on a laptop to region pools and pull agents. PostgreSQL 15 or later is required; RabbitMQ is needed only for distributed roles. See Architecture and roles.

Alongside The relationship
Prometheus, Grafana, Loki, Tempo Not a replacement. cerbix stores none of your telemetry and has no query language. It reads PromQL as a check source and exports its own cerbix_* metrics. Keep those tools for why; cerbix answers is this service reliable by our definition, and what are we doing about it.
Checkly, Pingdom, Uptime Kuma Synthetic checks overlap. There a monitor is the final object; here it is an observation, and the object is a Service with a versioned definition, a budget and a response — plus probing from inside the perimeter, multi-tenancy and self-hosting.
PagerDuty, Opsgenie cerbix carries rotations and escalation for its own incidents. It is not a hub for signals from other systems.
Nobl9, Pyrra, Sloth Those compute SLOs over your metrics. cerbix produces the observations itself, keeps one derived timeline per Service, and versions the definition rather than only the threshold.

Stated plainly, because the boundary is a feature. cerbix does not serve arbitrary time-series queries, store generic telemetry, support user-defined downsampling, expose a query language, act as a metrics backend or provide a service catalog. There is no trace or log ingestion and no automatic root-cause analysis: it offers correlation candidates over a dependency graph that you declared.