Architecture and deployment topologies
f5240f5Report a problem ↗cerbix ships as one static Go binary. The same binary runs every part of the system; the --role flag decides which part a process runs. This page explains what each role does and needs, how a check travels from the scheduler to an alert, and which deployment topology fits your environment.
Connections, from the side that opens each one:
- Browser opens HTTPS to your ingress: UI, REST, SSE.
- CI pipeline opens HTTPS to your ingress: cerbix gate check and change record with an API token.
- Alertmanager (optional) opens HTTPS to your ingress to deliver inbound alerts by webhook.
- Your ingress proxies plain HTTP to the api role on port 8080.
- api connects to PostgreSQL over SQL and writes heartbeats, status and outbox events.
- api connects to RabbitMQ and consumes results from checks.results.
- scheduler connects to PostgreSQL, which also holds the leader lock between the two schedulers.
- scheduler connects to RabbitMQ and publishes jobs to checks.jobs.<region>.
- The core connects out to the OIDC provider for sign-in and to notification channels — email, Slack, Telegram, webhook — delivered from the outbox.
- Workers in region core connect to RabbitMQ over AMQP: jobs come in, results go out.
- Workers in region eu-central connect to RabbitMQ over AMQP the same way.
- Workers in region core probe your services reachable from the core.
- Workers in region eu-central probe your services in eu-central.
- The agent in dc-private probes private services that are only reachable inside its segment.
- The agent opens HTTPS to your ingress and pulls its jobs; the remote segment needs no broker and accepts no inbound connections.
Path of one check:
- scheduler publishes the job
- the region’s worker takes it
- probe
- result back to the broker
- api ingests it
- heartbeat, status and outbox event in one transaction
- notification delivered
One binary, five roles
Section titled “One binary, five roles”Start any role with cerbix serve:
cerbix serve --config /etc/cerbix/config.yaml --role allcerbix serve --config /etc/cerbix/worker.yaml --role worker --region eu-central--role defaults to all. --region applies to worker and agent; an empty value means the core region.
| Role | What it runs | Needs PostgreSQL | Needs RabbitMQ |
|---|---|---|---|
all |
Everything in one process: scheduler, prober pool, result ingest, REST API, SSE, web UI, outbox delivery and Monitoring as Code providers. Uses the in-process dispatcher. | yes | no |
api |
REST API, SSE stream, web UI, /auth/*, the check-result consumer, outbox delivery and Monitoring as Code providers. |
yes | yes |
scheduler |
The leader-elected scheduler: publishes due jobs, runs rollups, retention, re-notification, burn-rate evaluation, SLA reports, region-worker alerts, escalations and service-reliability maintenance. Also delivers the outbox. | yes | yes |
worker |
A stateless prober pool that consumes the jobs of one region. | no, except the core worker if you run composite monitors |
yes |
agent |
An HTTP-pull prober for a network segment with no broker access. Talks to the api role over outbound HTTPS only. |
no | no |
Every role starts an operational HTTP server with /healthz, /readyz and /metrics. Only all and api serve the web UI and the API, so only those belong behind your ingress.
A distributed role (api, scheduler, worker) refuses to start without rabbitmq.url and logs rabbitmq_required. An agent refuses to start without pull.server_url and pull.token.
How a check travels
Section titled “How a check travels”- The scheduler leader finds a due monitor and publishes a job through the dispatcher.
- An executor (an in-process worker, a
workerprocess or anagent) runs the probe behind the SSRF guard. - The result returns to the ingest consumer, which writes the heartbeat and updates the monitor status.
- A status change and the notification event commit in the same transaction. The outbox delivers the notification afterwards.
Dispatcher: in-process or AMQP
Section titled “Dispatcher: in-process or AMQP”The all role uses an in-process dispatcher: buffered channels inside one process, no broker. The distributed roles use RabbitMQ. cerbix publishes to named queues through the default exchange; it declares no topic exchange of its own.
| Queue family | Purpose |
|---|---|
checks.jobs.<region> and versioned checks.jobs.v2.<region> … checks.jobs.v4.<region> |
Scheduled jobs, one queue per region and carrier generation. |
checks.tests.<region> (and v2, v3) |
“Test connection” requests, answered through a temporary reply queue. |
checks.results |
Results from every region; no TTL. |
checks.dead |
Bodies an executor or the consumer refused, kept for inspection. |
A job carries a per-message TTL of about one monitor interval. If no executor consumes it, the job expires and the next scheduler tick issues a new one, so a dead region does not build an unbounded backlog. If the broker connection drops, the dispatcher reconnects with exponential backoff (1 s to 30 s) and re-subscribes its consumers without a process restart. Watch cerbix_broker_up.
Region affinity
Section titled “Region affinity”Every monitor belongs to exactly one region. The default is core; a region name matches ^[a-z0-9-]{1,40}$. Affinity is strict: a job for region eu-central is never run by another region’s executor, and there is no fallback to core. A private target is often reachable from one segment only, and probing it from elsewhere would report a false DOWN.
In the all role every region except the pull-served ones runs in the local process. Regions listed in pull.regions still go to their agents. See Geo-distributed probers.
Transactional outbox
Section titled “Transactional outbox”The heartbeat, the status flip, the auto-incident and the notification event are written in one PostgreSQL transaction. The outbox worker then delivers with retry, backoff and a dead-letter state. Claims use FOR UPDATE SKIP LOCKED, so the outbox runs safely on every all, api and scheduler replica. worker and agent never deliver notifications.
Suppression (instance-wide silence, dependency suppression, escalation rules) applies at delivery time. Facts keep recording while delivery is muted. A global admin can list and replay dead-lettered events at GET /api/v1/admin/outbox/dead.
Scheduler leader election
Section titled “Scheduler leader election”Scheduler replicas compete for a PostgreSQL session advisory lock. One replica holds it and schedules; the others wait. The leader re-checks the lock every 5 s on its held connection and steps down at once if that connection is lost. A standby then takes over. The leader keeps an in-memory snapshot of enabled monitors and reloads it every 15 s; a committed Monitoring as Code apply wakes it earlier. The cerbix_scheduler_leader gauge shows which replica leads.
Run two scheduler replicas for availability. Adding more does not add throughput: one leader does the work.
Storage
Section titled “Storage”PostgreSQL 15 or newer is required. The repository images use PostgreSQL 16. cerbix migrate reads the server version before it applies anything and refuses an older server with an explanation. Migrations are embedded in the binary and every role with a database.dsn applies them at startup; concurrent starts are serialized by an advisory lock.
Adaptive heartbeat storage
Section titled “Adaptive heartbeat storage”Raw heartbeats use one of two layouts. cerbix detects the layout when it opens the database.
| Condition | heartbeats layout |
Retention |
|---|---|---|
The timescaledb extension is installed and heartbeats is a hypertable |
Hypertable with one-day chunks, native compression after 7 days | Leader drops expired chunks |
| Plain PostgreSQL | Declarative daily RANGE partitions plus a DEFAULT partition, so an insert never fails for a missing partition |
Leader creates dated partitions, drops expired ones and purges expired rows from DEFAULT |
The conversion to a hypertable happens in a migration. Install the extension before cerbix first migrates the database; adding it later leaves heartbeats in partition mode. heartbeats.retention_days sets the raw window (default 30, minimum 2). The daily rollup has the same schema in both modes, and service-reliability facts are partitioned by month in both modes.
Deployment topologies
Section titled “Deployment topologies”Single node
Section titled “Single node”One cerbix serve --role all process and PostgreSQL. No broker is needed. This is the path described in Install, with Docker Compose or a systemd unit. Put a TLS-terminating reverse proxy in front of port 8080.
Distributed
Section titled “Distributed”Run the roles as separate deployments against one PostgreSQL and one RabbitMQ:
api: N stateless replicas behind a load balancer. Sessions live in PostgreSQL, so sticky sessions are not required. Every replica that runs a file provider must see identical provider directories.scheduler: two replicas, one active.worker: M replicas per region; add replicas as probe load grows. Queue prefetch spreads jobs across them.
Route all ingress traffic to the api replicas. /metrics shares the listener with the API, so restrict that path at the proxy or scrape it over an internal network.
Geo-distributed
Section titled “Geo-distributed”Keep the core (API, scheduler, PostgreSQL, RabbitMQ) in one place and run probers inside the segments you need to observe: worker processes on the central RabbitMQ, or agent processes where exposing the broker is not an option. See Geo-distributed probers.