Skip to content
v0.3.8GitHub

Architecture and deployment topologies

Verified against cerbix f5240f5Report a problem ↗

cerbix ships as one static Go binary. The same binary runs every part of the system; the --role flag decides which part a process runs. This page explains what each role does and needs, how a check travels from the scheduler to an alert, and which deployment topology fits your environment.

cerbix architecture — distributed and geo topologyClients reach the api role through your ingress. The api, a leader-elected scheduler, PostgreSQL and RabbitMQ form the core. Worker pools per region take jobs from RabbitMQ and probe your services; an agent in a segment without broker access pulls its jobs over HTTPS. The full list of connections and the path of one check follow the figure.CORE · ONE NETWORKREGION POOLS · AMQPREMOTE SEGMENT · NO BROKERBrowseroperators, viewersCI pipelinecerbix CLI, API tokenAlertmanageroptional, inbound alertsYour ingress — TLS termination (Traefik, nginx, Caddy). Restrict /metrics here.api × NREST · SSE · web UI · result ingest · outbox delivery · Monitoring as CodePostgreSQL 15+facts, settings, outbox, pull jobsRabbitMQ — distributed roles onlychecks.jobs.<region> · checks.results · checks.deadscheduler × 2one leader, one standbyOutbound from coreOIDC providersign-in (any issuer)Notification channelsemail · Slack · Telegram ·webhook, delivered from theoutboxworker × M--region coreworker × M--region eu-centralYour servicesreachable from coreYour servicesin eu-centralagent--region dc-privatePrivate servicesonly reachable insideUI · REST · SSEgate check · change recordwebhookHTTP :8080SQLconsumes resultsleader lockjobsAMQP · jobs in, results outprobeprobeprobeHTTPS pulloutbound onlyevery role serves/healthz · /readyz · /metrics1234567cerbix architecture — distributed and geo topologyOne binary, five roles. Arrows point from the side that opens the connection; moving dots show where data flows.cerbix process (--role)your infrastructure or externalSingle node: one cerbix serve --role all process runs api, scheduler and probers in-process — no RabbitMQ needed. Agents connect the same way in bothtopologies. Region names are examples.Path of one check ① scheduler publishes the job → ② the region’s worker takes it → ③ probe → ④ result back to the broker → ⑤ api ingests it⑥ heartbeat, status and outbox event in one transaction → ⑦ notification delivered

Connections, from the side that opens each one:

  • Browser opens HTTPS to your ingress: UI, REST, SSE.
  • CI pipeline opens HTTPS to your ingress: cerbix gate check and change record with an API token.
  • Alertmanager (optional) opens HTTPS to your ingress to deliver inbound alerts by webhook.
  • Your ingress proxies plain HTTP to the api role on port 8080.
  • api connects to PostgreSQL over SQL and writes heartbeats, status and outbox events.
  • api connects to RabbitMQ and consumes results from checks.results.
  • scheduler connects to PostgreSQL, which also holds the leader lock between the two schedulers.
  • scheduler connects to RabbitMQ and publishes jobs to checks.jobs.<region>.
  • The core connects out to the OIDC provider for sign-in and to notification channels — email, Slack, Telegram, webhook — delivered from the outbox.
  • Workers in region core connect to RabbitMQ over AMQP: jobs come in, results go out.
  • Workers in region eu-central connect to RabbitMQ over AMQP the same way.
  • Workers in region core probe your services reachable from the core.
  • Workers in region eu-central probe your services in eu-central.
  • The agent in dc-private probes private services that are only reachable inside its segment.
  • The agent opens HTTPS to your ingress and pulls its jobs; the remote segment needs no broker and accepts no inbound connections.

Path of one check:

  1. scheduler publishes the job
  2. the region’s worker takes it
  3. probe
  4. result back to the broker
  5. api ingests it
  6. heartbeat, status and outbox event in one transaction
  7. notification delivered

Start any role with cerbix serve:

Terminal window
cerbix serve --config /etc/cerbix/config.yaml --role all
cerbix serve --config /etc/cerbix/worker.yaml --role worker --region eu-central

--role defaults to all. --region applies to worker and agent; an empty value means the core region.

Role What it runs Needs PostgreSQL Needs RabbitMQ
all Everything in one process: scheduler, prober pool, result ingest, REST API, SSE, web UI, outbox delivery and Monitoring as Code providers. Uses the in-process dispatcher. yes no
api REST API, SSE stream, web UI, /auth/*, the check-result consumer, outbox delivery and Monitoring as Code providers. yes yes
scheduler The leader-elected scheduler: publishes due jobs, runs rollups, retention, re-notification, burn-rate evaluation, SLA reports, region-worker alerts, escalations and service-reliability maintenance. Also delivers the outbox. yes yes
worker A stateless prober pool that consumes the jobs of one region. no, except the core worker if you run composite monitors yes
agent An HTTP-pull prober for a network segment with no broker access. Talks to the api role over outbound HTTPS only. no no

Every role starts an operational HTTP server with /healthz, /readyz and /metrics. Only all and api serve the web UI and the API, so only those belong behind your ingress.

A distributed role (api, scheduler, worker) refuses to start without rabbitmq.url and logs rabbitmq_required. An agent refuses to start without pull.server_url and pull.token.

  1. The scheduler leader finds a due monitor and publishes a job through the dispatcher.
  2. An executor (an in-process worker, a worker process or an agent) runs the probe behind the SSRF guard.
  3. The result returns to the ingest consumer, which writes the heartbeat and updates the monitor status.
  4. A status change and the notification event commit in the same transaction. The outbox delivers the notification afterwards.

The all role uses an in-process dispatcher: buffered channels inside one process, no broker. The distributed roles use RabbitMQ. cerbix publishes to named queues through the default exchange; it declares no topic exchange of its own.

Queue family Purpose
checks.jobs.<region> and versioned checks.jobs.v2.<region> … checks.jobs.v4.<region> Scheduled jobs, one queue per region and carrier generation.
checks.tests.<region> (and v2, v3) “Test connection” requests, answered through a temporary reply queue.
checks.results Results from every region; no TTL.
checks.dead Bodies an executor or the consumer refused, kept for inspection.

A job carries a per-message TTL of about one monitor interval. If no executor consumes it, the job expires and the next scheduler tick issues a new one, so a dead region does not build an unbounded backlog. If the broker connection drops, the dispatcher reconnects with exponential backoff (1 s to 30 s) and re-subscribes its consumers without a process restart. Watch cerbix_broker_up.

Every monitor belongs to exactly one region. The default is core; a region name matches ^[a-z0-9-]{1,40}$. Affinity is strict: a job for region eu-central is never run by another region’s executor, and there is no fallback to core. A private target is often reachable from one segment only, and probing it from elsewhere would report a false DOWN.

In the all role every region except the pull-served ones runs in the local process. Regions listed in pull.regions still go to their agents. See Geo-distributed probers.

The heartbeat, the status flip, the auto-incident and the notification event are written in one PostgreSQL transaction. The outbox worker then delivers with retry, backoff and a dead-letter state. Claims use FOR UPDATE SKIP LOCKED, so the outbox runs safely on every all, api and scheduler replica. worker and agent never deliver notifications.

Suppression (instance-wide silence, dependency suppression, escalation rules) applies at delivery time. Facts keep recording while delivery is muted. A global admin can list and replay dead-lettered events at GET /api/v1/admin/outbox/dead.

Scheduler replicas compete for a PostgreSQL session advisory lock. One replica holds it and schedules; the others wait. The leader re-checks the lock every 5 s on its held connection and steps down at once if that connection is lost. A standby then takes over. The leader keeps an in-memory snapshot of enabled monitors and reloads it every 15 s; a committed Monitoring as Code apply wakes it earlier. The cerbix_scheduler_leader gauge shows which replica leads.

Run two scheduler replicas for availability. Adding more does not add throughput: one leader does the work.

PostgreSQL 15 or newer is required. The repository images use PostgreSQL 16. cerbix migrate reads the server version before it applies anything and refuses an older server with an explanation. Migrations are embedded in the binary and every role with a database.dsn applies them at startup; concurrent starts are serialized by an advisory lock.

Raw heartbeats use one of two layouts. cerbix detects the layout when it opens the database.

Condition heartbeats layout Retention
The timescaledb extension is installed and heartbeats is a hypertable Hypertable with one-day chunks, native compression after 7 days Leader drops expired chunks
Plain PostgreSQL Declarative daily RANGE partitions plus a DEFAULT partition, so an insert never fails for a missing partition Leader creates dated partitions, drops expired ones and purges expired rows from DEFAULT

The conversion to a hypertable happens in a migration. Install the extension before cerbix first migrates the database; adding it later leaves heartbeats in partition mode. heartbeats.retention_days sets the raw window (default 30, minimum 2). The daily rollup has the same schema in both modes, and service-reliability facts are partitioned by month in both modes.

One cerbix serve --role all process and PostgreSQL. No broker is needed. This is the path described in Install, with Docker Compose or a systemd unit. Put a TLS-terminating reverse proxy in front of port 8080.

Run the roles as separate deployments against one PostgreSQL and one RabbitMQ:

  • api: N stateless replicas behind a load balancer. Sessions live in PostgreSQL, so sticky sessions are not required. Every replica that runs a file provider must see identical provider directories.
  • scheduler: two replicas, one active.
  • worker: M replicas per region; add replicas as probe load grows. Queue prefetch spreads jobs across them.

Route all ingress traffic to the api replicas. /metrics shares the listener with the API, so restrict that path at the proxy or scrape it over an internal network.

Keep the core (API, scheduler, PostgreSQL, RabbitMQ) in one place and run probers inside the segments you need to observe: worker processes on the central RabbitMQ, or agent processes where exposing the broker is not an option. See Geo-distributed probers.