Skip to content
v0.3.8GitHub

Geo-distributed probers

Verified against cerbix f5240f5Report a problem ↗

A private target is often reachable from one network segment only. cerbix probes it from inside that segment with a prober that runs there, while the core (API, scheduler, PostgreSQL) stays in one place. This page explains the two prober transports, how to deploy each, and how cerbix alerts when a region loses its prober.

AMQP worker pool HTTP-pull agent
Command cerbix serve --role worker --region <name> cerbix serve --role agent --region <name>
Connects to The central RabbitMQ The central api role over HTTPS
Network direction Worker to broker (AMQP or AMQPS) Outbound HTTPS only
Needs PostgreSQL No No
Liveness signal A consumer on the region’s job queue, read from the RabbitMQ management API An agent heartbeat within the last 45 s
Test connection route RPC on checks.tests.<region> A database queue the agent polls

Use a worker pool when the segment can reach your broker through a tunnel or a TLS listener. Use an agent when exposing the broker to the segment is not an option.

Every monitor has a region field. The default is core, and a region name matches ^[a-z0-9-]{1,40}$. A monitor runs only in its own region: there is no fallback to core, because probing a private target from the wrong segment would report a false DOWN. Composite monitors always run in core.

The region picker reads GET /api/v1/regions. It lists regions that monitors use plus regions with a live executor, each marked live or not.

A remote worker needs only the central broker URL and a probe policy:

# worker.yaml (example)
log: { level: info, format: json }
rabbitmq:
url: "amqps://cerbix:${RABBITMQ_PASSWORD}@rabbitmq.example.com:5671/"
prober:
allow_private_ips: true # the default; lets the worker reach targets inside the segment
Terminal window
cerbix serve --config worker.yaml --role worker --region eu-central

The worker consumes only eu-central jobs. Scale it by adding replicas. The network path to RabbitMQ is yours to provide: a WireGuard tunnel, an amqps:// listener or an AMQP proxy. Do not expose the broker to the internet without TLS.

Liveness comes from the RabbitMQ management API. cerbix derives its address from rabbitmq.url (amqp to port 15672, amqps to port 15671) unless you set rabbitmq.management_url on the api and scheduler roles.

  1. Declare the region pull-served on the central api and scheduler roles. The scheduler then writes that region’s jobs to a database queue instead of RabbitMQ:

    pull:
    regions: ["dc-east"]
  2. Issue an agent token. Pick one kind:

    Token Where it lives Scope
    pull.token Central config Catch-all: any region
    pull.agents: [{region, token}] Central config One region
    Database token Settings → Administration → Agent tokens, or POST /api/v1/agent-tokens with {"name", "region"} (global admin) One region; revoke with DELETE /api/v1/agent-tokens/{tokenID} without a redeploy

    The secret of a database token appears once, in the create response. The agent endpoints are mounted only when the central config has pull.token, pull.agents or pull.regions.

  3. Configure the agent:

    # agent.yaml (example)
    server:
    listen: ":8080" # ops endpoints only
    pull:
    server_url: "https://cerbix.example.com"
    token: "${CERBIX_AGENT_TOKEN}"
    prober:
    allow_private_ips: true
  4. Start it in the segment:

    Terminal window
    cerbix serve --config agent.yaml --role agent --region dc-east
  • Claims jobs with a long poll. The API holds an empty claim for up to 20 s and answers as soon as a job for the region arrives.
  • Leases claimed jobs for 30 s. If an agent dies mid-batch, the jobs become claimable again.
  • Posts results to the same ingest path as AMQP results. A result for a monitor outside the token’s region is refused with 403.
  • Sends a heartbeat every 15 s. The heartbeat is the region’s liveness signal.
  • Buffers up to 10,000 results in memory while the API is unreachable. After reconnecting it replays them as history: they fill the SLA gap, but they do not open incidents or send alerts.
  • Negotiates the job carrier generation with the API and falls back to an older route when the API predates it, so agents and core can upgrade in either order.

Watch cerbix_pull_jobs_pending{region} and cerbix_pull_agent_lag_seconds{region}. A lag that keeps growing means an agent heartbeats but does not drain its queue.

The scheduler leader checks every 30 s that each region with enabled, non-push monitors has a live executor. Live means a RabbitMQ consumer on the region’s job queue, or an agent heartbeat within 45 s. A region that stays without an executor for 90 s triggers one message to every enabled notification channel of each affected project. When an executor returns, a recovery message follows. The grace period absorbs restarts, and a failed management lookup skips the check instead of reporting every region missing. Instance-wide silence suppresses these messages.

“Test connection” (POST /api/v1/projects/{projectID}/monitors/test) runs one probe in the monitor’s region, through the same transport as its scheduled jobs. A pull region’s agent picks up the test within about a second; the API waits for the result for the monitor timeout plus 6 s. Push and composite monitors cannot be tested.

If nothing in the region answers, the request fails with 502, not with a DOWN result:

Message Meaning
no worker queue for region "R" (AMQP 312 NO_ROUTE), immediately Nobody consumes the test queue the API published to. Often the API and the region’s workers disagree on secrets.dispatch_envelope.
no worker responded in region "R", after the timeout A worker took the request and returned nothing.
no agent responded in region "R" No agent in the pull region answered before the deadline.

When monitors carry credentials from the project secret inventory, core seals each job’s credentials for its region. Each executor holds only its own region’s key; the at-rest master key never leaves core.

# Core roles (all, api, scheduler): the at-rest key plus one keyring per region
security:
encryption_key: "${CERBIX_ENCRYPTION_KEY}"
dispatch:
regions:
dc-east: { primary: { id: "dc-east-2026a", key: "${DC_EAST_DISPATCH_KEY}" } }
secrets: { enabled: true, dispatch_envelope: "enforced" }
# The dc-east executor: its own keyring only
security:
dispatch:
regions:
dc-east: { primary: { id: "dc-east-2026a", key: "${DC_EAST_DISPATCH_KEY}" } }
secrets: { enabled: true, dispatch_envelope: "enforced" }

A key is a base64-encoded 32-byte value; an id matches ^[a-z][a-z0-9-]{0,62}$.

To rotate a regional key, follow Rotate a dispatch key in the runbook.