Skip to content
v0.3.8GitHub

Monitoring as Code

Verified against cerbix f5240f5Report a problem ↗

A file provider watches a directory of YAML bundles and reconciles them into cerbix monitors at runtime. You change monitors by changing files; no restart or reload is needed. This page covers the bundle format, ownership, how reconciliation behaves when files are wrong or missing, and how to operate providers in a cluster.

  • One bundle file describes the monitors of one project.
  • Files are the desired state for the monitors a provider owns. PostgreSQL stays the store for runtime state and history.
  • Only the api and all roles run providers. The scheduler and workers never read YAML; they see ordinary monitors.
  • A provider never creates organizations or projects. The slugs a bundle names must already exist.

Providers are declared in the static config under providers.file.<name>. Changing a provider’s name, scope, directory or limits needs a restart; the files inside the directory do not.

providers:
file:
platform:
directory: /etc/cerbix/monitoring.d
scope:
type: instance
payments:
directory: /etc/cerbix/payments.d
orphan_grace_period: 10m
scope:
type: project
organization: acme
project: payments
Key Default Allowed
directory — (required) Absolute path, not /, must not overlap another provider’s directory
scope.type — (required) instance, organization (needs scope.organization), project (needs scope.organization and scope.project)
debounce 2s 100ms to 30s
resync_interval 30s 5s to 1h
orphan_grace_period 30s 0 to 24h; 0 disables on the first valid absence
limits.max_files 1000 Positive, bounded
limits.max_file_bytes 1048576 Positive, bounded
limits.max_total_bytes 16777216 Positive, bounded
limits.max_monitors_per_bundle 1000 Positive, bounded
limits.max_managed_monitors 5000 Positive, bounded

A provider name matches ^[a-z][a-z0-9-]{0,39}$; up to 64 providers are allowed. At startup, an api or all process exits with file_provider_startup_failed if a configured directory is missing or unreadable. Mount it read-only at the exact configured path.

The scope decides which header fields a bundle may carry:

Provider scope Bundle organization Bundle project
instance required required
organization forbidden required
project forbidden forbidden

Example data: a bundle for the payments project under an instance-scoped provider.

format: 1
organization: acme
project: payments
monitors:
checkout-api:
name: Checkout API
type: http
target: https://checkout.example.com/healthz
interval: 30s
timeout: 5s
failure_threshold: 3
conditions:
- "[STATUS] == 200"
- "[RESPONSE_TIME] < 800"
tags: [payments, tier-1]
depends_on: [checkout-db]
checkout-db:
name: Checkout database socket
type: tcp
target: db.example.internal:5432
interval: 30s
timeout: 5s
nightly-billing:
name: Nightly billing job
type: push
interval: 24h
grace: 1h

Rules the decoder enforces:

  • Only regular .yaml and .yml files directly in the directory are read; dotfiles and subdirectories are skipped. One YAML document per file.
  • format is required. Format 1 declares monitors; format 2 adds a services map and a per-monitor slug (see Services).
  • The key under monitors is the monitor’s stable identity (UID), matching ^[a-z][a-z0-9-]{0,62}$. Changing name updates the same monitor; changing the UID orphans the old one and creates a new one. Changing type for an existing UID is rejected.
  • Monitor fields: name (required), type (required), description, target, method, interval, timeout, retries, failure_threshold, confirm_interval, renotify, grace, conditions, tags, region, enabled, auto_incident, depends_on, and a typed settings map.
  • Defaults when a field is absent: method: GET, interval: 60s, timeout: 10s, retries: 0, failure_threshold: 1, region: core, enabled: true, auto_incident: true.
  • Durations are whole seconds: 30s, 1m, 1h. 1500ms is rejected.
  • depends_on names UIDs in the same bundle. Missing, self or cyclic references reject the bundle.
  • Unknown fields and server-owned fields (id, status and similar) reject the bundle.
  • Supported types: http, tcp, icmp, dns, tls, grpc, websocket, ssh, push, postgres, mysql, redis, rabbitmq, promql and async_canary. composite and synthetic are not available through files. See Check types.

Ownership is per monitor. A project can mix file-managed monitors with monitors created in the UI or API. A provider compares a bundle only against the monitors it owns, never against the whole project, and never adopts an existing monitor by name.

A file-managed monitor is read-only through normal CRUD. Editing, pausing or deleting it returns 409:

{
"error": "managed_by_file",
"management": { "source": "file", "provider": "platform", "uid": "checkout-api", "path": "payments.yaml", "read_only": true }
}

Live status, heartbeats, SLA and incidents keep working as for any monitor. Deleting a project or organization that owns file-managed resources also returns 409 managed_by_file.

Change detection polls the directory signature (names, sizes, modification times), debounces, and rescans. A full resync runs every resync_interval regardless.

Situation Result
Valid change One PostgreSQL transaction per project applies monitors, dependency edges, provenance, the bundle generation and an audit record (file_provider.apply).
Comments, key order, file rename, a touch No monitor write; the execution revision does not move. A rename only updates the recorded source path.
Changed probe settings Update through the normal config-write path; the execution revision increments and the scheduler is notified in the same transaction.
Only depends_on changed Dependency graph updated; no execution-revision bump.
Invalid bundle Rejected as a whole. The project keeps its last-known-good state and its orphan clock is frozen.
Invalid file whose project cannot be determined Orphaning is suspended for the whole provider for that scan; other valid bundles still create and update.
Two files claim one project Both are rejected; the project keeps its last-known-good state.
Directory missing or unreadable after startup Not treated as deletion. Last-known-good stays.
A UID or whole file absent from a valid scan The monitor is marked orphaned. After orphan_grace_period, a later valid scan disables it. History is never deleted.
The UID returns Restored with the same monitor ID and push token; the orphan mark clears.

A provider never hard-deletes monitors. If you narrow a provider’s scope, monitors outside the new scope stay running and owned; they are not orphaned.

Every api replica runs a candidate loop for each provider, and exactly one replica applies at a time. Leadership is a PostgreSQL advisory lock per provider. The apply transaction runs on the leader’s locked connection, so a replica that loses the lock cannot commit. Followers stay ready.

  • GET /api/v1/admin/file-providers (global admin) returns {bundles, providers}: per-bundle status, last error and generation, plus this process’s runtime view of each provider (leadership, last scan, last success, counts). Leadership is process-local, so query each replica to find the leader. Add ?provider=<name> to filter.
  • GET /api/v1/organizations/{orgID}/file-providers (organization admin) returns {bundles} for that organization only.

Metrics carry only the provider label: cerbix_file_provider_leader, cerbix_file_provider_reconcile_total{outcome} (applied, noop, rejected, error), cerbix_file_provider_reconcile_duration_seconds, cerbix_file_provider_last_success_timestamp_seconds, cerbix_file_provider_managed_monitors, cerbix_file_provider_orphaned_monitors and cerbix_file_provider_bundle_errors.

The repository ships three Prometheus rules in docker/alerts/monitoring-as-code.rules.yml: no clean reconcile for over 90 s, bundle errors for 15 min, and no leader for 10 min. Retune the first to about three times your resync_interval.