Monitoring as Code
f5240f5Report a problem ↗A file provider watches a directory of YAML bundles and reconciles them into cerbix monitors at runtime. You change monitors by changing files; no restart or reload is needed. This page covers the bundle format, ownership, how reconciliation behaves when files are wrong or missing, and how to operate providers in a cluster.
How it works
Section titled “How it works”- One bundle file describes the monitors of one project.
- Files are the desired state for the monitors a provider owns. PostgreSQL stays the store for runtime state and history.
- Only the
apiandallroles run providers. The scheduler and workers never read YAML; they see ordinary monitors. - A provider never creates organizations or projects. The slugs a bundle names must already exist.
Configure a provider
Section titled “Configure a provider”Providers are declared in the static config under providers.file.<name>. Changing a provider’s name, scope, directory or limits needs a restart; the files inside the directory do not.
providers: file: platform: directory: /etc/cerbix/monitoring.d scope: type: instance payments: directory: /etc/cerbix/payments.d orphan_grace_period: 10m scope: type: project organization: acme project: payments| Key | Default | Allowed |
|---|---|---|
directory |
— (required) | Absolute path, not /, must not overlap another provider’s directory |
scope.type |
— (required) | instance, organization (needs scope.organization), project (needs scope.organization and scope.project) |
debounce |
2s |
100ms to 30s |
resync_interval |
30s |
5s to 1h |
orphan_grace_period |
30s |
0 to 24h; 0 disables on the first valid absence |
limits.max_files |
1000 |
Positive, bounded |
limits.max_file_bytes |
1048576 |
Positive, bounded |
limits.max_total_bytes |
16777216 |
Positive, bounded |
limits.max_monitors_per_bundle |
1000 |
Positive, bounded |
limits.max_managed_monitors |
5000 |
Positive, bounded |
A provider name matches ^[a-z][a-z0-9-]{0,39}$; up to 64 providers are allowed. At startup, an api or all process exits with file_provider_startup_failed if a configured directory is missing or unreadable. Mount it read-only at the exact configured path.
The scope decides which header fields a bundle may carry:
| Provider scope | Bundle organization |
Bundle project |
|---|---|---|
instance |
required | required |
organization |
forbidden | required |
project |
forbidden | forbidden |
Write a bundle
Section titled “Write a bundle”Example data: a bundle for the payments project under an instance-scoped provider.
format: 1organization: acmeproject: payments
monitors: checkout-api: name: Checkout API type: http target: https://checkout.example.com/healthz interval: 30s timeout: 5s failure_threshold: 3 conditions: - "[STATUS] == 200" - "[RESPONSE_TIME] < 800" tags: [payments, tier-1] depends_on: [checkout-db]
checkout-db: name: Checkout database socket type: tcp target: db.example.internal:5432 interval: 30s timeout: 5s
nightly-billing: name: Nightly billing job type: push interval: 24h grace: 1hRules the decoder enforces:
- Only regular
.yamland.ymlfiles directly in the directory are read; dotfiles and subdirectories are skipped. One YAML document per file. formatis required. Format1declares monitors; format2adds aservicesmap and a per-monitorslug(see Services).- The key under
monitorsis the monitor’s stable identity (UID), matching^[a-z][a-z0-9-]{0,62}$. Changingnameupdates the same monitor; changing the UID orphans the old one and creates a new one. Changingtypefor an existing UID is rejected. - Monitor fields:
name(required),type(required),description,target,method,interval,timeout,retries,failure_threshold,confirm_interval,renotify,grace,conditions,tags,region,enabled,auto_incident,depends_on, and a typedsettingsmap. - Defaults when a field is absent:
method: GET,interval: 60s,timeout: 10s,retries: 0,failure_threshold: 1,region: core,enabled: true,auto_incident: true. - Durations are whole seconds:
30s,1m,1h.1500msis rejected. depends_onnames UIDs in the same bundle. Missing, self or cyclic references reject the bundle.- Unknown fields and server-owned fields (
id,statusand similar) reject the bundle. - Supported types:
http,tcp,icmp,dns,tls,grpc,websocket,ssh,push,postgres,mysql,redis,rabbitmq,promqlandasync_canary.compositeandsyntheticare not available through files. See Check types.
Ownership
Section titled “Ownership”Ownership is per monitor. A project can mix file-managed monitors with monitors created in the UI or API. A provider compares a bundle only against the monitors it owns, never against the whole project, and never adopts an existing monitor by name.
A file-managed monitor is read-only through normal CRUD. Editing, pausing or deleting it returns 409:
{ "error": "managed_by_file", "management": { "source": "file", "provider": "platform", "uid": "checkout-api", "path": "payments.yaml", "read_only": true }}Live status, heartbeats, SLA and incidents keep working as for any monitor. Deleting a project or organization that owns file-managed resources also returns 409 managed_by_file.
Reconciliation semantics
Section titled “Reconciliation semantics”Change detection polls the directory signature (names, sizes, modification times), debounces, and rescans. A full resync runs every resync_interval regardless.
| Situation | Result |
|---|---|
| Valid change | One PostgreSQL transaction per project applies monitors, dependency edges, provenance, the bundle generation and an audit record (file_provider.apply). |
| Comments, key order, file rename, a touch | No monitor write; the execution revision does not move. A rename only updates the recorded source path. |
| Changed probe settings | Update through the normal config-write path; the execution revision increments and the scheduler is notified in the same transaction. |
Only depends_on changed |
Dependency graph updated; no execution-revision bump. |
| Invalid bundle | Rejected as a whole. The project keeps its last-known-good state and its orphan clock is frozen. |
| Invalid file whose project cannot be determined | Orphaning is suspended for the whole provider for that scan; other valid bundles still create and update. |
| Two files claim one project | Both are rejected; the project keeps its last-known-good state. |
| Directory missing or unreadable after startup | Not treated as deletion. Last-known-good stays. |
| A UID or whole file absent from a valid scan | The monitor is marked orphaned. After orphan_grace_period, a later valid scan disables it. History is never deleted. |
| The UID returns | Restored with the same monitor ID and push token; the orphan mark clears. |
A provider never hard-deletes monitors. If you narrow a provider’s scope, monitors outside the new scope stay running and owned; they are not orphaned.
High availability
Section titled “High availability”Every api replica runs a candidate loop for each provider, and exactly one replica applies at a time. Leadership is a PostgreSQL advisory lock per provider. The apply transaction runs on the leader’s locked connection, so a replica that loses the lock cannot commit. Followers stay ready.
Diagnostics and alerts
Section titled “Diagnostics and alerts”GET /api/v1/admin/file-providers(global admin) returns{bundles, providers}: per-bundle status, last error and generation, plus this process’s runtime view of each provider (leadership, last scan, last success, counts). Leadership is process-local, so query each replica to find the leader. Add?provider=<name>to filter.GET /api/v1/organizations/{orgID}/file-providers(organization admin) returns{bundles}for that organization only.
Metrics carry only the provider label: cerbix_file_provider_leader, cerbix_file_provider_reconcile_total{outcome} (applied, noop, rejected, error), cerbix_file_provider_reconcile_duration_seconds, cerbix_file_provider_last_success_timestamp_seconds, cerbix_file_provider_managed_monitors, cerbix_file_provider_orphaned_monitors and cerbix_file_provider_bundle_errors.
The repository ships three Prometheus rules in docker/alerts/monitoring-as-code.rules.yml: no clean reconcile for over 90 s, bundle errors for 15 min, and no leader for 10 min. Retune the first to about three times your resync_interval.