Skip to content
v0.3.8GitHub

Metrics and alerts

Verified against cerbix f5240f5Report a problem ↗

Every cerbix process exposes health, readiness and Prometheus metrics on its HTTP listener. This page lists the endpoints, every cerbix_* metric family and the alerting rules shipped in the repository.

Every role, including agent, serves these on server.listen (default :8080). The paths are configurable.

Endpoint Config key Response
/healthz server.healthz_path Always 200 {"status":"ok"} while the process serves HTTP. Use it for liveness.
/readyz server.readyz_path 200 {"status":"ready"}, or 503 {"status":"not_ready","error":"<reason>"}. Use it for readiness.
/metrics server.metrics_path Prometheus text format, version 0.0.4.

/readyz returns 503 in these cases:

Condition Reported reason
The process is starting or shutting down. shutting down during shutdown
The periodic database ping (every 10 s) fails. database unreachable
A worker or agent cannot decrypt credential envelopes with its keyring. credential envelope decrypt unavailable
Service-reliability work is wedged (scheduler leader). The wedge reason
A service alerting evaluator has stalled (scheduler leader). The stall reason

cerbix_ready reports the same verdict as /readyz.

Labels are low-cardinality by design: no metric carries a tenant, project, service, monitor or token label. Those details go to logs.

Metric Type Labels Meaning
cerbix_build_info gauge version, commit, go_version, role Always 1; carries build metadata.
cerbix_up gauge — Always 1 while serving.
cerbix_ready gauge — 1 when ready, 0 otherwise.
cerbix_uptime_seconds gauge — Seconds since process start.
cerbix_database_up gauge — Database reachable. Exported only when a database is configured.
cerbix_broker_up gauge — AMQP broker reachable. Exported only by processes that use RabbitMQ.
cerbix_scheduler_leader gauge — 1 when this process holds scheduler leadership, 0 on standby.
cerbix_dispatch_shared_trust gauge — 1 when one fallback dispatch key can open several regions’ credential payloads.
Metric Type Labels Meaning
cerbix_checks_total counter result (up, down) Monitor checks recorded.
cerbix_result_quarantined_total counter reason Results set aside without touching state, for example future_timestamp.
cerbix_result_ignored_total counter reason Results not applied to live state: out_of_order, outside_retention.
cerbix_result_rejected_total counter reason Results rejected with no insert: missing_timestamp, stale_revision, missing_revision.
cerbix_result_clock_skew_total counter origin, reason Accepted results with an anomalous client clock.
cerbix_result_missing_revision_total counter — Scheduled results without a revision, applied in observe mode.
cerbix_result_observed_before_issue_total counter — Results accepted with an observation time before the job’s issue time, within result.allowed_skew.
cerbix_executor_probe_error_total counter reason Credential-envelope execution errors, such as no_dispatch_key, unknown_key_id, decrypt_auth_failed.
cerbix_secret_resolution_failed_total counter reason Credential materialization or capability refusals before dispatch, such as no_capable_executor.
cerbix_dispatch_transport_failed_total counter reason Jobs that could not be handed to their transport.
cerbix_canary_stage_total counter kind, stage, outcome Canary executions by workflow kind, deciding stage and outcome class.
cerbix_canary_dispatch_refused_total counter reason Canary runs not dispatched.
cerbix_pull_jobs_pending gauge region Unclaimed HTTP-pull jobs per region.
cerbix_pull_agent_lag_seconds gauge region Age of the oldest unclaimed pull job per region.
cerbix_expected_runs_updates gauge — Updates to expected-run rows across retained partitions. Falls when a partition is dropped.
cerbix_expected_runs_hot_updates gauge — The subset of those updates that were HOT.
cerbix_expected_runs_hot_update_ratio gauge — HOT updates over all updates. Not exported until an update has been observed.
Metric Type Labels Meaning
cerbix_incidents_opened_total counter — Incidents opened through the API.
cerbix_outbox_delivered_total counter — Outbox events delivered.
cerbix_outbox_dead_total counter — Outbox events parked as dead after exhausting retries.
cerbix_alert_suppressed_total counter topic, reason Monitor alerts withheld because a service covers that signal.
cerbix_alert_delegation_fail_open_total counter reason Deliveries that paged because service coverage could not be confirmed.
cerbix_escalation_steps_total counter subject (monitor, service) On-call ladder steps enqueued.
cerbix_status_page_component_unreadable_total counter — Status-page components whose service could not be read.
Metric Type Labels Meaning
cerbix_service_repair_ranges gauge state (pending, running, error) Durable repair ranges by state.
cerbix_service_repair_outcomes_total counter outcome, reason Repair ranges reaching a lifecycle outcome, by range reason.
cerbix_service_watermark_lag_seconds gauge — Worst sealed-watermark lag across services. Steady state is near the late-arrival grace, never 0.
cerbix_service_slices_total counter outcome (worked, empty, error) Leader materialization slices.
cerbix_service_wedged gauge — 1 when service-reliability work is wedged. Fails /readyz.
cerbix_service_unrecomputable_rejections_total counter — Mutations refused because their range cannot be recomputed.
cerbix_service_epoch_fanout_total counter — Evaluation epochs created by execution-changing writes.
cerbix_service_late_arrivals_total counter — Heartbeats excluded because they arrived behind the seal.
cerbix_service_late_arrival_overflow_total counter — Late-arrival example slots that overflowed their bound.
cerbix_service_fact_maintenance_failing gauge — 1 when the last fact-partition maintenance pass failed.
cerbix_service_fact_maintenance_last_success_timestamp_seconds gauge — Unix time of the last successful maintenance pass.
cerbix_service_impact_links_total counter role New incident-to-service impact links.
cerbix_service_impact_correlation_failures_total counter — Failed impact-correlation attempts; each retries through the outbox.
cerbix_service_impact_witness_overflow_total counter — Open incidents beyond the per-service correlation bound.

All families here are labelled by signal (health or burn) plus at most one fixed enum.

Metric Type Labels Meaning
cerbix_service_alert_evaluations_total counter signal, outcome Evaluations by outcome (ok, error, skipped).
cerbix_service_alert_emitted_total counter signal, edge Alert edges enqueued (onset, close).
cerbix_service_alert_withheld_total counter signal, reason Onsets not announced because nothing could receive them.
cerbix_service_alert_delegation_total counter signal, state Delegation outcomes: armed, disarmed, degraded.
cerbix_service_alert_undeliverable_total counter signal Announcements with no recipient left.
cerbix_service_alert_recipient_missing_total counter — Snapshot recipients whose channel no longer exists.
cerbix_service_incidents_total counter action (opened, resolved) Service incidents opened or resolved automatically.
cerbix_service_alert_active gauge signal Open alert episodes.
cerbix_service_alert_backlog gauge signal Services or burn targets due for evaluation.
cerbix_service_alert_last_success_seconds gauge signal Unix time of the last successful evaluation pass.
cerbix_service_alert_lag_seconds gauge signal Staleness of the oldest verdict in the last successful pass.
Metric Type Labels Meaning
cerbix_gate_decisions_total counter state, action, policy_source, window_mode, overridden Gate decisions. NOT_CONFIGURED carries action="none", policy_source="none", window_mode="none".
cerbix_gate_evaluate_rejected_total counter reason Evaluations refused with 429: process_inflight, principal_inflight, process_rate, principal_rate.
cerbix_gate_evaluate_errors_total counter kind Admitted evaluations that failed: snapshot_conflict, timeout, ledger_unwritable, error.
cerbix_gate_decision_duration_seconds histogram — Evaluation wall time. Buckets 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30 s.
cerbix_gate_maintenance_errors_total counter kind Ledger maintenance failures: lock_timeout, statement_timeout, partition_identity, error.
cerbix_gate_decisions_writable_horizon_seconds gauge — Seconds until the newest ledger partition ends. Decisions stop being accepted at 0.
cerbix_gate_decisions_partitions_pending_drop gauge — Ledger partitions past retention and not yet dropped.
cerbix_gate_decisions_oldest_partition_age_seconds gauge — Age of the oldest attached partition past the cutoff; 0 when none.
cerbix_gate_decisions_bytes gauge — Total size of ledger partitions not yet dropped.
Metric Type Labels Meaning
cerbix_changes_recorded_total counter kind, phase, outcome (recorded, replayed) Change phases accepted.
cerbix_change_record_rejected_total counter reason Records refused, by refusal code or rate-limit reason.
cerbix_change_correlations_total counter role (own_service, upstream) Incident-to-change links inserted.
cerbix_change_correlation_errors_total counter — Failed correlation attempts. The incident opens regardless.
cerbix_change_compare_total counter outcome (figure, withheld, pending) Before/after comparisons served.
cerbix_changes_retained gauge — Change rows kept after the last retention pass.
Metric Type Labels Meaning
cerbix_audit_retention_passes_total counter result (deleted, empty, lock_busy, error, budget) Retention passes.
cerbix_audit_retention_rows_deleted_total counter — Audit rows deleted.
cerbix_audit_retention_purge_interval_seconds gauge — Configured interval between passes.
cerbix_audit_retention_configured_timestamp_seconds gauge — When retention was configured in this process.
cerbix_audit_retention_last_success_timestamp_seconds gauge — Database time of the last successful pass.
cerbix_audit_retention_oldest_expired_seconds gauge — How far past the cutoff the oldest surviving expired row is.

Every family carries a provider label, bounded by the configured providers. See Monitoring as code.

Metric Type Extra labels Meaning
cerbix_file_provider_leader gauge — 1 when this process leads the provider’s reconcile.
cerbix_file_provider_reconcile_total counter outcome (applied, noop, rejected, error) Reconciles.
cerbix_file_provider_reconcile_duration_seconds gauge — Duration of the last reconcile.
cerbix_file_provider_last_success_timestamp_seconds gauge — Unix time of the last clean reconcile.
cerbix_file_provider_managed_monitors gauge — Monitors owned by the provider.
cerbix_file_provider_orphaned_monitors gauge — Owned monitors that are orphaned.
cerbix_file_provider_bundle_errors gauge — Bundles or files rejected in the last scan.

The repository ships Prometheus rule files. Load them through rule_files: and validate them with promtool check rules <file>.

Alert File Fires when For Severity
CerbixAuditRetentionStale docs/alerts.yaml No successful audit-retention pass for two purge intervals, including when none has ever succeeded. 15m warning
CerbixAuditRetentionBacklog docs/alerts.yaml Expired audit rows survive longer than two purge intervals past the cutoff, or the backlog gauge is still absent two purge intervals after configuration. 15m warning
CerbixFileProviderStaleSuccess docker/alerts/monitoring-as-code.rules.yml No clean reconcile for a provider in over 90 s. 10m warning
CerbixFileProviderBundleErrors docker/alerts/monitoring-as-code.rules.yml A provider rejected bundles on its last scan. 15m warning
CerbixFileProviderNoLeader docker/alerts/monitoring-as-code.rules.yml No replica leads a provider’s reconcile. 10m critical
CerbixCredentialDispatchUnavailable docker/alerts/secret-inventory.rules.yml cerbix_secret_resolution_failed_total{reason="no_capable_executor"} increased in 10 min. 5m critical
CerbixCredentialEnvelopeFailures docker/alerts/secret-inventory.rules.yml cerbix_executor_probe_error_total increased for no_dispatch_key, unknown_key_id or decrypt_auth_failed. 5m critical
CerbixDispatchSharedTrustEnabled docker/alerts/secret-inventory.rules.yml cerbix_dispatch_shared_trust == 1. 1m info

The 90 s threshold of CerbixFileProviderStaleSuccess assumes the default 30 s resync_interval. Set it to about three times your configured interval. The Runbook lists further suggested expressions for service reliability and alerting that are not shipped as rule files.

Rule files on GitHub: docs/alerts.yaml, docker/alerts/monitoring-as-code.rules.yml, docker/alerts/secret-inventory.rules.yml.