Every cerbix process exposes health, readiness and Prometheus metrics on its HTTP listener. This page lists the endpoints, every cerbix_* metric family and the alerting rules shipped in the repository.
Every role, including agent, serves these on server.listen (default :8080). The paths are configurable.
| Endpoint |
Config key |
Response |
/healthz |
server.healthz_path |
Always 200 {"status":"ok"} while the process serves HTTP. Use it for liveness. |
/readyz |
server.readyz_path |
200 {"status":"ready"}, or 503 {"status":"not_ready","error":"<reason>"}. Use it for readiness. |
/metrics |
server.metrics_path |
Prometheus text format, version 0.0.4. |
/readyz returns 503 in these cases:
| Condition |
Reported reason |
| The process is starting or shutting down. |
shutting down during shutdown |
| The periodic database ping (every 10 s) fails. |
database unreachable |
| A worker or agent cannot decrypt credential envelopes with its keyring. |
credential envelope decrypt unavailable |
| Service-reliability work is wedged (scheduler leader). |
The wedge reason |
| A service alerting evaluator has stalled (scheduler leader). |
The stall reason |
cerbix_ready reports the same verdict as /readyz.
Labels are low-cardinality by design: no metric carries a tenant, project, service, monitor or token label. Those details go to logs.
| Metric |
Type |
Labels |
Meaning |
cerbix_build_info |
gauge |
version, commit, go_version, role |
Always 1; carries build metadata. |
cerbix_up |
gauge |
— |
Always 1 while serving. |
cerbix_ready |
gauge |
— |
1 when ready, 0 otherwise. |
cerbix_uptime_seconds |
gauge |
— |
Seconds since process start. |
cerbix_database_up |
gauge |
— |
Database reachable. Exported only when a database is configured. |
cerbix_broker_up |
gauge |
— |
AMQP broker reachable. Exported only by processes that use RabbitMQ. |
cerbix_scheduler_leader |
gauge |
— |
1 when this process holds scheduler leadership, 0 on standby. |
cerbix_dispatch_shared_trust |
gauge |
— |
1 when one fallback dispatch key can open several regions’ credential payloads. |
| Metric |
Type |
Labels |
Meaning |
cerbix_checks_total |
counter |
result (up, down) |
Monitor checks recorded. |
cerbix_result_quarantined_total |
counter |
reason |
Results set aside without touching state, for example future_timestamp. |
cerbix_result_ignored_total |
counter |
reason |
Results not applied to live state: out_of_order, outside_retention. |
cerbix_result_rejected_total |
counter |
reason |
Results rejected with no insert: missing_timestamp, stale_revision, missing_revision. |
cerbix_result_clock_skew_total |
counter |
origin, reason |
Accepted results with an anomalous client clock. |
cerbix_result_missing_revision_total |
counter |
— |
Scheduled results without a revision, applied in observe mode. |
cerbix_result_observed_before_issue_total |
counter |
— |
Results accepted with an observation time before the job’s issue time, within result.allowed_skew. |
cerbix_executor_probe_error_total |
counter |
reason |
Credential-envelope execution errors, such as no_dispatch_key, unknown_key_id, decrypt_auth_failed. |
cerbix_secret_resolution_failed_total |
counter |
reason |
Credential materialization or capability refusals before dispatch, such as no_capable_executor. |
cerbix_dispatch_transport_failed_total |
counter |
reason |
Jobs that could not be handed to their transport. |
cerbix_canary_stage_total |
counter |
kind, stage, outcome |
Canary executions by workflow kind, deciding stage and outcome class. |
cerbix_canary_dispatch_refused_total |
counter |
reason |
Canary runs not dispatched. |
cerbix_pull_jobs_pending |
gauge |
region |
Unclaimed HTTP-pull jobs per region. |
cerbix_pull_agent_lag_seconds |
gauge |
region |
Age of the oldest unclaimed pull job per region. |
cerbix_expected_runs_updates |
gauge |
— |
Updates to expected-run rows across retained partitions. Falls when a partition is dropped. |
cerbix_expected_runs_hot_updates |
gauge |
— |
The subset of those updates that were HOT. |
cerbix_expected_runs_hot_update_ratio |
gauge |
— |
HOT updates over all updates. Not exported until an update has been observed. |
| Metric |
Type |
Labels |
Meaning |
cerbix_incidents_opened_total |
counter |
— |
Incidents opened through the API. |
cerbix_outbox_delivered_total |
counter |
— |
Outbox events delivered. |
cerbix_outbox_dead_total |
counter |
— |
Outbox events parked as dead after exhausting retries. |
cerbix_alert_suppressed_total |
counter |
topic, reason |
Monitor alerts withheld because a service covers that signal. |
cerbix_alert_delegation_fail_open_total |
counter |
reason |
Deliveries that paged because service coverage could not be confirmed. |
cerbix_escalation_steps_total |
counter |
subject (monitor, service) |
On-call ladder steps enqueued. |
cerbix_status_page_component_unreadable_total |
counter |
— |
Status-page components whose service could not be read. |
| Metric |
Type |
Labels |
Meaning |
cerbix_service_repair_ranges |
gauge |
state (pending, running, error) |
Durable repair ranges by state. |
cerbix_service_repair_outcomes_total |
counter |
outcome, reason |
Repair ranges reaching a lifecycle outcome, by range reason. |
cerbix_service_watermark_lag_seconds |
gauge |
— |
Worst sealed-watermark lag across services. Steady state is near the late-arrival grace, never 0. |
cerbix_service_slices_total |
counter |
outcome (worked, empty, error) |
Leader materialization slices. |
cerbix_service_wedged |
gauge |
— |
1 when service-reliability work is wedged. Fails /readyz. |
cerbix_service_unrecomputable_rejections_total |
counter |
— |
Mutations refused because their range cannot be recomputed. |
cerbix_service_epoch_fanout_total |
counter |
— |
Evaluation epochs created by execution-changing writes. |
cerbix_service_late_arrivals_total |
counter |
— |
Heartbeats excluded because they arrived behind the seal. |
cerbix_service_late_arrival_overflow_total |
counter |
— |
Late-arrival example slots that overflowed their bound. |
cerbix_service_fact_maintenance_failing |
gauge |
— |
1 when the last fact-partition maintenance pass failed. |
cerbix_service_fact_maintenance_last_success_timestamp_seconds |
gauge |
— |
Unix time of the last successful maintenance pass. |
cerbix_service_impact_links_total |
counter |
role |
New incident-to-service impact links. |
cerbix_service_impact_correlation_failures_total |
counter |
— |
Failed impact-correlation attempts; each retries through the outbox. |
cerbix_service_impact_witness_overflow_total |
counter |
— |
Open incidents beyond the per-service correlation bound. |
All families here are labelled by signal (health or burn) plus at most one fixed enum.
| Metric |
Type |
Labels |
Meaning |
cerbix_service_alert_evaluations_total |
counter |
signal, outcome |
Evaluations by outcome (ok, error, skipped). |
cerbix_service_alert_emitted_total |
counter |
signal, edge |
Alert edges enqueued (onset, close). |
cerbix_service_alert_withheld_total |
counter |
signal, reason |
Onsets not announced because nothing could receive them. |
cerbix_service_alert_delegation_total |
counter |
signal, state |
Delegation outcomes: armed, disarmed, degraded. |
cerbix_service_alert_undeliverable_total |
counter |
signal |
Announcements with no recipient left. |
cerbix_service_alert_recipient_missing_total |
counter |
— |
Snapshot recipients whose channel no longer exists. |
cerbix_service_incidents_total |
counter |
action (opened, resolved) |
Service incidents opened or resolved automatically. |
cerbix_service_alert_active |
gauge |
signal |
Open alert episodes. |
cerbix_service_alert_backlog |
gauge |
signal |
Services or burn targets due for evaluation. |
cerbix_service_alert_last_success_seconds |
gauge |
signal |
Unix time of the last successful evaluation pass. |
cerbix_service_alert_lag_seconds |
gauge |
signal |
Staleness of the oldest verdict in the last successful pass. |
| Metric |
Type |
Labels |
Meaning |
cerbix_gate_decisions_total |
counter |
state, action, policy_source, window_mode, overridden |
Gate decisions. NOT_CONFIGURED carries action="none", policy_source="none", window_mode="none". |
cerbix_gate_evaluate_rejected_total |
counter |
reason |
Evaluations refused with 429: process_inflight, principal_inflight, process_rate, principal_rate. |
cerbix_gate_evaluate_errors_total |
counter |
kind |
Admitted evaluations that failed: snapshot_conflict, timeout, ledger_unwritable, error. |
cerbix_gate_decision_duration_seconds |
histogram |
— |
Evaluation wall time. Buckets 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30 s. |
cerbix_gate_maintenance_errors_total |
counter |
kind |
Ledger maintenance failures: lock_timeout, statement_timeout, partition_identity, error. |
cerbix_gate_decisions_writable_horizon_seconds |
gauge |
— |
Seconds until the newest ledger partition ends. Decisions stop being accepted at 0. |
cerbix_gate_decisions_partitions_pending_drop |
gauge |
— |
Ledger partitions past retention and not yet dropped. |
cerbix_gate_decisions_oldest_partition_age_seconds |
gauge |
— |
Age of the oldest attached partition past the cutoff; 0 when none. |
cerbix_gate_decisions_bytes |
gauge |
— |
Total size of ledger partitions not yet dropped. |
| Metric |
Type |
Labels |
Meaning |
cerbix_changes_recorded_total |
counter |
kind, phase, outcome (recorded, replayed) |
Change phases accepted. |
cerbix_change_record_rejected_total |
counter |
reason |
Records refused, by refusal code or rate-limit reason. |
cerbix_change_correlations_total |
counter |
role (own_service, upstream) |
Incident-to-change links inserted. |
cerbix_change_correlation_errors_total |
counter |
— |
Failed correlation attempts. The incident opens regardless. |
cerbix_change_compare_total |
counter |
outcome (figure, withheld, pending) |
Before/after comparisons served. |
cerbix_changes_retained |
gauge |
— |
Change rows kept after the last retention pass. |
| Metric |
Type |
Labels |
Meaning |
cerbix_audit_retention_passes_total |
counter |
result (deleted, empty, lock_busy, error, budget) |
Retention passes. |
cerbix_audit_retention_rows_deleted_total |
counter |
— |
Audit rows deleted. |
cerbix_audit_retention_purge_interval_seconds |
gauge |
— |
Configured interval between passes. |
cerbix_audit_retention_configured_timestamp_seconds |
gauge |
— |
When retention was configured in this process. |
cerbix_audit_retention_last_success_timestamp_seconds |
gauge |
— |
Database time of the last successful pass. |
cerbix_audit_retention_oldest_expired_seconds |
gauge |
— |
How far past the cutoff the oldest surviving expired row is. |
Every family carries a provider label, bounded by the configured providers. See Monitoring as code.
| Metric |
Type |
Extra labels |
Meaning |
cerbix_file_provider_leader |
gauge |
— |
1 when this process leads the provider’s reconcile. |
cerbix_file_provider_reconcile_total |
counter |
outcome (applied, noop, rejected, error) |
Reconciles. |
cerbix_file_provider_reconcile_duration_seconds |
gauge |
— |
Duration of the last reconcile. |
cerbix_file_provider_last_success_timestamp_seconds |
gauge |
— |
Unix time of the last clean reconcile. |
cerbix_file_provider_managed_monitors |
gauge |
— |
Monitors owned by the provider. |
cerbix_file_provider_orphaned_monitors |
gauge |
— |
Owned monitors that are orphaned. |
cerbix_file_provider_bundle_errors |
gauge |
— |
Bundles or files rejected in the last scan. |
The repository ships Prometheus rule files. Load them through rule_files: and validate them with promtool check rules <file>.
| Alert |
File |
Fires when |
For |
Severity |
CerbixAuditRetentionStale |
docs/alerts.yaml |
No successful audit-retention pass for two purge intervals, including when none has ever succeeded. |
15m |
warning |
CerbixAuditRetentionBacklog |
docs/alerts.yaml |
Expired audit rows survive longer than two purge intervals past the cutoff, or the backlog gauge is still absent two purge intervals after configuration. |
15m |
warning |
CerbixFileProviderStaleSuccess |
docker/alerts/monitoring-as-code.rules.yml |
No clean reconcile for a provider in over 90 s. |
10m |
warning |
CerbixFileProviderBundleErrors |
docker/alerts/monitoring-as-code.rules.yml |
A provider rejected bundles on its last scan. |
15m |
warning |
CerbixFileProviderNoLeader |
docker/alerts/monitoring-as-code.rules.yml |
No replica leads a provider’s reconcile. |
10m |
critical |
CerbixCredentialDispatchUnavailable |
docker/alerts/secret-inventory.rules.yml |
cerbix_secret_resolution_failed_total{reason="no_capable_executor"} increased in 10 min. |
5m |
critical |
CerbixCredentialEnvelopeFailures |
docker/alerts/secret-inventory.rules.yml |
cerbix_executor_probe_error_total increased for no_dispatch_key, unknown_key_id or decrypt_auth_failed. |
5m |
critical |
CerbixDispatchSharedTrustEnabled |
docker/alerts/secret-inventory.rules.yml |
cerbix_dispatch_shared_trust == 1. |
1m |
info |
The 90 s threshold of CerbixFileProviderStaleSuccess assumes the default 30 s resync_interval. Set it to about three times your configured interval. The Runbook lists further suggested expressions for service reliability and alerting that are not shipped as rule files.
Rule files on GitHub: docs/alerts.yaml, docker/alerts/monitoring-as-code.rules.yml, docker/alerts/secret-inventory.rules.yml.