Skip to content
v0.3.8GitHub

Change intelligence

Verified against cerbix f5240f5Report a problem ↗

Change intelligence lets a pipeline record that a release, rollback or flag change happened to a Service. The Service’s sealed facts then answer three questions: what changed and when, which changes preceded an incident, and how the SLI read before and after. The feature stores events, not a deployment catalog, and never claims that a change caused anything — only that it preceded something, with the lag stated.

Terminal window
export CERBIX_URL=https://cerbix.example.com
export CERBIX_TOKEN="<api-token>"
cerbix change record --project <project-id> --service <service-id> \
--kind deploy --phase succeeded \
--source github-actions --external-id 123456 --ref v1.2.3
Flag Rule
--project, --service Required IDs.
--kind Required: deploy, rollback or flag. Record a configuration change as deploy with a --ref.
--phase Required: started, succeeded, failed or cancelled.
--source Required slug of the reporting system, matching ^[a-z0-9][a-z0-9-]{0,63}$.
--external-id Required ID of the change at the source, 1–128 characters, case-sensitive.
--ref Optional label up to 128 characters, such as a version or commit.
--url Optional https:// link up to 512 characters; http:// is refused.
--decision Optional gate decision_id this release rested on. It must be a decision of this Service.
--at When the phase occurred, RFC3339. Defaults to the invocation instant.
--json Print the API response verbatim.
--timeout Overall request deadline, default 10s.

CERBIX_URL, CERBIX_TOKEN and CERBIX_CA_FILE work exactly as for cerbix gate check: environment only, TLS always verified, one request, no retries, no redirects. Text fields are normalized to Unicode NFC and trimmed on the server; control characters are refused.

The command calls POST /api/v1/projects/{projectID}/services/{serviceID}/changes, which requires the change:record action (editor and above).

A change is identified by Service, source and external-id. Its phases follow one order:

  • started may be followed by exactly one terminal phase: succeeded, failed or cancelled.
  • A terminal phase without started is accepted, for pipelines that only report the end.
  • A second terminal phase, or started after a terminal, is 409 phase_order.
  • A phase with a different kind from the rest of the change is 409 kind_mismatch.
  • A terminal phase earlier than its started is 400 occurred_at_before_start.
  • --at more than change.max_past (default 24 h) behind or change.max_future (default 5 min) ahead of the server clock is 400 occurred_at_out_of_bounds.

Writes for one identity are serialized, so two jobs reporting different terminal phases at the same moment cannot both succeed. Use a new external-id for every run: reusing one from an earlier run mixes two runs into one change, and its phases are refused with 409.

Sending the same phase again with identical kind, occurred_at, ref, url and decision_id returns 200 with the original row and prints replayed. A replay with any of those fields changed is 409 phase_exists, naming the field.

stdout is one line:

recorded change=<id> kind=<kind> phase=<phase>
replayed change=<id> kind=<kind> phase=<phase>

Refusals are printed verbatim on stderr.

Exit When
0 Recorded (201) or replayed (200)
2 Refused by the contract (400, 404, 409) or a usage error, such as a missing required flag
1 Missing variable, transport, TLS or timeout failure, 401, 403, 429 (with Retry-After printed), 5xx, a redirect, or a malformed response

Use a token with "role": "editor" and "actions": ["gate:evaluate", "change:record"]. This example links the release to its gate decision and records both ends of the deploy. Example data:

env:
CERBIX_URL: https://cerbix.example.com
CERBIX_TOKEN: ${{ secrets.CERBIX_TOKEN }}
CHANGE_ID: ${{ github.run_id }}-${{ github.run_attempt }}
steps:
- name: Reliability gate
id: gate
run: |
cerbix gate check --project "${{ vars.CERBIX_PROJECT }}" --service "${{ vars.CERBIX_SERVICE }}" --json > gate.json
echo "decision_id=$(jq -r .decision_id gate.json)" >> "$GITHUB_OUTPUT"
- name: Record deploy start
run: |
cerbix change record --project "${{ vars.CERBIX_PROJECT }}" --service "${{ vars.CERBIX_SERVICE }}" \
--kind deploy --phase started --source github-actions --external-id "$CHANGE_ID" \
--ref "${{ github.sha }}" --decision "${{ steps.gate.outputs.decision_id }}"
# … deploy steps …
- name: Record deploy result
if: always() && steps.gate.outcome == 'success'
run: |
cerbix change record --project "${{ vars.CERBIX_PROJECT }}" --service "${{ vars.CERBIX_SERVICE }}" \
--kind deploy --phase "${{ job.status == 'success' && 'succeeded' || job.status == 'cancelled' && 'cancelled' || 'failed' }}" \
--source github-actions --external-id "$CHANGE_ID" --ref "${{ github.sha }}"

The gate step fails the job on BLOCK (exit 2) and on NOT_CONFIGURED (exit 4). Handle exit 4 explicitly, as shown on the Release gate page, if Services without a policy should deploy.

Timeline. GET /api/v1/projects/{projectID}/services/{serviceID}/changes?from=<RFC3339>&to=<RFC3339> returns one group per change with its phases nested, the linked gate decision and the incidents it preceded. The range is at most 92 days. Optional kind (repeatable), source, limit (1–200, default 50) and cursor narrow it. A linked decision shows its state and action while the ledger still holds it, and aged_out after. The Service page shows the same timeline.

Changes that preceded an incident. When a Service auto-incident opens, cerbix links every change whose latest phase falls within change.correlation_window (default 60 min) before the open, on that Service (own_service) and on Services the incident marks as probable roots (upstream). It adds one 🚀 Changes: note to the timeline, naming up to change.correlation_note_max (default 5) changes and counting the rest. Links are fixed at open; a change recorded later is not added. Read them at GET /api/v1/projects/{projectID}/incidents/{incidentID}/changes. If correlation fails, the incident opens and resolves as usual and the failure is counted in cerbix_change_correlation_errors_total.

Before and after. GET …/services/{serviceID}/changes/compare?source=<slug>&external_id=<id>&horizon=1h compares the SLI over [T − h, T) and [T, T + h), where T is the terminal phase’s time floored to the minute and horizon is 15m, 1h (default), 6h or 24h. Each side is exactly one of:

Shape Meaning
Figure: availability, good_seconds, bad_seconds, unknown_seconds, excluded_seconds, buckets Sum of sealed buckets, the reliability page’s own arithmetic.
pending: true with sealed_through The side ends after the seal watermark. Ask again once it passes.
withheld: definition_changed, undecidable (with detail) or no_facts The reliability page would not show a number for this range either.

delta (after minus before, in availability points) is present only when both sides are figures. A change with only started has no comparison: 404 no_terminal_phase. Never read a pending or withheld side as zero. See How the SLI is computed.

Changes are kept for change.retention_days (default 400 days), judged by each change’s latest phase. Deleting a Service deletes its changes; existing 🚀 Changes: notes remain as text.