Changelog
teamlead-com/cerbix · CHANGELOG.md · v0.3.8Edit on GitHub ↗All notable changes to cerbix will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[v0.3.8] - 2026-10-08
Section titled “[v0.3.8] - 2026-10-08”✨ Added
Section titled “✨ Added”- Incident history on status pages. A status page now lists its ten most recent past incidents, and a View incident history link opens a history page that walks the last 90 days month by month (calendar months in UTC), up to 50 incidents at a time with Show more. Unlisted pages keep their token through the link. The page no longer grows with every resolved incident: what a page or history response carries is bounded, its incident lists stop early on the new index, and uncached public history requests are limited (HTTP 429 with
Retry-Afterbeyond the limit). The per-month counts still read the 90-day window (D-0270).
Changed
Section titled “Changed”- API:
recent_incidentson a status page render holds at most the 10 newest past incidents, with a newrecent_incidents_moreflag; the rest are served byGET /api/v1/public/status-pages/{slug}/historyandGET /api/v1/status-pages/{pageID}/history. A client that readrecent_incidentsas “every incident resolved in the last 90 days” should page through the history instead. Feeds, webhooks and subscriptions are unchanged.
🩹 Fixed
Section titled “🩹 Fixed”- Reliability readouts are no longer cut off by their card. Hovering a bucket near either edge of a service’s reliability strip or segment lane showed only part of its readout, because the card clips its contents. The readout is no longer clipped by the card and stays inside the window whenever it fits there, moving in from the edge, or above or over the strip when it has to.
- Response-time readouts are steady. A recorded check on the monitor page can be hovered from a few pixels away instead of only exactly on its dot (the nearest check wins), and the card no longer grows and shrinks as the readout appears and disappears, however long the readout is.
- The escalation-policy row on a service’s paging card lines up with the rest of the card instead of running to its edge in larger type.
Upgrade notes
Section titled “Upgrade notes”- One schema migration (00110): an index on resolved incidents, applied automatically on startup (or by
cerbix migrate). It is built withoutCONCURRENTLY, so writes toincidentswait while it builds — a fraction of a second at typical sizes; reads are unaffected and no role needs stopping. The index changes no data, so rolling back to v0.3.7’s binary is safe. No configuration change.
[v0.3.7] - 2026-10-07
Section titled “[v0.3.7] - 2026-10-07”🩹 Fixed
Section titled “🩹 Fixed”- Service reliability no longer stalls on services with a long history. The sealed-through watermark was recomputed over every bucket since the service’s era began, on every scheduler slice; once that took longer than the slice’s commit reserve (about 50 000 buckets — roughly five weeks of a service’s history), every slice rolled back and materialization stopped for good, for every service, since the stuck one was always picked first. The recompute now starts at the current watermark and looks at most one day ahead. A stalled installation resumes on its own after upgrading (D-0269).
Upgrade notes
Section titled “Upgrade notes”- No schema migration, no configuration change.
[v0.3.6] - 2026-10-06
Section titled “[v0.3.6] - 2026-10-06”🔒 Security
Section titled “🔒 Security”- Frontend dependencies patched.
vue3.5.43 fixes GHSA-g2v6-rqmx-r4w6 in@vue/server-renderer(XSS through an attribute name containing a carriage return) andsource-map-js1.2.2 fixes GHSA-68fv-2mgg-jv7q (event-loop denial of service); both were flagged as high bynpm auditon v0.3.5’s dependencies. The embedded SPA is rebuilt with them. google.golang.org/grpc1.83.2. Clears GO-2026-6443 (in an imported package; govulncheck finds no call path from cerbix).
🩹 Fixed
Section titled “🩹 Fixed”- The Security workflow scans again.
govulncheckis pinned to v1.7.0, the newest release that runs on the repository’s Go 1.25 toolchain;@latest(v1.8.0) needs Go 1.26 and had stopped installing, so nothing was being scanned. Threegolang.org/x/cryptoadvisories (GO-2026-6355, GO-2026-6354, GO-2026-5932) remain until the Go 1.26 upgrade; govulncheck reports none of them as called.
Upgrade notes
Section titled “Upgrade notes”- No schema migration, no configuration change.
[v0.3.5] - 2026-10-06
Section titled “[v0.3.5] - 2026-10-06”✨ Added
Section titled “✨ Added”-
Structured Cobra CLI. The command tree now provides grouped top-level help and detailed command-specific help. Documented commands, canonical flags, defaults, environment variables and gate/change result exit codes remain unchanged. Leaf commands reject previously ignored positional arguments and unsupported shorthand clusters with usage exit code 2 before side effects; cluster-looking string-flag values remain literal. A version JSON writer failure now reports
version:and exits 1 instead of the native parser’s silent exit 0; successful JSON still exits 0. Viper is not used. -
Cross-brand identity and shell UX. cerbix now uses the Sealed C default mark with custom-logo priority and contrast-safe instance accents. The SPA adds responsive navigation access, accessible overlay and focus behavior, atomic organization/project transitions, a stale-response-safe SearchBox, truthful Dashboard loading/no-data/error states, and explicit breadcrumb, theme and announcement semantics. Desktop layout, status colors, reliability formulas and backend/API contracts remain unchanged.
Changed
Section titled “Changed”- Service burn rules need a short window of at least 300 s. A service burn window ends at the sealed watermark, which trails now by the 120 s late-arrival grace plus up to a bucket, so a shorter short window could never — or only intermittently — contain sealed time. The API now refuses it with 400; monitor burn rules keep their 60 s floor. Rules stored earlier are not migrated and keep their behaviour until they are next saved (D-0267).
🔒 Security
Section titled “🔒 Security”- Remote CLI URL errors redact credentials.
gate checkandchange recordno longer echo a rejectedCERBIX_URLvalue or parser detail that could contain embedded userinfo or query credentials.
🩹 Fixed
Section titled “🩹 Fixed”-
Cobra parser boundaries fail closed. Commands must occupy their canonical raw positions; undocumented shorthand clusters (including
-hh=falseand-hhh=false) now fail with usage exit 2 before an executor or config/network access, while registered string-flag values remain literal. Version positionals, including"", take priority over unsupported clusters and malformed help before JSON encoding. Help assignments accept only exact lowercasetrue/false, not pflag aliases such as1ort; string-flag values remain literal. Ordinary positional and compound-help path errors precede any help policy; only then does any canonical false-help assignment win for ordinary commands regardless of repeated flag order;versionkeeps false-help as a compatibility no-op while any parseable true-help wins; compound help is idempotent; hidden completion protocols are unreachable; and the Windows mousetrap is disabled. -
Architecture and operator documentation now follows the current runtime. HTTP-pull diagrams use the real
/api/v1/agent/*routes; heartbeat storage is described as adaptive and the PostgreSQL 15+ floor is separated from the repository’s PostgreSQL 16 images; README build and deploy/change boundaries match the supported workflow; and the top-level configuration table includes the current service, audit, and expected-run ledger sections. -
Wave 2 documentation closes the remaining domain and operator gaps. Architecture/overview/README now cover the current-domain partial ERDs, operator recovery CLI, build/toolchain contexts, and shipped shell/brand behavior; docs-check adds guards against stale routes, queues, storage and heartbeat-ERD claims.
-
Instance branding and detail reads are fenced. Custom accents choose readable
--accent-inkvalues and clear all inline accent properties when removed. Incident Detail rejects stale-generation and foreign-project responses before publishing incident data or actions. -
Service reliability materialization no longer globally sorts retained heartbeat history for carry-in. The historical sample-and-hold lookup now performs one bounded, index-backed latest prior observation lookup per declared monitor, preserving the exact carry-in semantics while preventing a large retained history from consuming the Service materialization slice. This is a data-compatible change and requires no database migration.
-
SLA windows longer than raw heartbeat retention cover the whole window. With the default 30-day
heartbeats.retention_days, the 90-day SLA (and any window longer than a lowered retention) silently covered only the retained days. Availability now adds the daily rollup for days whose raw heartbeats were purged — on monitor and project SLA, the weekly SLA report, the 90-day uptime of a monitor-backed status-page component, and monitor burn windows. A window with no data older than its start still shows its number and says where the data starts: the UI marks it “since DD.MM.YYYY UTC”, and the API addsdata_from,latency_from(latency stays raw) and the public status-page fielduptime_since(D-0268). -
SLA “budget left” reads the budget, not all time. An untouched 99.9 % budget showed 0 % and At risk because the view rendered
remaining_ratio(a fraction of all time); it now shows 100 % and Meeting (D-0267). -
Late heartbeats correct every sealed bucket they reach. A heartbeat arriving behind the service watermark now repairs every sealed bucket up to the monitor’s next observation, not only its own minute, and the repair runner makes progress on long ranges instead of retrying one oversized batch; a statement timeout at the tail of a slice no longer counts as a failure (D-0267).
-
A re-enabled push monitor is not instantly BAD in service reliability. A ping from before the re-enable is no longer evidence, matching the monitor’s own dead-man timer (D-0267).
🧪 Tests
Section titled “🧪 Tests”- Gate fixtures now match schema 109. Test rows and assertions were aligned with the persisted
policy source/owner fields, all-window fields, and the schema-109 payload limits. PostgreSQL 16
declarative-partition and TimescaleDB normal/race
internal/storesuites pass.
Upgrade notes
Section titled “Upgrade notes”- This release contains no schema migration and does not require an offline database upgrade. Existing installations retain their current data and watermarks.
- Saving a service burn rule whose short window is below 300 s now returns 400; widen it to 300 s or more when you next edit that target.
- On plain PostgreSQL (no TimescaleDB), long SLA windows of large projects read noticeably slower than before (measured locally: a 500-monitor project’s four windows about 2.1–2.5 s instead of about 1.1 s); TimescaleDB is unaffected in the same measurement (D-0268).
[v0.3.0] - 2026-09-22
Section titled “[v0.3.0] - 2026-09-22”This minor release combines the status-page service-component repair from iter-0181 with the
alert-routing tenant-boundary hardening from iter-0182, bounded audit retention from iter-0184,
project-level inherited release-gate policy from iter-0185, and worst-of-all-windows gate evaluation
from iter-0186, the evidence-driven onboarding journey from iter-0187, and the service-first public
status-page incident experience from iter-0189, hardened in iter-0190, plus the incident-aware
overall-status composition from iter-0191, hardened in iter-0192.
✨ Added
Section titled “✨ Added”-
Bounded audit-log retention (iter-0184, FR-033 / NFR-027). Strict instance-wide retention removes only database-clock-expired organization and global audit rows in fenced, bounded batches. The scheduler, metrics, alerts, recovery guidance, and migration index ship together; audit APIs and retained-history semantics do not change.
-
Project-level inherited release-gate policy (iter-0185, FR-034 / NFR-028). Projects can define one versioned gate policy. A service resolves its effective document in explicit service → project → not-configured order, records the immutable source tuple in decisions and overrides, and exposes the result through API, CLI-compatible schema, metrics, and the Settings / service SPA. Editing a project policy revokes only overrides that inherited that project revision.
-
Worst-of-all-windows release-gate evaluation (iter-0186, FR-035 / NFR-029). A schema-v2 project or service policy can evaluate every configured standard SLO target in one snapshot. The decision keeps healthy and unhealthy windows together, applies BLOCK → unavailable → WARN precedence without averaging or early exit, persists complete per-window evidence, and renders mode switching, inventory, inherited policy, UNKNOWN and override states in the SPA.
-
Evidence-driven onboarding (iter-0187, FR-036 / NFR-030). The Dashboard now offers a non-blocking organization → project → monitor → first persisted heartbeat guide. Progress is recomputed from tenant-scoped server facts, monitor creation stays in the existing typed form, push credentials remain on monitor detail, and UP or DOWN is accepted as the first useful result without mislabeling a failed target as healthy. Existing installations keep their normal KPI, availability and monitor-card Dashboard; manual re-entry opens only a compact setup summary. Automatic-dismissal preference is browser-local and scoped by user and selected tenant context.
-
Service-first public status pages (iter-0189, FR-037 / NFR-031). Current service/component state now precedes incident detail. Active incidents render as compact, keyboard-operable full-row accordions with deterministic impact grouping only above eight rows, one lifecycle badge, one impact badge, truthful opened/updated metadata, a two-line latest-update preview, and the real timeline on expansion. Page-local affected-component relations support direct incident focus without exposing internal monitor, service, project, actor, update, or postmortem identifiers.
🩹 Fixed
Section titled “🩹 Fixed”-
Active status-page incidents can no longer coexist with a false green all-clear (iter-0191, hardened in iter-0192; FR-038 / NFR-032). Public and authenticated-preview heroes compose measured component health with the existing public active-incident list in one pure frontend helper. Explicit
summary_stateowns the component truth and all-clear eligibility; the legacy summary fallback runs only when state is absent. Measured impairment or maintenance keeps its headline, while otherwise active incidents lead with count and order-independent worst impact. Major and Minor render warning attention, Critical renders danger attention, and impactnoneremains neutral rather than green. Complete supporting copy retains no-data, empty-page and unmeasured-component disclosure. Component rows, incident lifecycle, API/OpenAPI, persistence, cache, metrics, alerts, SLA/SLO and timestamp semantics remain unchanged. -
Status-page post-close hardening is isolated in iter-0190. The scheduled-maintenance title is restored to the logical heading hierarchy without changing its visual style. Backend incident projection now reuses one ordered page-local monitor/service index, and the SPA reuses one computed component-to-incident map for counts and navigation instead of repeatedly scanning both lists. The closed iter-0189 report remains an immutable delivery snapshot.
-
The scheduler readiness race gate is deterministic again. Its live-leader test no longer cancels the scheduler before assertions or mistakes the startup fail-closed
ready=0state for the intended evaluator-lag verdict. Cancellation and goroutine joining now happen in test cleanup, and the wait requires the specificlaggingreason. Production scheduler behaviour is unchanged. -
Status-page components can be bound to services again. The API contract and SPA already sent
service_id, and the store already persisted it, but the HTTP request decoder did not accept the field. Because unknown JSON fields are rejected, every valid service-backed component request failed as400 invalid JSON body. The handler now acceptsservice_id, validates both binding ID formats at the transport boundary, and returns the created component with its derivedservicesource. -
Binding failures now keep their correct HTTP meaning. A missing or unauthorized monitor or service returns the same non-oracular
400 binding not found; a status page deleted during the create transaction remains a real404; and database or infrastructure failures are logged and returned as500rather than being mislabeled as caller errors. -
The live status-page regression cleans up after partial failures. Its service fixture is now registered for cleanup immediately, so a page-create or response-decoding failure does not leave test data behind.
🔒 Security
Section titled “🔒 Security”-
Status-page binding ownership is now enforced transactionally by the store. Monitor and service bindings must belong to the page’s organization and, for project-scoped pages, to the page’s project. A retained monitor/service pair must resolve to one compatible source project. HTTP and non-HTTP callers therefore use the same validation owner, without a separate monitor-only preflight that could drift or expose cross-tenant existence.
-
Alert-routing references are now same-project persistence invariants. Monitor/channel links, escalation-policy targets, on-call schedule participants and updates, override channels, and frozen service-escalation snapshots are validated at store and schema boundaries. Updates match both object ID and project ID, and runtime escalation resolution carries the incident project through policy, schedule, and channel reads so a pre-existing malformed reference cannot deliver across projects.
-
Missing and foreign routing IDs share one generic client refusal. The API maps the store-owned refusal to
400without revealing whether an identifier exists in another tenant. No target ID is added to metrics or logs as a label.
⚙️ Upgrade notes
Section titled “⚙️ Upgrade notes”-
Migration
00106_alert_routing_tenancy.sqlbackfills project identity for relational routing edges, adds composite tenant foreign keys, and installs JSONB tenant guards for policies, schedules, and frozen snapshots. It deletes no heartbeat, incident, audit, or historical routing data and performs no silent repair. -
The migration intentionally fails fast if deployed data already contains a cross-project routing reference. Diagnose the offending row read-only, repair it explicitly, and rerun the migration; the runbook contains the operator procedure. Valid existing routing data requires no manual action.
-
No configuration key or public API schema is removed. The status-page change makes the server honor the already-published
service_idcontract; valid same-project alert-routing behavior is unchanged.
iter-0181 / D-0249 and iter-0182 / D-0250 · full Go tests, targeted race, vet, build,
documentation checks and configured lint green · live status-page Playwright regression included in
71 passed / 1 skipped / 0 failed · PostgreSQL 16 migration and direct-SQL tenant regressions green,
including the full DB-backed internal/store package. iter-0184 completed its delivery gates;
iter-0185 and iter-0186 pass their PostgreSQL migration/snapshot regressions, full Go and SPA
suites, generated-schema parity, Dockerized lint, build, docs checks, and the scoped live gate UI
Playwright suite. iter-0191 / iter-0192 pass 56 focused helper/component tests, 764 frontend tests,
full Go/race/build/vet/lint, generated API and embedded-SPA parity, and rebuilt-stack public/preview
Playwright at desktop and 430 px with exact copy, accessible heading and no horizontal overflow.
[v0.2.0] - 2026-09-08
Section titled “[v0.2.0] - 2026-09-08”One defect, found by an independent reviewer about an hour after v0.1.9 was published, and
repaired under its own iteration rather than left for the next one.
Why a minor bump for one fix, and why this section was numbered twice. It was released as
v0.1.9.1 first, and that tag and release were withdrawn by the owner: a four-component version is
not semantic versioning, so tooling that orders by semver — GitHub’s own “latest” among it — would
never have ranked it. The content is unchanged; only the number is. v0.2.0 rather than v0.1.10
is the owner’s choice, and the fix it carries is a behaviour restoration rather than an addition.
🩹 Fixed
Section titled “🩹 Fixed”-
An
async_canarythat declared a credential binding could not run at all. Every dispatch was refused at the executor gate and the monitor reported nothing else.The scheduler nominated monitors for credential materialization by asking whether the monitor’s TYPE carries a static credential schema. An
async_canaryhas none — its bindings are declared inside the workflow document — so a canary with a declared binding took the plain branch and was stamped with a generation that carries no envelope, while the expected-field set named the binding’s field anyway. The executor then refused the job, correctly, for a field the producer had never been asked to supply. If you declared a binding on a canary inv0.1.9, it never ran; after this release it runs, with no change to the monitor.The predicate the nomination asks is now a monitor-level one rather than a type-level one. The schema half of it is unchanged on purpose: the tempting definition — “expects at least one envelope field” — reads false for
promql, which has a schema and no required field, and would have repaired one type by breaking another. -
And Test Connection on such a canary now says what to do instead of failing far away. The same substitution was made on the test path: it builds a generation-1 job with no envelope for any type whose credential requirement is decided by its schema, and a canary’s is not. The pre-save refusal that a SYNTHETIC scenario binding has had since FR-028 — “save the monitor before testing it” — had no canary twin, so pressing Test on a canary with a binding produced a message about a missing envelope field from the far side of the dispatch. It now names the binding and the way forward, before any probe is attempted.
-
And the envelope could not be BUILT for such a canary either, one level below the nomination. The digest that binds a credential to the exact execution asks for the canonical value of every binding key; for a canary those keys are the workflow document, the run key and each reference, and the value lookup had arms for a synthetic scenario and none for a canary — so it fell through to the credential-schema resolver, which refuses a type that has no schema. Sealing failed and the monitor was refused with a reason naming decryption, which points an operator at a key problem that does not exist. This was found by writing the store/materializer/gate regression an independent reviewer refused to release without; the nomination repair alone would have shipped the feature still broken.
-
No migration, no configuration change, no API change. A canary that declares no binding, and every other monitor type, dispatches exactly as it did. The new refusal is a 400 on a path that previously reached a prober and failed there.
Why no test caught it
Section titled “Why no test caught it”Every test passed on the defective tree, including the full -race suite, the browser suite and the
geo topology suite. None of them dispatches a canary that declares a binding: the canary suites use
bindingless workflows and the credential suites use schema types. It is stated here rather than
quietly fixed because the gap is the useful part — a green gate is only evidence about what it
actually covers.
iter-0180, D-0248 · full -race 33 packages exit 0 (internal/store 763.4 s) · make dev-test
70 passed / 1 skipped on an image rebuilt from this candidate tree · make docs-check OK · THREE defects,
one substitution: a type-level predicate answering a monitor-level question, in the nomination, in the pre-save test refusal and in the canonical value lookup the digest depends on · the third was found only by TestACanaryBindingIsSealedIntoAGenerationThreeEnvelope, which runs against the real materializer and the real executor gate and proves generation 3 + EnvelopeV2 + the credential actually used — a store/materializer/gate seam, not an end-to-end one: it neither publishes nor probes · four regressions, each with its killing mutation — the nomination loop, and
the pre-save refusal whose canary arm was missing · the sibling was found by reading every product
call site of the type-level predicate against what it decides: eighteen sites, fifteen correct, two
already unions, one carrying the same defect · the nomination regression with its killing
mutation — restoring the old predicate fails the declared-binding case by name and leaves both
negative cases passing · what is NOT covered is recorded in iter-0180 §4c and not softened: no single process holds the
scheduler, the store and the prober at once, and the live cases that exist today do not demonstrate a successful canary run
[v0.1.9] - 2026-09-07
Section titled “[v0.1.9] - 2026-09-07”Everything since v0.1.8, in four parts and one set of fixes: the typed external canary and
the audit trail that were prepared under this number once before, the truthful-rendering
package, the expected-run ledger, and the fifty-two repairs an independent two-axis review
of this whole range then found in it.
This heading was [v0.1.9] - 2026-09-03 before, and it was withdrawn. The tag had been created
and then deleted by hand before the review that gates it closed (8ee023c docs(v0.1.9): the tag is dropped, not moved, until the review closes), while the dated heading outlived it and announced a
release that did not exist. It is back by the owner’s decision, with the later work folded in and
the date moved to when that decision was taken.
Why that history is kept here: a version heading can make a claim falsely, and this one did.
The withdrawn heading announced a release whose tag had been deleted by hand, and the file went on
saying so for days. The record of that stays; the sentence that reported the repository’s tag state
does not, because THIS SECTION IS THE RELEASE BODY — .github/workflows/build.yml publishes it
verbatim for the matching tag — and a release announcing that its own tag does not exist would be
the same defect wearing the opposite sign.
Part 1 — the canary work, prepared under this number once before
Section titled “Part 1 — the canary work, prepared under this number once before”Four things in one release, in the order the owner set: a typed external canary for async API journeys, an audit trail for writes that change something — incidents and now monitors — the dependency sweep that had been waiting behind both, and one small improvement that came last, an optional description on a monitor. The canary is the largest single feature since FR-021 and it ships complete: all six phases, including the capability announcement and the typed UI form that an earlier draft of these notes listed as missing. No breaking change. Five migrations, all additive.
⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- The order of a canary rollout no longer matters. A canary reaches only an executor that
ANNOUNCED it can run one, so declaring a canary before a region is upgraded is safe: the run is
refused with one bounded DOWN per attempt —
no_capable_runnerwhen nothing there announced a canary runner,capability_mismatchwhen what is there speaks another version — and it starts working by itself when the executors arrive. An earlier draft of these notes told you to upgrade the region first; that instruction is retired, not merely relaxed. - An AMQP
workerolder than this release never receives a canary at all. Canaries ride their own queue (checks.canary.<kind>@<version>.<region>), which only a current binary binds, and on the HTTP-pull transport a job carries the capability it requires while a claim carries what the agent declared. That is the point: in a mixed fleet the old executor cannot take a job it would fail. - A repeated incident acknowledgement is now a genuine no-op. It used to rewrite the incident’s
updated_aton every retry. The acknowledger and the instant stay the FIRST ones, and the modification time no longer moves. If anything of yours sorts incidents byupdated_atand relied on a re-acknowledgement bumping them, it will not any more. - Two identical Alertmanager deliveries that arrive at the same moment now both answer HTTP 200 —
one reporting
opened: 1(orresolved: 1) and the otherignored: 1. The loser of that race used to answer 500, while the sequential retry a millisecond later answered 200. If you alert on 5xx from the receiver, that source of noise is gone. - The
golangbuild image moves to 1.27.0 anddocker/setup-buildx-actionto 4.3.0. Neither changes the language version ingo.mod.
✨ Added
Section titled “✨ Added”A typed external canary for async API journeys (FR-029 / NFR-024)
A new monitor type, async_canary, runs ONE asynchronous transaction end to end and reports it as an
ordinary monitor: submit, take a correlation id, await a terminal outcome, assert the declared fields,
validate the cleanup boundary — and return ONE heartbeat carrying up/down, total latency, the failed
stage and a bounded code class. It is declared in a Monitoring-as-Code bundle as a nested, typed
workflow: block with closed unions and no free-form field anywhere: no settings map, no JSON string,
and a key the schema does not name refuses the whole bundle by name.
- The contract is a type, not a convention. Two submit kinds (
http_json,multipart_fixture), two completion kinds (sse,poll_json), two correlation sources, three assertion kinds, two restricted grammars, and every bound with a number.{{ correlation_id }}is legal in exactly one field — the completion URL — because nothing has produced an id at submit time. - Credentials are bindings, never values. A binding is declared once under
workflow.secretsand referenced by name at each position; the value is resolved at dispatch, delivered in the existing credential envelope, held in memory for one execution and wiped. A credential-bearing header accepts a binding and nothing else. What is not protected, stated plainly: a credential pasted into an ordinary header or under an innocuous body key is not detectable and is not refused — the same residual FR-028 named. - The URL policy is strict and has no off switch. HTTPS only; loopback, link-local, private ranges
and cloud metadata are refused AFTER DNS resolution and re-validated on every redirect hop. There is
no setting that relaxes it, in v1 or as a hidden flag, because a flag reachable in production is the
policy’s own bypass — tests reach a local fixture through an injected dialer at the seam, never
through a product option. The executor also drops EVERY binding-backed header on a cross-host
redirect, which
net/httpdoes not do: it stripsAuthorizationand has never heard ofx-api-key. - One run per monitor, four per region, decided by the scheduler at dispatch. A refused run writes
one ordinary DOWN with a bounded reason (
region_saturated,already_in_flight,no_capable_runner,capability_mismatch) and the monitor’s ownfailure_thresholddecides the flip, so the sample counts as unavailable in the SLI like any other DOWN. Every submit carries anIdempotency-Keyderived from the monitor, its execution revision and the scheduled window — whether a second submit with that key creates a second task is the target’s contract, not cerbix’s. - Nothing it touches leaks. No heartbeat, log line, error or metric label carries a URL, a body, a header value, a secret, the correlation id or the result object path.
- New metrics
cerbix_canary_stage_totalandcerbix_canary_dispatch_refused_total, both low-cardinality and carrying no monitor id, URL or correlation id.
An executor receives a canary only if it announced it can run one (the last of the six phases).
The announcement is a set of <kind>@<version> tokens, and an executor makes it from the runner it
actually has: a pull agent in its heartbeat and again on every claim, an AMQP worker by consuming
checks.canary.[v3.]<token>.<region>, and --role=all by construction. The scheduler dispatches only
into a region that announced the token the DOCUMENT needs, and a region that announced nothing gets
no_capable_runner while a version skew gets capability_mismatch — two reasons because the fix
differs: start a runner, or finish the upgrade. The filter is not treated as the whole barrier, for the
reason the credential envelopes already established: a capability CHECK does not stop a consumer from
consuming, so an incapable executor is made unable to RECEIVE the job — a queue it does not bind on
AMQP, and a per-claim capability filter on the pull transport.
A canary is declared in the SPA, in a bundle, or through the API. The typed form is typed only: no JSON editor on the create view or the read view, the five stages laid out as stages, bindings as stage 0, and every refusal met AT THE FIELD rather than as a 400 after Create. A saved canary reads back into the same form, with the binding halves recombined from the document’s marker and the flat reference key’s name.
A monitor says what it is for. A monitor carries an optional description — plain text, at most
200 characters counted as Unicode code points, so a Cyrillic sentence is as long as a Latin one to the
person writing it. It appears under the name on the monitor list (one line, the whole text as its
tooltip), as the beginning of a line on a dashboard panel, and in full on the monitor’s own page; it is
editable on the create and edit forms with a live count, and declarable in a Monitoring-as-Code bundle.
Every monitor that exists reads back with an empty description and every surface renders exactly as it
did — the absence of an element is asserted per surface, not assumed. It is deliberately absent from
public status pages, notification payloads and search.
🔒 Security
Section titled “🔒 Security”An incident write says who wrote it (FR-026 / NFR-021)
FR-022 promised, as its invariant 14, that “every write is audited with actor and tenant, in the
mutating transaction”. That was false of the PRODUCT rather than merely unimplemented: an incident
write left no audit_logs row, for a monitor incident or a service one. A member could resolve
someone else’s incident, publish a postmortem or acknowledge a page, and the organization’s audit trail
said nothing.
- Every incident mutation made by a PRINCIPAL now writes exactly one audit row in the same transaction — the manual create, the Alertmanager receiver’s create and its resolve, a status change, a note, the first acknowledgement, and a postmortem create or update. A committed change without its row cannot exist, and a rolled-back one leaves none.
- The vocabulary is five words —
incident.create,incident.status,incident.note,incident.acknowledge,incident.postmortem— and the target carries ids and both ends of a transition and never a body: no note text, postmortem text or alert annotation reaches the trail. - Machine writes are excluded by decision and enumerated: the reconciler’s auto-open and
auto-resolve, the service auto-incident and its resolve, the
⚡ Context:note, both⏸ Suppressed:writers,🚀 Changes:and🕸 Impact:. Their record is the incident’s own timeline, which namessystemas the author. Auditing them would bury a tenant’s log under a flapping service’s heartbeat. - No read is added. The rows appear in the organization’s existing audit listing and never in the
instance one.
docs/runbook.mdnow answers “who resolved this incident” directly. - Two AST guards hold the door surface, each driven by a fixture that contains the violation. The first version of one guard exempted the Alertmanager receiver by name — and the exemption hid the exact defect the requirement exists to prevent, because Alertmanager posts with a project-write token and a token is a principal.
And a monitor write says who wrote it, by the same rules
Creating, editing or deleting a monitor through the API or the SPA left no audit row either. It now
writes exactly one, in the mutating transaction, with three words — monitor.create,
monitor.update, monitor.delete.
- The target names the document and never its contents: the monitor’s id, slug, type and region;
enabled true→falsewhen the write was a pause or a resume, because that is the edit an operator asks about; and, for an async canary whose workflow declarescleanup.kind: nonewithacknowledged: true, the clausecleanup=none acknowledged— which is what makes that acknowledgement visible in the trail, as its requirement promised. No config value, credential reference, target URL, scenario or workflow body appears in a target, for any type. - What the row says is read from the row the writer HOLDS, not from what the caller passed: the
FOR UPDATEstatement for a delete, the returned row for an update. Independent review found the first version taking a whole monitor from the caller — a concurrent edit could make the audited region stale, and a careless caller could file the row under another project’s organization. The delete door takes an id now, so there is nowhere to put a foreign project. - A monitor applied from a Monitoring-as-Code bundle writes no row: it is a machine write, and its record is the bundle. Deleting a project leaves the organization’s trail intact.
🔧 Changed
Section titled “🔧 Changed”- A pull monitor with a timeout past 30 seconds is no longer re-claimable mid-probe. The agent’s claim lease was hardcoded at 30 s, so any pull monitor slower than that could be handed to a second agent while the first was still working. The lease is now per job. This is older than the canary; the canary is what made it visible.
async_canaryjoins themonitors_type_checkdatabase constraint. Every Go test passed while the real table refused the type, because no store test writes to it — a live E2E found it, and the type vocabulary is now asserted against the database itself so the next new type cannot repeat it.- A red CI job now names the tests that failed, as annotations the check-runs API serves without repository-admin rights. Two failures that had been unreadable from outside turned out to be tests measuring the MACHINE rather than the product: one read a legitimate deferral of a partition drop — correct behaviour while any older transaction is open — as a crash that never happened, and one gave a repair slice two seconds and failed when the closing write missed it, where production keeps the range under a sixty-second lease and resumes. Both tests are corrected; no product code changed.
📦 Dependencies
Section titled “📦 Dependencies”Seven of the eight open dependabot branches, verified one group and one major at a time (iter-0170).
- go-modules:
rabbitmq/amqp091-go1.13.0 → 1.14.0,google.golang.org/grpc1.83.0 → 1.83.1 - github-actions:
docker/setup-buildx-action4.2.0 → 4.3.0 (SHA-pinned) - docker: the
golangbuild image 1.26.6-bookworm → 1.27.0-bookworm - frontend:
@types/node→ 26.4.1,eslint→ 10.9.1,vue-tsc→ 3.3.11, and three MAJORS —jsdom26 → 30,vite6.4 → 8.2,vue-router4.6 → 5.3. The last two are one decision, not two: vue-router 5 declarespeerOptional vite@"^7.3.0 || ^8.0.0"and will not install on Vite 6. - Not taken: TypeScript 7.
vue-tsc3.3.11 cannot drive the native compiler port —npm run type-checkdies withERR_PACKAGE_PATH_NOT_EXPORTEDbefore checking a file. Upstream of this repository; the bump waits for avue-tscthat supports it.
The committed SPA snapshot (internal/web/dist) is regenerated with this: every asset filename hash
moved, because Vite 8 hashes differently.
API · Metrics · Schema
Section titled “API · Metrics · Schema”- API:
Monitor.descriptionon read, create and update (maxLength: 200, counted as code points; omitted on update leaves it unchanged,""clears it, and a longer value is 400 naming the field). The monitorconfigdescription documents theasync_canaryworkflowdocument and itscanary_secret_<binding>_refkeys; the server canonicalizes the document on write, so an API-created canary and a bundle-declared one store byte-identical documents. The agent protocol gains one capability declaration in both directions:workflow_kindsin the heartbeat andX-Cerbix-Workflow-Kindson every job claim, bounded and grammar-checked before it reaches a column. - Metrics:
cerbix_canary_stage_totalandcerbix_canary_dispatch_refused_total, both low-cardinality and carrying no monitor id, URL or correlation id. The refusal metric’s reasons are the four bounded ones a heartbeat can carry. - Schema: five additive migrations —
00095the canary in-flight lease,00096the per-job pull claim lease,00097async_canaryinmonitors_type_check,00098pull_jobs.workflow_kind(NULL for every job any agent may run),00099monitors.description(NOT NULL DEFAULT '', which is the whole compatibility promise in one default). No index added: nothing reads by a description, and search over it was declined.
Decisions D-0218 … D-0234 · FR-029, NFR-024, FR-026 (+ its §10 monitor amendment), NFR-021 and
FR-030 DONE · iterations iter-0168 … iter-0173, all CLOSED, every range approved by the independent
reviewer · full -race suite green (33 packages), vitest 46 files / 514 tests, browser suite 66 passed
/ 1 skipped, geo topology suite 12 passed · hosted CI green on the pushed head, all seven jobs,
including Backend (timescaledb hypertables) — the job whose two failures D-0225 could not read until
the annotations step made a red run legible
Part 2 — a rendering claims no more than the facts behind it (FR-031 / NFR-025)
Section titled “Part 2 — a rendering claims no more than the facts behind it (FR-031 / NFR-025)”Three surfaces drew two different states identically, so a reader could not tell them apart — and one of them quoted availability over a range it had never stored.
- The reliability timeline is on CLOCK TIME. Cells come from the requested range at the rollup
grain, so a step with nothing stored occupies its own real width in its own encoding instead of
vanishing and letting its neighbours stretch over it. Five encodings that cannot be confused:
unknownis a solid neutral slice (a decided verdict, never a status hue),not-storedis a hatch with an outline, and opacity meansprovisionaland nothing else. Height carries QUANTITY. A problem is never hidden: a slice too small to draw is MARKED non-geometrically and named in the readout rather than being silently rounded away. - A segment states its STORAGE verdict and withholds availability while that storage is
incomplete. This closed a real defect in the old numbers: a segment could quote
availability 100%over a range it never materialized. Coverage is still printed, as its own separately named fraction. - The Response time panel draws every heartbeat at its real timestamp — including the zero-latency failures that used to be filtered out and disappear — as points only, with no connecting stroke and no fill, because neither can span time nothing was measured in. Absence is drawn POSITIVELY by an observation ruler: one tick per recorded check, and the empty spans between them are focusable and say only what is known.
- The
/slaobjective card says which state it is in. Read-only with an explicit Edit, a guarded save, the draft cleared on success, and both writers gated on their load generation — so a save that started under one project can no longer land under another. - NFR-025 — no timestamp is rendered without naming its zone, and it is enforced rather than
fixed. One mechanism, a NAMED RENDERER PER SUBJECT and deliberately no generic date formatter,
with the UTC offset resolved AT the instant (a 30-day window in late March or October crosses a
DST change, so this is the ordinary case). Identity stays UTC; presentation is local and always
names its offset. Every hand-rolled date site in the SPA is gone: outside two modules — one
for renderings, one for keys, wire values and HTML control values — no product file calls
toISOStringor pulls a field off aDateat all. The suite runs green atTZ=UTCand atTZ=Asia/Yekaterinburg, which is what makes that a fact rather than a claim.
Decisions D-0235 … D-0236, D-0246 · FR-031, NFR-025a/b/c DONE · iterations iter-0174 and the NFR-025c half in iter-0178 · no Go code and no API change in the FR-031 half: all four surfaces are the SPA and the storage verdict reads a field the payload already carried
Part 3 — cerbix records that a run was expected (FR-032)
Section titled “Part 3 — cerbix records that a run was expected (FR-032)”Until now a check that never ran left nothing behind. A gap in the heartbeats meant “we do not know”, and the product could not tell “the monitor was fine and nothing was due” from “a run was due and never happened” — so no rendering could honestly connect two points across it.
- An expected-run ledger. The scheduler now persists the window it already computed:
expected_runs, range-partitioned bydue_at, one row per(monitor, due window), with amonitor_schedulesegment history behind it so a monitor whose interval changed is read against the interval that was in force. Every window ends in a verdict — covered, covered late, missed, withheld, reserved — and the ledger withholds rather than guesses whenever it cannot defend an answer. - RESERVE → PUBLISH → CONFIRM. The window is recorded BEFORE the job leaves the process, and
issued_atis written only by the CONFIRM that follows a successful publish. A running instance produced a falseexpected_never_issuedunder the older advance-then-publish ordering; this is the fix, and only the windows whose reservation was proved durable are published. - A fourth carrier generation that announces itself.
checks.jobs.v4.<region>and its HTTP-pull twin carry job identity; a region gets ledger-eligible jobs only where an executor has ANNOUNCED it can consume them, so a half-upgraded fleet degrades to “not eligible” instead of losing runs. Rolling back the migration REFUSES while generation-4 rows are pending rather than discarding them. - The panel may draw a connecting stroke again — but only across a span the ledger says was covered. That is the whole point of the requirement: the stroke returns as a claim the product can defend.
- Reading it:
GET …/expected-runswith a versioned cursor, aledger_fromfence so nothing before the ledger existed is presented as a missed run, retention that includes the DEFAULT partition, and an unlabelled HOT-ratio gauge sampled by the leader.
Decisions D-0237 … D-0245 · FR-032 DONE · iterations iter-0174 … iter-0177 · schema 00100
monitor_execution_revisions, 00101 the generation-4 pull boundary, 00102 the ledger itself,
00103 the reservation columns — all additive
Fixed — two things only a live distributed stack could show
Section titled “Fixed — two things only a live distributed stack could show”-
A credentialed Test Connection failed in the distributed topology with
no worker queue for region "core". The API enforced credential envelopes and the core worker’s config did not, so the API published on a carrier that worker had never bound. The two halves of one topology now agree, and a regression fails if they ever disagree again — in either direction, and also if the executor is handed the at-rest master key it must never hold. -
Every generation-3 Test Connection was dead-lettered by the consumer bound to its own queue. One function serves both envelope test carriers and its admission check compared against a hardcoded generation, so the newer carrier accepted nothing. Because the caller is answered only on success, this surfaced as a ten-second timeout reported as
no worker responded in region …— the same message an empty region gives. The queue is the generation now, as it already was on the jobs path, and a table-driven regression publishes on every carrier an executor serves. -
An envelope OLDER than its carrier defines was accepted on every path. The mapping is explicit — carrier generation 2 carries envelope v1, generation 3 carries v2 — and the producer had implemented it exactly since generation 3 shipped, while every consumer enforced only a floor. So a generation-1 envelope on a generation-3 carrier opened normally and its credential reached the prober WITHOUT the execution-body binding that generation exists to add: the property that stops a credential being replayed against a different target. One owner holds the mapping now, the gate every executor crosses refuses a mismatch, and both AMQP consumers dead-letter it.
-
A refused dispatch is answered instead of being met with silence. It used to publish nothing, so the caller waited out its RPC timeout and reported
no worker responded in region …— the same sentence an EMPTY region gives, which made a refused delivery indistinguishable from a dead region. It now returns a typed reason immediately, from the same bounded vocabulary executors already answer with, and the poison body is still dead-lettered for inspection. If you alert on Test Connection latency, refusals stop costing a full timeout. -
Every operator-typed timestamp control now names the zone it is read in. Maintenance windows, gate overrides, backfill instants and the instance-silence deadline are entered in controls whose HTML value carries no offset, and they said only “Starts”, “Ends”, “Until”, “from”. The date ranges on the change and gate-decision views are read as UTC calendar days and said only “From”/“To”. Each now states its interpretation, with the offset resolved AT the instant typed, so a window entered across a DST change is not labelled with the wrong one.
iter-0178, D-0246 · make dev-test-distributed 11 passed / 1 skipped · geo topology 14
passed, including credentialed dispatch on both remote transports — a path no geo run exercised
before · make secret-smoke and make mac-smoke green · full -race 33 packages exit 0 ·
iter-0178 CLOSED by the owner
Part 4 — the repairs a two-axis review of this range found (audit gap package 3)
Section titled “Part 4 — the repairs a two-axis review of this range found (audit gap package 3)”Fifty-two findings from an independent review of everything since v0.1.8, in seven clusters. The
package adds no requirement: it repairs FR-020, FR-026, FR-029, FR-030, FR-031/NFR-025 and FR-032,
which is why it carries an iteration number and no FR-. Three of the findings were P0.
⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- Migration
00104can REFUSE the upgrade, deliberately. It narrows thepull_testscarrier ceiling to generations 1–3 and stops with a count if the table holds a row atprotocol_version = 4. No build emits a generation-4 TEST carrier, so such a row was not written by this product, and the migration will not discard it for you: silently deleting an unexplained row destroys the evidence of whatever wrote it. Nothing is lost when this happens — the database stays at00103— andrunbook.mdcarries the query to inspect the rows and the reason the asymmetry withpull_jobsis correct. - Migration
00105dropsexpected_runs_job_idx, which served no query in the tree: every statement that mentionsjob_idalso constrainsdue_at, and the primary key is(monitor_id, due_at). ON DELETE SET NULLcolumn-list migrations: SIX, not five.00093added the sixth when the reliability gate landed, and only the README and the refusal message were corrected then. Eight other statements — the overview, the runbook’s count and its list, two rows ofstatus.md, two passages ofdecisions.md, and a test comment — still said five. The PostgreSQL 15 floor itself is unchanged; what was wrong was every document’s count of why.
🩹 Fixed
Section titled “🩹 Fixed”- A credentialed-by-SCHEMA monitor with no credential value was stamped with the wrong carrier
(P0). One producer branch skipped the stamp, so the ledger read those windows as
unknownand coverage was lost, silently, for a whole class of monitors — probes ran and results correlated. - An external canary kept its binding headers across a scheme-downgrading redirect (P0). The hop
is now refused before any header operation, and the origin comparison includes the scheme, so a
redirect from
httpstohttpon the same host and port is cross-origin rather than “the same place”. - An undispatchable canary reported nothing at all (P0). A canary in a region with no capable runner now flips the monitor DOWN through the ordinary result pipeline instead of sitting on a queue until its TTL — the indefinite pending the invariant forbids.
- A canary workflow with no completion block dereferenced a nil pointer, which one crafted message could use to take down a region’s whole prober pool. It is refused structurally now.
- The expected-run ledger’s read API listed rows it declared unanswerable. The retained floor
ignored the DEFAULT partition, so rows stored there were returned by the list while
ledger_fromsaid the ledger could not speak for that span — two answers, one of them provably wrong. The retention cutoff is also midnight-aligned everywhere, matching the purge that drops partitions on that boundary. - A batch of expectation advances failed as a whole when one item was unrepresentable. The healthy items reserve and publish; the rejected one is named in the log with its reason.
- The reliability timeline drew cells wider than the windows they belonged to, so a covered cell painted over the missed windows beside it and the pointer targets overlapped identically — the readout named a window the pointer was not over. Cell width now comes from the window grid, and each pointer target is bounded by half the space to its neighbour, so two targets can touch and never overlap.
- An on-call rotation anchor moved by the viewer’s offset on a cosmetic edit. The control is pre-filled in the viewer’s zone and saved back as an instant; CI now runs the frontend suite a second time in a non-UTC zone, without which the test proved nothing.
- Six more rendering repairs: a storage verdict printed from a series that had not been read; a panel subtitle describing strokes it had not drawn; a legend naming all observations where the figure was computed from measured ones; a timeout declared in scale from an unrelated average; a UTC day range that hid its clock; and a retention bound the API description had hardcoded.
🔧 Changed
Section titled “🔧 Changed”- The carrier decision has one owner, and the pull-region exclusion one reader. The scheduler held three independent resolvers for “what can this region’s executors consume”; they now share one record, and a source guard fails by name if any of them goes back to reading the map directly.
- Four guards were widened to what they claim. The
Test…citation guard reads Go comments in both spellings; the declared-door guard follows embedded interfaces at any depth and reports an element it cannot read rather than treating it as absence; the PG15 migration guard derives its sites from the tree instead of a hardcoded pair; and the rendering guard matches the locale formatter with or withoutnew.
iter-0179, D-0247 · 52 of 52 discharge rows built, each behaviour row with a recorded
killing mutation · independent review COMPLETE, 52 of 52, no open findings · full -race 33
packages exit 0 (internal/store 755.8 s) · make docs-check OK · SPA 683 tests at UTC and
Asia/Kolkata · make dev-test 70 passed / 1 skipped on an image rebuilt from the CLOSING tree,
verified by CONTENT rather than by build timestamp — the live server serves the closing SPA build
and answers the pre-E4 asset with the index fallback · make geo-test 14 passed,
run by the independent reviewer on a stack raised for him · iter-0179 CLOSED by the owner, on his
word and not on the review’s result
[v0.1.8] - 2026-09-03
Section titled “[v0.1.8] - 2026-09-03”A credential inside a synthetic monitor’s scenario used to be stored in cleartext, returned to every principal who could READ the monitor, skipped by key rotation, and echoed into the heartbeat message with the whole request URL whenever a step’s transport failed. This release closes all four, in three stages that were built and reviewed separately, and gives the editor a way to declare a credential without ever typing one. It BREAKS a synthetic monitor that keeps a token in an authorization-style header — read the upgrade note. No migration.
This was a spec-versus-code defect, not a new feature: func-oncall-synthetic-pull.md FR-SYN-1
already promised “encryption like the other types” and its §217 promised inclusion in reencrypt, and
D-0090 promised the failure message “never echoes bodies/headers”. None of the three was true of the
code, and the SYN requirements had no rows in docs/status.md at all, which is how a false promise
survived unnoticed. They have rows now.
⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- A synthetic monitor whose credential-bearing header holds a LITERAL now fails validation on its
next write. The affected headers are the finite credential-bearing set —
authorization,proxy-authorization,cookie,x-api-key,api-key,x-auth-token,auth-token,x-access-token,access-token,private-token— whose value must now be exactly one{{secret:<binding>}}placeholder. Existing monitors keep probing: nothing re-validates on read, so the refusal appears the next time such a monitor is EDITED, and it names the step and the header without echoing the value. Move the token into the project secret inventory and reference it as a binding (b3c99b6) - No migration, and no readiness coupling. The scenario stays in
monitors.config; only its form changes, to ciphertext. An instance with nosecurity.encryption_keystarts, reports ready and keeps probing every monitor it was probing: what it refuses is a scenario WRITE it cannot protect. Legacy plaintext scenarios stay plaintext until a key is supplied — that cost is stated rather than traded for an outage (cc90e4a) - Probe failure messages changed shape for every monitor type, not only synthetic. A transport
failure now reads as a bounded class plus the target’s host (
dns: no such host (api.internal)) instead of Go’s error text. If you have alert rules or dashboards matching onheartbeat.msgsubstrings, check them. Nothing in the product parses that field to make a decision (e6a3db8)
🔒 Security
Section titled “🔒 Security”A credential in a scenario is a secret at rest, on read, and in the record (FR-028 / NFR-023)
- Stage 0 — no probe result carries a request URL, for any type.
net/httpembeds the request URL in every error it returns, soMsg: err.Error()published whatever the target’s query string carried — reproduced on an ordinaryhttpmonitor and onpromql, not only on synthetic. Every failure is now composed from a bounded class plus a host, asserted per type through the real prober registry with a secret planted in the URL, the query, a header and the body (e6a3db8) - Stage 1 — the scenario is ciphertext at rest, and withheld from anyone who cannot write the monitor. One secret set became three classifications by MEANING (encrypted-at-rest, write-only-on-read, writer-only-display) and the store reader became an explicit MODE named by what it decrypts. A viewer receives no scenario at all — not plaintext, not ciphertext — through a store call chosen after the authorization decision rather than by redacting a decrypted document. Key rotation covers scenarios, and an idempotent, compare-and-set, NON-fatal startup backfill converts existing rows (cc90e4a)
- Stage 2 — a declared credential is a NAMED BINDING resolved from the project inventory. The
document carries
{{secret:<binding>}}and the secret’s NAME lives in a flat config key,scenario_secret_<binding>_ref, so rename, delete-counting and rotation run on the pathpassword_refalready runs on. The value is delivered in the credential envelope, substituted into the scenario by the executor, and the substituted copy is dropped when the probe ends (b3c99b6) - A moved placeholder cannot survive a valid envelope. The scenario and its reference keys became
EXECUTION BINDING KEYS, so
EnvelopeV2’s body digest covers the stored document: an attacker with a valid envelope who rewrites a step’s URL produces a different document and the AEAD fails before any request. A binding therefore REQUIRES a body-bound envelope, and a region on an older carrier gets a per-monitorcarrier_too_oldrather than a job that looks protected and is not (b3c99b6) - The binding belongs to a synthetic monitor and to nothing else, decided once at the write boundary and gated in the store, in both derived key sets, and at the dispatch gate — which refuses rather than ignores such a job on any other type, permanently, on carrier integrity (084b49d) (4fdedff)
What this does NOT claim, stated because a security note that overstates is worse than none. A literal secret is not detectable by shape, so the enforceable rule is the header-NAME one. A credential pasted into a header nobody would call a credential header, or into a body, is still legal: it is encrypted at rest, withheld from viewers and kept out of probe results, and it travels to the prober inside the ordinary job rather than in an envelope. Buying the stronger property needs a restrictive typed request model for the scenario, which is an owner’s decision and separate work — and a heuristic over VALUES is refused rather than deferred (900aa1b)
New Features
Section titled “New Features”Declaring a binding in the UI
- The synthetic monitor form has a Scenario secrets panel above the steps: pick a project secret, name the binding, and the row then shows which secret fills it, where it is used, and the flat key it is stored as. A binding name is never displayed without the secret it resolves to (3a67106)
- A credential-bearing header stops being a free-text field. Its value control is a binding selector from the first keystroke of the header name — empty and disabled with “Add a binding first” until one exists — so the rule is met before a token is pasted rather than after a failed save (95ee945)
- “Save before test” is stated at the button. A scenario carrying bindings is deliberately not testable before it is saved: that path builds an unsaved monitor with no envelope, so a placeholder would travel to the target as literal text. The Test control is disabled with the reason and the way forward; a credential-free scenario still tests unchanged (3a67106)
- What is NOT protected is shown too: a pasted-looking value in an ordinary header gets a hint offering the inventory — never a refusal, because cerbix cannot tell a credential from data there (3a67106)
Bug Fixes
Section titled “Bug Fixes”- A synthetic monitor could not be created from the SPA at all.
canSubmitrequired a target for every active type while the form deliberately hides that field forsynthetic— whoseNeedsTarget()is false — so Create monitor was permanently disabled and unexplained. Found by a new component test, and it survived because no unit test and no browser test had ever submitted this type (3a67106) - The form’s default scenario opened INVALID. The scaffold shipped
Authorization: Bearer {{token}}, which the new rule refuses, so merely choosing Synthetic produced a form whose own example could not be saved. It now demonstratesextract→ interpolate with an id used in the next step’s path, which is whatextractis for (95ee945) - A malformed binding reference was silently ignored.
scenario_secret_Login_ref— capitalised, so outside the grammar — meant the operator declared a binding, saw no error, and shipped a scenario the credential was never wired into. It is refused by name now (084b49d)
Improvements
Section titled “Improvements”Documentation that stopped lying
- FR-SYN-1, NFR-SYN-2 and AC-SYN-2 corrected in place, and the five SYN requirements given the
docs/status.mdrows they never had — including two gaps stated rather than implied: no test names the whole-scenario deadline, and no browser test puts a synthetic monitor on a GEO worker (4e6def8) - The binding key is documented in
openapi.yamlon all three monitor config schemas. The feature was API-reachable and undocumented, which makes it unusable rather than merely un-designed (084b49d) - The runbook gained an FR-028 section covering what is protected at rest, on read and in the record; the detection query and one-edit repair for a stale reference; and — as plainly — what is not protected (258adb5)
Repository hygiene
- The SPA build runs as the developer, not as root. Every docker build wrote
frontend/node_modulesand its caches as root, so any tool later run by a developer failed with EACCES on a tree it owned nothing in — it stopped an independent reviewer from running the frontend tests at all.make spa-snapshotnow passes--userand an in-container npm cache (c321931) - The first browser coverage a synthetic monitor has ever had (
e2e/tests/synthetic-bindings.spec.ts): declare a binding through the UI, meet the refusal, see the test blocked, save, and find the flat key and the placeholder on the wire with no value anywhere (3a67106)
API · Metrics · Schema
Section titled “API · Metrics · Schema”- API: the monitor
configdescription documentsscenario_secret_<binding>_refand its rules on create, update and test; ascenario_secret_*key on any non-synthetic type is refused with 400 naming the key; test-before-save answers 400 naming the binding for a scenario that declares one. - Metrics: unchanged. A binding that cannot be materialized surfaces as the existing per-monitor
reason (
carrier_too_old,missing_reference,decrypt_failed), never as a readiness flip. - Schema: no migration and no new column.
monitor_secret_refscarries a binding through the samesetting_keyshape aspassword_ref. The scenario’s VALUE inmonitors.configbecomes self-describing ciphertext, converted in place by the startup backfill.
15 commits · decisions D-0216 (+ addendum) and D-0217 · FR-028 and NFR-023 DONE, FR-SYN-1/2/3 and
NFR-SYN-1/2 given their first rows · iterations iter-0166 and iter-0167, both CLOSED and both approved
by the independent reviewer after five and four rounds · full -race suite green (33 packages, with
internal/store re-run idle to separate a pre-existing load-dependent FR-024 flake), vitest 38 files /
418 tests, browser suite 60 passed / 1 skipped
[v0.1.7] - 2026-09-02
Section titled “[v0.1.7] - 2026-09-02”PromQL grew up: it is expressible in a Monitoring-as-Code bundle and can authenticate to a Prometheus behind basic auth. One security fix rides with it, and it BREAKS a configuration that works today — read the upgrade note first. No migration.
⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- A monitor target may no longer carry credentials in its URL userinfo.
https://user:pass@prom.internal:9090is refused on every surface now, not only in bundles. It used to work — Go’snet/httpturns such a URL into anAuthorization: Basicheader on its own — and the password was stored as plaintext inmonitors.target, which the API’s redaction does not blank, so every VIEWER of the project could read it in the monitor list. Stored monitors keep probing: nothing re-validates on read, so the refusal appears the next time such a monitor is EDITED. Move the credential to the type’s own settings (promqlnow has them) or to a target without userinfo before editing (f827d96) - No migration, no configuration change. A
promqlmonitor with noauth_modebehaves exactly as before, at validation and at the dispatch gate — the default is resolved, never written, so no canonical hash moves and no monitor is rescheduled (754e5dc) - Worth knowing if you deploy the release BINARIES: the
v0.1.6assets were compiled with Go 1.25.12, which carries seven known standard-library vulnerabilities this release closes by moving the pin. The container image was never affected — it is built ongolang:1.26.6— so an installation running the image was not exposed by this (be75d51)
New Features
Section titled “New Features”PromQL: bundles and basic auth
- A
promqlmonitor can live in a Monitoring-as-Code bundle. It carriesquery— required, bounded at 1024 characters, never blank. The type was excluded because a 2026-08 classification grouped it with the credentialed types; the prober had in fact never had a credential slot, so what stood between it and a bundle was one typed field (39245fe) - Optional HTTP basic auth against Prometheus, through
auth_mode: none | basic.basicrequires ausernameand a credential — a project-secret reference in a bundle, value or reference in the UI/API — and the prober sends the header only when a username is configured, so an unauthenticated Prometheus is never handed an empty one. Bearer tokens and mTLS are deliberately NOT supported: basic is what Prometheus implements natively, and for anything else the answer remains a regional agent plus an unauthenticated path for the prober (754e5dc) - In the SPA: the PromQL section gained an Authentication selector, a username field and the same credential control the database types use — value or project secret, with the dangling-secret warning (05db534)
Security
Section titled “Security”- The Go pin carries the standard-library fixes
govulncheckasks for. Under the version the workflows install fromgo.mod(1.25.12) it reported SEVEN called vulnerabilities —crypto/tls×2,net/http×2,encoding/xml,encoding/asn1,html/template— reached through ordinary product paths: the HTTP prober’sClient.Do, the mailer’s TLS handshake, the operational server’sListenAndServe. All seven are fixed in 1.25.13, and one line moves the scan AND the release binaries, because all three workflows read the version fromgo.mod(be75d51) - The full-history secret scan is clean, without weakening it. Its seven findings were read one by
one and are fabricated test fixtures — two sequential-hex AES keys, the same two base64-encoded, and
an
actor_labelin a committed CLI transcript, which is the server’s derived name for a bearer and never the token. Nothing needed rotating. The allowlist entries are anchored to those exact literals, so any other high-entropy string in the same files still fails the scan (7a9350d)
Improvements
Section titled “Improvements”Correctness
- The file provider no longer conflates “has settings” with “is credentialed”. One entry point routes a type to the credential registry, to its own schema, or to an error naming the key — which is what the specification has said in words since the beginning (39245fe)
- A blank-after-trim
queryis refused. A presence check accepts" ", and a whitespace-only expression makes the prober report “no query configured” on every run: a monitor that is configured, scheduled and permanently meaningless (754e5dc) - A rejection reason is decided by a typed error, not by matching message text, so a reworded validator can no longer silently change which reason a bundle is refused with (754e5dc)
Repository hygiene
- The documented test command is the one that finishes. CI has passed
-timeout 40msince iter-0163 becauseinternal/storeneeds about eleven minutes under-race;make race,CLAUDE.mdandAGENTS.mddid not, so a developer running the documented command metpanic: test timed outnaming whichever test was running — a name that sends the reader after the wrong bug (9608fbf) - Dependabot got quieter without hiding anything: monthly instead of weekly, all MAJOR bumps in one PR per ecosystem instead of one each, lower open-PR limits, and a cooldown so a version released yesterday is not proposed today (18e3564)
- A roadmap exists (
docs/roadmap.md) and the PRD points at it; the red Security workflow is its first item (6e89c84, 1efd5dc) - A flake left the suite by being understood rather than retried. A scheduler test waited on the gate pass’s sink events and then asserted on the leader gauge, which the leadership loop publishes on an unordered path — so under load it read the state before the step-down landed. The assertion now waits for the signal it asserts on, with the same requirement and its own deadline (aa3cbfb)
API · Metrics · Schema
Section titled “API · Metrics · Schema”- API: the monitor
configdescription now names promql’s settings (query, and withauth_mode: "basic"a username pluspassword|password_ref); a monitor target with URL userinfo is refused with 400 on create and update. - Metrics: unchanged.
- Schema: no migration.
promqljoins the credential registry, somonitor_secret_refscovers itspassword_refthrough the same normalization every credentialed type uses.
14 commits · decisions D-0145 (addendum) and D-0215 · no requirement row: the type boundary is
FR-017’s, and extending it is an addendum to its decision rather than a new requirement · full -race
suite green under the new pin, browser suite 59 passed / 1 skipped, govulncheck and gitleaks both
clean
[v0.1.6] - 2026-09-01
Section titled “[v0.1.6] - 2026-09-01”Two requirements that turn reliability facts into something a pipeline can act on: a release gate
that answers whether the error budget allows a deploy, and change intelligence that records the
deploy went and lets the service’s own facts say what followed. Eighty-five commits since v0.1.5.
⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- Two additive migrations, no data repair, nothing to stop.
00093creates the gate’s policy, override and decision tables plus the partition registry;00094creates the change tables and addsapi_tokens.actions. Runcerbix migratewith the new binary and start as usual — every role applies migrations on startup anyway. On a database where no service has a gate policy, nothing evaluates and nothing is written (16eecaf, c6c74dd) 00094builds a UNIQUE index onincidents. The constraint(id, project_id)is what lets an incident↔change link be tenant-safe by the schema rather than by a query. It is built non-concurrently and takes a brief exclusive lock onincidents; the guard skips the build entirely if an equivalent constraint already exists (c6c74dd)- The scheduler leader gains one background pass. Gate-ledger partition maintenance runs on its own
fenced advisory session — daily partitions created a week ahead, dropped past
gate.decision_retention_days(default 90). Watchcerbix_gate_decisions_writable_horizon_secondsandcerbix_gate_decisions_partitions_pending_drop; a healthy pass keeps the first well above zero (a6a9915) - Existing API tokens are unchanged. The new
actionsallow-list is optional: omitted ornullmeans the token’s role decides, exactly as before. A token is only ever NARROWED by it — the list is intersected with the role, never added to it (7260e68) gate.*andchange.*ship with defaults, so no YAML edit is required. The ones worth knowing: gate decisions retained 90 days and purged hourly; change groups retained by whole identity; the record endpoint bounded at 300 requests/minute per process and 30 per principal (16eecaf, c6c74dd)
New Features
Section titled “New Features”Reliability Gate — a deploy asks whether the error budget allows it (FR-024)
- One call, one machine-readable answer.
cerbix gate check --project <id> --service <id>(orPOST /api/v1/projects/{p}/services/{s}/gate) returns an observedstate—ALLOW,WARN,BLOCK,UNKNOWN,NOT_CONFIGURED— an effectiveaction, every matching reason and the evidence under it. The CLI’s exit code follows the action (0allow/warn,2block,4not configured,1transport/auth), so a CI step is one command; credentials come from the environment only (d926568, 229b33e) - What blocks is declared, not guessed. A per-service policy names ONE SLO window and assigns each
clause of a closed vocabulary — budget exhausted, budget consumed over a threshold, page-burn firing,
ticket-burn firing, an open service incident — to
block,warnorignore, with a mandatoryunknown_behaviorand a seal-lag bound past which the budget is UNAVAILABLE rather than quoted stale (8cb12d6) - The gate derives nothing. The service row, the policy, the active override, the report, the burn
latches, the incident and the ledger write all happen in ONE
REPEATABLE READtransaction whose first snapshot-bearing statement is the decision’sevaluated_at— so a gate answer and the service page cannot disagree about the same instant (8cb12d6, b7518b9) - An override changes the action and never the facts. At most seven days, one active per service, bound to the policy revision, project-admin only; the decision still records the state that was observed and that somebody overrode it (229b33e)
- Every decision is an immutable ledger row, in a daily-partitioned bounded table, readable by id after the service is renamed or deleted — the moment the evidence is actually wanted. Partition maintenance runs as a fenced pass with a 30-second lifecycle and ownership proved by marker rather than by OID (a6a9915, 24e901f)
- In the SPA: a
Release gatecard on the service page with the policy editor, the latest decision and the override panel; aGate decisionsbrowser with a server-side state filter and keyset paging; the by-id record; and the per-service override history. Opening a page never creates a decision — only a pipeline does (9793758, 84adac1, 3fc19e3)
Change Intelligence — the pipeline says what it changed (FR-025)
- A change is a fact about time, not a catalog.
cerbix change record(or onePOST) reports adeploy,rollbackorflagin one of its phases —started,succeeded,failed,cancelled— under an external identity(source, external_id), optionally naming the gate decision the release rested on. Phases are append-only and idempotent: an identical retry returns the original row, a contradictory one is refused by name, and two runners reporting different endings for one run cannot both pass (c6c74dd, 74ce187) - The service timeline is a bounded
[from, to)read of change groups with an opaque cursor that never returns a group twice, each group carrying its live gate decision and the incidents it preceded (7260e68) - Incident correlation says “preceded”, never “caused”. When a service auto-incident opens, the
changes within the correlation window on that service and on its
probable_rootupstreams are linked and named in one🚀 Changes:note. It is fail-open in both directions: the incident opens and resolves exactly as before whatever the correlation does (7260e68, 3b43b70, f5654d7) - Before/after SLI around a change, computed from SEALED buckets through the same query the
reliability page uses — never a second implementation. Each side is a figure, or withheld with the
page’s own word, or
pendinguntil the seal reaches it; a delta only when both sides are figures (c6c74dd, 5cace76) - A CI token can be narrower than a role. An optional
actionsallow-list on an API token is intersected with the role inauthz.Can, so a pipeline token can be exactly “ask the gate, record a change” and nothing else (7260e68, a1cfe00) - In the SPA: a
Changescard beside the release gate with terminal-only marks on the facts strip, the timeline view, the comparison view, andPreceded byon the incident page (d2e0f3c, 42ae4ab)
Notification channels are edited in place
- A channel’s name and config are editable, so rotating a bot token or a hook URL no longer means delete-and-recreate — which silently dropped every monitor link, escalation step and alert route pointing at the channel. A secret left blank keeps the stored value, because the API never sends one out; the merged config is validated, so an edit cannot leave a channel undeliverable (3e76791)
Improvements
Section titled “Improvements”UI
- The alerting panel keeps an operator’s unsaved edits when a late prop arrives instead of discarding them (dd83bfa)
- The gate ledger’s state filter is the server’s, so a page of results is a page of matches and the cursor continues the filtered set (84adac1)
Documentation & Gates
- Every surface now says what the product is. D-0174’s positioning — a service reliability platform, not “uptime & SLA monitoring” — reached the README at iter-0160 and stopped there; the CLI’s help, the OpenAPI description, the overview, the onboarding doc, the systemd unit and the OIDC client all still carried the pre-FR-021 framing. Claims that the repository is private are gone with them; they told a reader to authenticate for things that need no authentication (92edc5b)
- The PRD describes the product as it is now — services, their incidents, their escalation ladder and change intelligence — and points at a roadmap that exists (4066527, 59d05c4)
make docs-checkcompares the FR-025 acceptance map as a SET and refuses the spellings the design retired, in the specification and in every living document (056eff5, b320290)- The incident-audit gap has a requirement. FR-026 / NFR-021 are specified and approved at revision four after three review rounds; no product code changes until the iteration opens (9208171)
API · Metrics · Schema
Section titled “API · Metrics · Schema”- API: eleven gate routes (the decision, the policy CRUD with
expected_revision, the override lifecycle and history, the project-scoped ledger read and listing); four change routes (record, timeline, compare, an incident’s preceding changes);ApiToken.actions;PATCH /api/v1/notification-channels/{id}now acceptsnameandconfigas well asenabled; the service detail carries itssla_targetsinventory. - Metrics: the gate family —
cerbix_gate_decisions_total{state,action,overridden},cerbix_gate_decision_duration_seconds(this project’s first histogram),cerbix_gate_evaluate_rejected_total,cerbix_gate_evaluate_errors_total,cerbix_gate_maintenance_errors_totaland four ledger gauges; and the change family —cerbix_change_correlations_total,cerbix_change_correlation_errors_total,cerbix_change_compare_total,cerbix_change_record_rejected_total(db16dfa). - Schema: migrations
00093–00094— gate policies, overrides, the daily-partitioned decision ledger and its ownership registry;service_changes,incident_changes,api_tokens.actions, andUNIQUE (id, project_id)onincidents.
85 commits · decisions D-0188…D-0214 · independent review: every FR-024 range approved, FR-025
approved as four effective slices plus the live-evidence correction · full account in
docs/iterations/iter-0163.md, docs/iterations/iter-0164.md, docs/iterations/iter-0165.md
[v0.1.5] - 2026-08-28
Section titled “[v0.1.5] - 2026-08-28”⚠️ Upgrade Notes
Section titled “⚠️ Upgrade Notes”- Stop the outbox owners before migrating. Roles
all,apiandschedulerrun the outbox worker;workerandagentdo not and can keep running. Runcerbix migrateonce with the new binary, then start the owners. Migration00088hands ownership of a class of outbox rows to the database and cannot reach a delivery an old worker already has in flight. Skipping the stop loses nothing; it risks out-of-order delivery for at most one already-claimed batch (≤50 rows) per old owner (fc6608d, 88ec4ad) - Migration
00090repairs incident data written by earlier versions. Incidents resolved and then walked backwards, and service incidents stranded open after their alert ended, are resolved with a🔧 Repaired:timeline note. Two classes are reported in the migration output but deliberately left alone — member snapshots that may name a not-yet-governing revision, and auto-incidents with no monitor or service left — because fixing them would mean guessing at history. Queries and the manual procedure are indocs/runbook.md. On a database that never hit the races the migration is a no-op (af7a7a2) - PostgreSQL 15 or newer is required. 14 is not supported and will not be;
cerbix migraterefuses before applying anything (e53244b) - Armed services keep their coverage across the upgrade.
00089seeds delivery tracking as “delivered” for existing rows, so no member monitors start paging in the minute after the upgrade (7a9e87d)
New Features
Section titled “New Features”Service Alerting: Coverage Means Somebody Was Told
- A route must name a live channel. A schedule pointing at a deleted or disabled channel no longer counts as a route; members keep paging instead of being silenced in favour of a replacement that can reach nobody (fb9e94b, 08a8b31)
- An alert nobody can receive is withheld, not sent into the void. Counted as
cerbix_service_alert_withheld_total{signal,reason}withunroutableorno_governing_revision, and announced as soon as a route exists — on both the health and burn signals (44dfdad, a3aecfc, 19d4943) - Coverage requires a delivered announcement. A service suppresses its members only once its own alert reached at least one recipient — a channel row that exists is not a delivery, and a 500 from the only channel is not a delivery either (7a9e87d, 5ec101b, 35de54d)
- Coverage follows the state the service is in. A delivered DEGRADED announcement no longer covers a service that has since observed DOWN and is still confirming it (6e6339f, 46c50de)
- An outage nobody heard about is announced again once there is somebody to tell — channel deleted mid-flight, every send failed, or retries exhausted — with a fresh recipient list and a new episode. A partial delivery that reached some recipients is never re-sent to them (23aa3cc, 425a59f, 2b2594c)
- One vocabulary for “why did my monitor page”. The service badge and
cerbix_alert_delegation_fail_open_total{reason}are computed from the same clause evaluation. New valuesno_owning_service,onset_pending,onset_undelivered,latch_inconsistent;stale_leaseon the burn arm now means an expired lease and nothing else (18e3aef, 3192ebb, 6e0b6a4, 4c4bae3)
Service Escalation: Repeat Cadence
- Services can repeat the last escalation step.
renotify_secondson the service — 0 is off and the default, otherwise 60..86400 — read live, so turning it down mid-incident takes effect immediately. Available in the UI, the API and monitoring-as-code bundles (alerting.renotify_seconds). Previouslyrepeat_laston a service-attached policy did nothing (90bf146, 425a59f) - An incident climbs the ladder it started with. The escalation policy is snapshotted when the incident opens, so editing a policy mid-outage cannot re-time a page already in flight. Incidents open across the upgrade start their ladder at the upgrade instant instead of firing every overdue step at once (0c6e5dd, d3f428f, 118c55c)
Outbox: Ordered and Bounded Delivery
- Incident webhooks are dispatched in order. Every
incident.*payload carriesseq; the claim will not release an event while an earlier one for the same incident is undelivered, and a dead predecessor blocks too. Arrival order over the network is not promised —(incident.id, seq)is what lets a receiver dedupe and order, and the runbook describes two receiver strategies (554e609, 8ed191a, 02bc005, 4762f1a, 64349a0) - A delivery is bounded by the lease that authorised it. A deposed worker can no longer keep an HTTP request or SMTP session open while the new owner sends the same event. The lease is measured in database time, the settle is never bounded, and a claim whose turn came after its lease is handed back with its attempt refunded (b17e0e2, 425a59f, 2b2594c)
Improvements
Section titled “Improvements”Incident Lifecycle
resolvedis terminal and status only moves forward, enforced in the write rather than in a stale read. A plain comment keeps the current status instead of carrying whatever the client last saw (b21b09f, 183b9f8, 451406b)- Service incidents are a full lifecycle. They announce open and resolve to webhooks and status-page subscribers, and resolve when their alert ends — including through disown and delete, which used to leave them open forever (e405d85, 744980f)
- The postmortem names the declaration that governed the outage, not the newest one; a foreign revision is refused (0a8ba9b, 335d4db)
- Every incident write stamps its times after the row lock, so a writer that waited cannot date its action before the wait (3ad2a4b, 5ec101b)
Status Pages
- A page, its feed and its subscriber mail agree on what the page reports. A page made only of Service components used to show an incident and email nobody about it. One project axis now, with the subscriber query as its exact inverse (c7ce059)
UI
- Search hits bring their workspace. Opening a monitor or incident from another project switches to it, so edit controls are not hidden from a legitimate editor; detail views follow the URL when navigating between two of the same kind (2dbc7ce)
- A partly unmeasured hour no longer renders green on the Reliability card (27c63f7)
- The service picker in status-page components works again, and the escalation form says what “repeat last step” does on a service (a26260d, a331af0)
Documentation & Gates
make docs-checkcompares the FR-021 invariant set in the spec against the discharge map exactly — missing, extra, duplicate or skipped numbers all fail; invariants 92–105 moved into the spec (46c50de, 35de54d)- Four documents that still described the product as it was before FR-022/FR-023 shipped are corrected, and a gate catches the class (afc8cd1)
API · Metrics · Schema
Section titled “API · Metrics · Schema”- API:
ServiceAlertPolicy.renotify_seconds; alerting-statereasongainsonset_pending,onset_undelivered,latch_inconsistent;incident.*webhook payloads carryseq. - Metrics: new
cerbix_service_alert_withheld_total{signal,reason};cerbix_alert_delegation_fail_open_total{reason}also emitserror,record_failed,unspecifiedfor lookup-level failures. - Schema: migrations
00085–00092— incident escalation snapshots,incidents.event_seq,CHECK ((status = 'resolved') = (resolved_at IS NOT NULL)),delivered_seq/undelivered_seqon both service latch tables,services.renotify_seconds,undeliveredas an episode close reason.
57 commits · decisions D-0175…D-0187 · independent review: product approved at 35de54d, follow-on work at eba2b69 · full account in docs/iterations/iter-0161.md
[v0.1.5-beta.2] - 2026-08-19
Section titled “[v0.1.5-beta.2] - 2026-08-19”Three defects found by running v0.1.5-beta.1 in production rather than by a test. The first blocked the
upgrade outright.
- PostgreSQL 15 is now enforced instead of assumed. A production upgrade to
v0.1.5-beta.1on PostgreSQL 14 applied00061…00069and died on00070withsyntax error at or near "(": five migrations use the column-listON DELETE SET NULL (col)form introduced in PG15. Every document, image and CI job already said 16, but nothing checked it and nothing said it out loud.cerbix migratenow readsserver_version_numbefore the first file and refuses with the version, the requirement and the fact that nothing was applied. README,runbook.mdandoverview.mdstate the requirement; the runbook also carries the recovery note for a system left partially migrated (00065makesmonitors.slugNOT NULL, which an older binary does not write). - A status-page component could not be created from a Service. The service picker rendered a blank
option AND carried no value, because the view read
sv.name/sv.idfrom a list endpoint that answersServiceSummary(the service wrapped with its rollup counts). Anas Service[]cast on the load is what stopped the compiler from reporting it, and the test fixture repeated the same wrong shape, so six passing tests never saw it. - A stale claim on the service page. The footer said availability, the error budget and the burn rate “arrive with the next iteration”. They arrived in iter-0144 and the page already renders them; what was actually missing on a fresh service is sealed facts and a declared objective, which is what it says now.
[v0.1.5-beta.1] - 2026-08-19
Section titled “[v0.1.5-beta.1] - 2026-08-19”242 commits since v0.1.0-beta.5. This is the release where a Service becomes the object
reliability is defined on, measured for, and paged about — and where the product’s own
positioning was corrected to match (D-0174).
- Service reliability, phases 3–5 (FR-021 / NFR-016, D-0169) — closed against an ENFORCED
discharge map of 91 acceptance invariants and 24 required scenarios, each naming a test that
exists;
make docs-checkfails if a number lacks a row or a row cites a test the tree lacks.- dependency impact graph — a same-project service DAG (schema-enforced tenancy, bounded, outside the declaration), symmetric open-time correlation into structured incident↔service links with 🕸 timeline notes. It annotates and links: it records candidates, never elects a culprit, and never suppresses or hides.
- status-page projection (§15.0) — a component renders from ONE of three sources under a
discriminator (
monitor/service/manual) with the replaced binding kept dormant for revert. Public-output change: the page summary is worst-of-MEASURED plus an unmeasured count, and measurement ABSENT is the public statusno_data— neveroperational. - alerting ownership (§16) — a Service can be the thing that pages, and its declared SLI members can stop delivering their own alerts for the same failure. Suppression is per SIGNAL (live health, sealed burn), per POLARITY (onset-like only; a recovery is never suppressed) and only while a replacement is demonstrably ARMED. Anything ambiguous fails open — the member pages.
- service burn alerting — arbitrary long/short windows over one burn-math owner, with the hold matrix and the watermark every number was computed from.
- Service incidents (FR-022 / NFR-017, D-0170, D-0171) — an incident can be an incident OF a Service. At most one anchor, enforced by CHECK; opened in the SAME transaction as the announcement and resolved by its close; never on a burn breach; at most one open auto-incident per service; a member snapshot a postmortem can still name after the world moved; impact links through the service graph, with no link ever naming its own subject. Closed against 16 invariants + 16 scenarios.
- Escalation for services (FR-023 / NFR-018, D-0172, D-0173) — a Service with an escalation policy escalates its own auto-opened incident: steps from the incident’s start, durable progress, acknowledgement or resolution ends it, every step names the SERVICE. The ladder fails closed where delegation fails open, and the service graph does not pause it. Closed against 16 invariants + 19 scenarios.
- Project-level SLO objective and an instance-wide audit surface for a global admin’s own actions, which had been recorded for months and shown nowhere.
job_idcorrelation end to end, andobserved_atordering that refuses a result older than the issue it answers.- A write path for a service’s escalation policy — the column existed since phase 5 and was reachable only at create time or from a file provider; a change to who gets woken is now audited with what moved, inside the mutating transaction.
cerbix_service_incidents_total{action}andcerbix_escalation_steps_total{subject}— the on-call ladder had no metric at all before, only a log line.
Changed
Section titled “Changed”- Positioning (D-0174) — cerbix is a service reliability platform, not “uptime & SLA monitoring”. The README states its NON-GOALS publicly, quoted from the specification: no arbitrary time-series queries, no generic telemetry, no query language, no metrics backend, no service catalog, no trace or log ingestion, and no automatic root-cause analysis.
ServiceDetail.reliabilitystaysnullby design and now SAYS so: SLO, error budget and burn rate live onGET …/services/{id}/reliabilitywith the honesty context a bare number would lack. The previous description claimed they were unbuilt.- Dependencies —
golang.org/x/crypto0.55.0,x/net0.58.0,x/text0.41.0; build image golang 1.26.6; pinia 4.0.3; and four frontend majors (@vitejs/plugin-vue6,npm-run-all29,unplugin-auto-import21,@vueuse/core14), each merged and verified one at a time.
- A public leak found by reviewing our own invariant:
Incident.PublicRedactedclearedproject_id,monitor_id, the external key and the ack actor — and not theservice_idadded days earlier, so every unauthenticated render of a page with a service incident shipped the service’s internal UUID. - CI had never run on this line of work: tests triggered only on
pull_request,docs-checkran only by hand, one storage mode was covered, and readiness was awaited withpg_isready, which this project’s own discipline calls a non-barrier. All four repaired. - Two flaky tests fixed at their cause, not re-run: a fence test that bet a budget refusal on 5ms of wall clock (2 failures in 12 runs under load), and scheduler telemetry waits that asserted more series than they waited for.
Migrations
Section titled “Migrations”24 new migrations (00061…00084), forward-only and applied automatically by every role.
Compatibility
Section titled “Compatibility”FR-021 §17 makes backward compatibility an acceptance criterion, not a footnote: zero Services is
a valid installation state, every existing Monitor stays valid without a service, bundle format 1
stays valid, existing composites and monitor SLOs keep their semantics. The one intentional
break is the public status-page output described above — a consumer that read operational for
an unmeasured component will now read no_data.
[v0.1.0-beta.1 … v0.1.0-beta.5] — released 2026-07-25 … 2026-08-12
Section titled “[v0.1.0-beta.1 … v0.1.0-beta.5] — released 2026-07-25 … 2026-08-12”Corrected on 2026-08-19. This block sat under
[Unreleased]while its contents were shipping across the fivev0.1.0-beta.*tags — nobody moved it out. It is relabelled rather than rewritten, because it is a record of what happened, not a plan. The per-beta split is not reconstructed here: the tag messages carry it, and inventing a division after the fact would be a worse claim than admitting the block covers the whole beta train.
- Geo-Distributed HTTP Pull Agent (
--role agent) with Long-Polling (LISTEN/NOTIFY), Edge Ring-Buffer (bufferCap=10000), and historical backfill (POST /agent/backfill). - Observability for Pull Transport: Prometheus gauges
cerbix_pull_jobs_pending{region}andcerbix_pull_agent_lag_seconds{region}with automatic lagging alerts. - Database Agent Tokens: Table
agent_tokensand Admin API (POST/GET/DELETE /api/v1/agent-tokens) for issuing, listing, and revoking agent tokens without redeploy. - Region Scoping: Enforcement of
monitor.region == agent.regionon/agent/resultsand/agent/backfill(403 Forbidden on mismatch). - 16 Prober Types: Added probers for PostgreSQL, MySQL, Redis, RabbitMQ, PromQL, gRPC, WebSocket, SSH, DNS, TLS cert expiry, Composite, and Synthetic multi-step HTTP scenarios.
- On-Call & Escalation Engine: Escalation ladders, on-call schedules with vacation overrides, and acknowledge-to-stop incident handling.
- Prometheus Alertmanager Receiver: Inbound webhook receiver (
firing-> auto-incident,resolved-> auto-close by fingerprint). - Instance Settings Framework: Database singleton
instance_settingsfor branding, auth policies, SMTP mailer, and global silence toggle.
Changed
Section titled “Changed”- OIDC Provider Independence: Identity provider is now any OpenID Connect issuer (Keycloak, Auth0, Okta, Google, Entra ID) discovered via
oidc.issuer. - Database Schema: Renamed
keycloak_subtooidc_sub. - Heartbeat Retention & Partitioning: Switched
heartbeatsto native daily RANGE partitioning with automatic retention purging (retention_days).
Security
Section titled “Security”- SSRF Guard: Prober target resolution validated by
prober.Guard, blocking cloud metadata (169.254.169.254) and link-local ranges by default. - AES-256-GCM Secrets at Rest: Keyring encryption for webhook secrets and channel credentials with zero-downtime key rotation (
cerbix reencrypt).
Note on the two entries below.
v0.35.0andv0.10.0belong to an EARLIER numbering scheme and correspond to no git tag in this repository — the tag line isv0.1.0-beta.*. They are kept because they document real work; treat their version numbers as historical labels, not as releases anybody can check out.
[v0.35.0] - 2026-08-01
Section titled “[v0.35.0] - 2026-08-01”- Global Search: Tenant-scoped search endpoint (
GET /api/v1/search) across monitors, projects, and incidents. - p95 Latency: SLA reporting enriched with
p95_latency_msviapercentile_cont(0.95). - UI 1:1 Design Sync: Vue 3 SPA views rebuilt matching modern dark-theme design artifacts with bespoke inline-SVG charts.
[v0.10.0] - 2026-07-25
Section titled “[v0.10.0] - 2026-07-25”- Transactional Outbox: Outbox delivery pipeline for notifications and incident webhooks with exponential backoff and dead-letter queue.
- SLA & SLO: Error budgets, maintenance window exclusion, and daily availability rollup aggregation.
- Single Binary Multi-Role Execution:
--role all|api|scheduler|workerprocess execution model.