Authoring Alerts#

An alert is written once, in the query registry, and becomes a rule through mz-monitoring-build gen-rules. An alert whose query is PromQL is installed as a PrometheusRule wherever it applies, and the Thanos ruler evaluates it. An alert whose query is LogQL is installed as a PrometheusRule labelled mzmon.materialize.cloud/flavor: logql, which the alloy-gateway writes into the Loki ruler, and the Loki ruler evaluates it. This page is the conventions a contributor follows when adding or changing an alert, and what the tooling checks on their behalf. Recording rules render through the same context and are Authoring Recording Rules. Where the alert goes once it fires is Alert Channels; why the stack is shaped this way is the alerting design doc.

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in RFC 2119.

From registry entry to installed rule#

StageWhereWhat happens
Authorpackages/queries/*.yaml, an alerts: entryThe alert names its query inline or by queryId, its severity and component labels, and its prose; the file’s alertLabels supply its audience
Validatebin/mz-monitoring-check check-queries (pre-commit)Schema validation, including every %%{…} placeholder name
Rendermake rules (mz-monitoring-build gen-rules)Renders through the alerting context, infers capabilities, and fails on any problem
Outputcharts/materialize-monitoring/pre-rendered/rules/One groups: file per registry file and ruler, PromQL in prometheus/ and LogQL in loki/, and one _index.yaml listing both
Installtemplates/alerts/prometheusrules.yaml, templates/alerts/lokirules.yamlFills the placeholders from values and keeps the rules that apply, the same way for both engines
Checkmake rules-checkpromtool check rules and a LogQL parse with logcli on several rendered scenarios, then promtool test rules

The generated files are committed, and CI fails when they are stale.

Where an alert goes#

An alert belongs in the file for the people who act on it.

FileaudienceHolds
materialize-alerts.yamlplatformThe Materialize deployment: environmentd, the system clusters, clusterd crashes, persist, auth and the console
materialize-log-alerts.yamlplatformThe Materialize deployment, detected in its log lines: panics and correctness violations
materialize-workload-alerts.yamlworkloadWhat runs on the deployment: user clusters’ freshness, hydration and sizing, and the sources feeding them
infra-alerts.yamlplatformThe Kubernetes platform under Materialize, and the monitoring stack
infra-log-alerts.yamlplatformThe monitoring stack, detected in its own log lines where its metrics travel the path that is failing

Each file sets audience for all of its alerts with alertLabels, and an alert MAY set its own. gen-rules rejects an alert whose audience is not platform or workload. A route matches on the label to send the two to different people.

Who acts on a signal can depend on the cluster it comes from. A user cluster falling behind is the workload owner’s to fix, and a system cluster falling behind is the platform’s. Such a signal MUST be two alerts, one in each file, each scoped by cluster id.

SeriesScoped by
environmentd’s per-collection series, such as mz_dataflow_wallclock_lag_secondsinstance_id=~"u.*" or "s.*"
clusterd’s own series, such as mz_metrics_resource_usagecluster_environmentd_materialize_cloud_cluster_id=~"u.*" or "s.*"
kube-state-metrics and cAdvisor seriesthe replica pod name, pod=~".*-cluster-u[0-9]+-replica-.*", lifted into the cluster-id label with label_replace where the alert names the cluster

The alerting context#

A dashboard renders a query against variables its viewer chooses. A ruler has no viewer, so alerts render through a third context, packages/mzmon-lib/src/query/rules/context.rs, which differs from the dashboard context in two ways.

Deployment-specific values render as placeholders. Rules are rendered once, at build time, while which namespaces hold Materialize, which workloads the cluster counts as core infrastructure, and which metric prefix Materialize uses are facts about one install. Those parameters render to a __mzmon_*__ token that the chart replaces at install time.

ParameterRenders toThe chart fills it from
mzEnvironmentNamespaceFilternamespace=~"__mzmon_environment_namespaces__"rules.namespaces.environment, else materialize.namespaces, else materialize-system.namespace
mzOperatorNamespaceFilternamespace=~"__mzmon_operator_namespaces__"rules.namespaces.operator, else materialize-operator.namespace
excludeMzDeploymentNamespaceFilternamespace!~"<operator>|<environments>"both of the above
excludeEnvironmentFilternamespace!~"__mzmon_excluded_namespaces__"rules.namespaces.exclude
mzSqlPrefix__mzmon_sql_prefix__materialize.deploymentMode: mz_, or v2_mz_ for cloud
infraCoreWorkloadList__mzmon_core_workloads__, a bare value: container=~"%%{infraCoreWorkloadList}"rules.infraWorkloads.core
infraImportantWorkloadList__mzmon_important_workloads__rules.infraWorkloads.important
infraNonessentialWorkloadList__mzmon_nonessential_workloads__rules.infraWorkloads.nonessential
infraDaemonsetWorkloadList__mzmon_daemonset_workloads__rules.infraWorkloads.daemonset
mzEnvironmentFiltermaterialize_cloud_organization_name=~".+"nothing; a rule covers every environment
consensusRdsResources, consensusCloudsqlResources, consensusAzurePostgresResources__mzmon_consensus_rds__ and so on, bare valuesexternalDependencies.consensus, per flavor; see Install-time facts

An empty namespace list or workload tier renders as a^, a regex that matches nothing, rather than as an empty string.

An infrastructure alert MUST read a workload tier rather than list container or Deployment names. Which workloads a cluster cannot run without is a fact about that cluster, and a list written into the expression is one no operator can correct. Each tier holds container and Deployment names alike, so the same parameter serves container=~ and deployment=~.

Selection parameters are absent. interval, range, rangeWindow, the log pickers and the generation filters mean “whatever the viewer chose”. A query that uses one fails to render in this context. An alert that needs a window MUST write it out, since the window is a decision about how long a condition persists, and it belongs to the alert.

The mzEnvironmentName function attaches materialize_cloud_organization_name to a series by joining on the given label, normally namespace, against the Materialize scrape targets. A namespace holding two environments stays unlabelled instead of failing the evaluation, and a series with no match passes through unchanged. The mzClusterName function attaches cluster_name from mz_cluster_info, joining on the namespace and the cluster-id label it is given. A cluster id is unique only within one environment, so the dashboard join, which keys on the id alone, is not used here. mzObjectName is not available to rules, because nothing yet scopes its catalog join to one environment.

Capabilities#

A rule installs only where every capability it requires is present. Capabilities name what a deployment contains, such as a CNI, a metadata-database flavor or an exporter, and never who operates it.

Most requirements are inferred. gen-rules reads the metrics an alert names and maps each through the ordered table in packages/mzmon-lib/src/query/rules/capability.rs, so an alert reading cilium_* requires cilium without saying so. A metric the table does not claim fails the build.

A LogQL alert names no metrics, so nothing is inferred for it. Its requirements are what it declares, and the chart installs it only where the release can deliver it: a Loki ruler, the alloy-gateway, and a rule store the ruler API can write to.

requires declares what metric names cannot show. An alert whose only metric is up MUST declare what it is about, since up exists for every target. An alert that depends on a label only some deployments add, such as a node label, SHOULD declare the capability that adds it. The effective set, inferred and declared, is recorded per rule in _index.yaml, which is how a reviewer sees it.

A recorded series is not a capability. An alert reading one, such as ext:consensus_up, installs wherever any one of the recording rules that could produce it installs, and _index.yaml records those under the alert’s reads. Reading a recorded series has the rules.

KindCapabilitiesPresent when
Derivedmaterialize, materialize-sql, materialize-operator, kube-state-metrics, cadvisor, node-exporter, loki, alloyThis chart runs the component and collects its metrics
Derivedcloudwatch, cloud-monitoring, azure-monitorThe gateway runs that provider pull, pipeline.metrics.provider.{cloudwatch,gcp,azure}
Explicitsynthetic-uptime, external-uptime, feature-flags, frontegg-auth, memory-limiter, crdb-dedicated, cilium, coredns, cert-manager, kubelet-metrics, swap-nodes, egress-gatewayAn operator lists it in rules.capabilities

Adding a capability means adding it to the Capability enum, to the capability enum in mzmon-query.schema.yaml, and, for a derived one, to mzmon.rules.derivedCapabilities in the chart. A test fails when the first two disagree. Nothing checks the chart’s list against them, so a derived capability’s derivation is added there by hand. A capability SHOULD NOT be added for a single rule; a tag that applies to one rule is a label pretending to be a category.

The default set#

enabledByDefault: true puts an alert in the default set, meaning it installs wherever it applies without an operator selecting it. Everything else installs only when named in rules.selected.

An alert MUST NOT enter the default set until its expression has been evaluated against a live self-managed install and does not fire falsely there. It SHOULD also have a unit test (see below), and it SHOULD have evidence that the condition matters, such as incident history or a Cloud counterpart that pages. Entering the default set commits the alert’s name; renaming it afterwards owes a changelog entry and, after 1.0, a deprecation cycle.

An alert whose normal duration depends on the workload SHOULD carry that duration in for, and its notes SHOULD say so. rules.overrides changes a rule’s for and labels per deployment, and never its expression. Hydration is the standing example: most clusters hydrate in minutes, and a large one can take hours with nothing wrong.

An alert that fires on workloads behind by design MUST stay out of the default set, however useful it is where it applies. A materialized view on a refresh schedule lags by up to its interval between refreshes, so an absolute freshness threshold on every user cluster pages on those clusters permanently.

What gen-rules rejects#

gen-rules reports every problem it finds and writes nothing until there are none.

CheckWhy
The name is kebab-case, and the group snake_caseNames become a committed surface
severity is critical, warning or notice, and component is setThe routing presets route by severity; an unknown one has no class
audience is platform or workloadRoutes match on it; a missing one sends the alert to neither audience’s receiver
for and keepFiringFor are Prometheus durationspromtool would reject the file, and the ruler with it
The query exists and has exactly one expression, PromQL or LogQL and not bothA rule is one expression, and its language decides which ruler evaluates it
A registry file’s alerts are all PromQL or all LogQLEach file installs as one PrometheusRule named after it, and one object cannot be for both rulers
The expression renders and parsesA group with one bad rule is dropped whole. PromQL is parsed here; LogQL is checked for shape here and parsed by make rules-check
A LogQL expression has a stream selector and a rangeWithout a range it is a log query, which parses and which the ruler refuses
No %%{…}, Grafana variable or unknown __mzmon_*__ token remainsA ruler resolves none of them, and the selector matches nothing
Every metric has a known sourceA rule reading a metric nothing produces never fires
A recorded series it reads has a recording rule that can produce what it selectsA selector every recording rule’s labels contradict never matches anything
deploymentMode is not a labelApplicability is requires, not a label nothing reads

The contract for a shipped alert#

These are the rules the audit behind this tooling found broken most often. The first two are checked mechanically; the rest are for the author and the reviewer.

  1. An alert MUST be able to fire on a stock self-managed install, or its capabilities MUST say what it needs.
  2. An alert MUST NOT name a namespace, release, job or container that belongs to one install. It uses the context’s parameters, or a suffix match such as job=~".*/.*materialize-clusterd".
  3. A per-environment alert SHOULD keep namespace in its by (…) clause. The routing tree groups by [alertname, cluster, namespace], and an alert that aggregates it away sends every environment’s condition as one notification.
  4. A ratio MUST guard its denominator. A container with no memory limit reports a limit of 0, and dividing by it gives +Inf, which fires forever.
  5. An alert on a terminated container MUST pair the exit code with a recent restart. kube_pod_container_status_last_terminated_exitcode persists for the life of the pod.
  6. A join between series of different granularity MUST name its labels with on (…). A pod-level and a container-level series never share a full label set.
  7. The name SHOULD describe the condition rather than its grade, so that re-grading does not force a rename.
  8. The summary MUST be actionable without following the runbook link, since an air-gapped install cannot follow it.

Log-derived alerts#

An alert whose query is LogQL reads Materialize’s log lines instead of its metrics. It is authored the same way, in a registry file of its own, and lands in pre-rendered/rules/loki/. It installs as a PrometheusRule labelled mzmon.materialize.cloud/flavor: logql, which the Thanos ruler’s importer leaves out and the alloy-gateway’s loki.rules.kubernetes writes into the Loki ruler through its API. It routes through the same Alertmanager, with the same labels, as a metric alert.

The range is the alert’s duration. Cloud’s clicked-in Loki rules used Grafana’s $__range, which no ruler can parse, and the alerting context refuses it. A log rule SHOULD be written as count_over_time(…[<range>]) > 0 with for: 0s, so it fires on the first evaluation at which a matching line falls inside the range. The range is then how long the alert keeps firing after the last matching line, and the notes SHOULD say so.

A log rule SHOULD match on structure before text. namespace, app, container and level are stream labels and belong in the selector. pod and the fields the pipeline extracts are structured metadata, which a rule filters with | field != "" and MAY group by. The pipeline already classifies a panic: a line starting thread '…' panicked at becomes level CRITICAL, with panic_thread and panic_location as structured metadata, so materialize-panic matches that rather than the word “panic”.

A line filter matches text Materialize does not promise to keep. A reworded message makes the rule silent rather than broken. A log rule that matches message text MUST name the source file that logs it, so a reviewer can check the message still exists. The durable fix is a stable error code in the log line, which the alerting design doc asks of Materialize.

A Loki rule reads one tenant. The rule set installs once per tenant in rules.logTenants, which defaults to the pipeline’s static tenant. Under non-static tenancy the chart cannot enumerate the tenants, and the render warns.

There is no unit test for a LogQL rule. Loki has no counterpart to promtool test rules. A log rule MUST instead be evaluated against a live install before it enters the default set, once as it stands and once with its range widened over a window in which the line is known to have been logged.

Testing an alert#

Unit tests. packages/queries/tests/<registry-file>.test.yaml holds promtool test rules cases, run by make rules-check against the rules the chart renders with every rule selected. A test SHOULD give one input that fires the alert and one that it previously misfired on, and assert on the ALERTS series rather than restating the annotations.

Live evaluation. Before an alert enters the default set, its expression SHOULD be run against a real install, read-only, with the placeholders substituted: once as it stands, to confirm it is quiet, and once over a window in which the condition is known to have happened, to confirm it fires. gcx metrics query does this without cluster access.

Cloud’s rules are prior art, not a source. Materialize Cloud’s rules (infra/prometheus/alerting.py and the Grafana-managed set, where its only log-derived rules live) are worth reading for thresholds and history. An alert MUST NOT be copied from them verbatim: they use Cloud’s namespaces, labels and job names, and some carry customer names that MUST NOT enter this repository.