Common Alerts#

These are the alerting rules the chart bundles, generated from the query registry. Each is installed only where the deployment has the capabilities it requires, and only the default set installs without being selected; see Configuring Alerting.

Rules outside the default set have not all been checked against a self-managed install. Evaluate one against the deployment before relying on it.

infra-alerts#

Alerting rules for the infrastructure Materialize depends on.

Each alert carries its firing expression as an inline query, is graded by a severity label (critical > warning > notice), and is tagged with a component label naming the subsystem it watches.

Most of these watch something a deployment may or may not contain – a CockroachDB metadata store, a Cilium CNI, an egress-gateway pool – and install only where it does. That is inferred from the metrics each alert reads, or declared with requires where the metric names cannot show it.

crdb-disk-usage-critical #

CockroachDB disk usage is above 90% and likely to impact Materialize. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: critical

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(
  crdb_dedicated_capacity_used / crdb_dedicated_capacity
) * 100 > 90

crdb-disk-usage-high #

CockroachDB disk usage is above 70% and may need a capacity increase. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(
  crdb_dedicated_capacity_used / crdb_dedicated_capacity
) * 100 > 70

crdb-disk-usage-elevated #

CockroachDB disk usage is above 30% and worth keeping an eye on. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: notice

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(
  crdb_dedicated_capacity_used / crdb_dedicated_capacity
) * 100 > 30

crdb-cpu-usage-critical #

CockroachDB CPU usage is critically high (>89%) and likely to impact Materialize. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(
  avg_over_time(crdb_dedicated_sys_cpu_combined_percent_normalized[30m]) * 100
) > 89

crdb-cpu-usage-high #

CockroachDB CPU usage is above 85% over 2h and may impact Materialize. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(
  avg_over_time(crdb_dedicated_sys_cpu_combined_percent_normalized[2h]) * 100
) > 85

crdb-query-latency #

CockroachDB p95 SQL service latency has been above 250ms for an extended period. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

histogram_quantile(0.95,
  sum by (le, node) (
    rate(crdb_dedicated_sql_service_latency_bucket[20m])
  )
) > 250 * 1000 * 1000

crdb-lsm-read-amplification-critical #

CockroachDB LSM read amplification is critically high (>150), indicating severe I/O overload. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_rocksdb_read_amplification) > 150

crdb-lsm-read-amplification-high #

CockroachDB LSM read amplification is elevated (>50), a sign writes may be outpacing compaction. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: notice

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_rocksdb_read_amplification) > 50

crdb-ranges-unavailable #

CockroachDB has unavailable ranges, which may indicate node failures or replication issues. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_ranges_unavailable) > 0

crdb-ranges-underreplicated #

CockroachDB has under-replicated ranges, which may indicate node failures or replication issues. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_ranges_underreplicated) > 0

crdb-sql-memory-rapid-growth #

CockroachDB SQL memory is growing faster than 8 MB/s, a sign of a runaway query heading toward OOM. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max by (node_id) (
  deriv(crdb_dedicated_sql_mem_distsql_current[5m])
) > 8 * 1000 * 1000

crdb-sql-memory-pressure-high #

CockroachDB distsql memory exceeds 18% of node RAM, indicating high user-query memory pressure. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: notice

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max by (node_id) (
  crdb_dedicated_sql_mem_distsql_current
)
/ on (node_id) max by (node_id) (
  crdb_dedicated_sys_totalmem
)
> 0.18

crdb-write-intent-accumulation-critical #

CockroachDB write-intent count has exceeded 10M for 10m, indicating large transactions holding locks. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: warning

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_intentcount) > 10 * 1000 * 1000

crdb-write-intent-accumulation-high #

CockroachDB write-intent count has exceeded 10M, a sign of large transactions accumulating locks. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: notice

Installed: only when named in rules.selected. Requires: crdb-dedicated.

max(crdb_dedicated_intentcount) > 10 * 1000 * 1000

crdb-backup-missing #

A CockroachDB backup may be missing — the last completed backup is more than 80 minutes old. Labels:
  • audience: platform
  • component: cockroachdb
  • severity: notice

Installed: only when named in rules.selected. Requires: crdb-dedicated.

(
  time() - max(max_over_time(crdb_dedicated_schedules_backup_last_completed_time[60m]))
) / 60 > 80

cilium-bpf-map-pressure #

A Cilium BPF map is filling up, which will cause networking issues if left unaddressed. Labels:
  • audience: platform
  • component: cilium
  • severity: warning

Installed: only when named in rules.selected. Requires: cilium.

max by (map_name, instance) (
  cilium_bpf_map_pressure
) > 0.7

cilium-drop-rate-elevated #

Cilium is dropping more packets than expected, which can indicate networking issues. Labels:
  • audience: platform
  • component: cilium
  • severity: notice

Installed: only when named in rules.selected. Requires: cilium.

max by (direction, reason, pod) (
  rate(cilium_drop_count_total[5m:1m])
) > 10

coredns-slow-queries #

CoreDNS p99 request latency is above 500ms, which can cause or accompany broader outages. Labels:
  • audience: platform
  • component: coredns
  • severity: notice

Installed: only when named in rules.selected. Requires: coredns.

histogram_quantile(0.99,
  sum by (le, service) (
    rate(coredns_dns_request_duration_seconds_bucket{zone="."}[5m])
  )
) > 0.5

node-unreachable #

A Kubernetes node is marked unreachable, which may indicate a node or network problem. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

kube_node_spec_taint{key="node.kubernetes.io/unreachable"}

k8s-horizontalpodautoscaler-replicas-critical #

A HorizontalPodAutoscaler is above 90% of its maximum replicas — nearly out of headroom to scale. Labels:
  • audience: platform
  • component: kubernetes
  • severity: critical

Installed: only when named in rules.selected. Requires: kube-state-metrics.

(
  kube_horizontalpodautoscaler_status_current_replicas
  / kube_horizontalpodautoscaler_spec_max_replicas
) > 0.9

k8s-horizontalpodautoscaler-replicas-high #

A HorizontalPodAutoscaler is above 70% of its maximum replicas. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

(
  kube_horizontalpodautoscaler_status_current_replicas
  / kube_horizontalpodautoscaler_spec_max_replicas
) > 0.7

k8s-container-disk-usage #

A container filesystem is above 70% usage and should be cleaned up or resized. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: by default, wherever it applies. Requires: cadvisor.

100 * (
  max by (node) (container_fs_usage_bytes)
  / max by (node) (container_fs_limit_bytes)
) > 70

k8s-node-disk-pressure #

A Kubernetes node is under disk pressure, which can cause pods to fail scheduling or be evicted. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

max by (node) (
  kube_node_status_condition{condition="DiskPressure", status="true"}
) > 0

k8s-volume-usage #

A Kubernetes persistent volume is above 70% usage and should be cleaned up or resized. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kubelet-metrics.

100 * (
  max by (namespace, persistentvolumeclaim) (
    kubelet_volume_stats_used_bytes{persistentvolumeclaim!~".*cluster-s2-.*|.*cluster-u.*"}
  )
  / max by (namespace, persistentvolumeclaim) (
    kubelet_volume_stats_capacity_bytes{persistentvolumeclaim!~".*cluster-s2-.*|.*cluster-u.*"}
  )
) > 70

alloy-log-drops #

Alloy is dropping log entries — logs are being lost right now. Labels:
  • audience: platform
  • component: loki
  • severity: warning

Installed: by default, wherever it applies. Requires: alloy.

sum by (namespace, job, reason) (
  rate(loki_write_dropped_entries_total[2m])
) > 0

loki-push-err-high #

Loki’s push endpoint is returning 10%+ write errors, so logs are being rejected. Labels:
  • audience: platform
  • component: loki
  • severity: warning

Installed: by default, wherever it applies. Requires: loki.

100 * (
  sum by (namespace, job) (
    rate(loki_request_duration_seconds_count{status_code=~"4.*|5.*", route=~".*push.*"}[2m])
  )
  / sum by (namespace, job) (
    rate(loki_request_duration_seconds_count{route=~".*push.*"}[2m])
  )
) > 10

loki-panics #

A Loki component has panicked. Labels:
  • audience: platform
  • component: loki
  • severity: notice

Installed: by default, wherever it applies. Requires: loki.

sum by (namespace, job) (
  increase(loki_panic_total[10m])
) > 0

loki-req-duration-high #

Loki p95 request duration is above 1s (excluding tail/long-poll routes). Labels:
  • audience: platform
  • component: loki
  • severity: notice

Installed: only when named in rules.selected. Requires: loki.

histogram_quantile(0.95,
  sum by (le, job) (
    rate(loki_request_duration_seconds_bucket{route!~"(?i).*tail.*|/schedulerpb.SchedulerForQuerier/QuerierLoop"}[5m])
  )
) > 1

loki-req-err-high #

Loki read requests are returning 10%+ 5xx errors. Labels:
  • audience: platform
  • component: loki
  • severity: notice

Installed: by default, wherever it applies. Requires: loki.

100 * (
  sum by (namespace, job, route) (
    rate(loki_request_duration_seconds_count{status_code=~"5.*"}[2m])
  )
  / sum by (namespace, job, route) (
    rate(loki_request_duration_seconds_count[2m])
  )
) > 10

loki-writer-err-high #

Loki write requests are erroring for 10%+ of requests. Labels:
  • audience: platform
  • component: loki
  • severity: notice

Installed: only when named in rules.selected. Requires: loki.

100 * (
  sum by (namespace, job, route) (
    rate(loki_request_duration_seconds_count{status_code="error"}[2m])
  )
  / sum by (namespace, job, route) (
    rate(loki_request_duration_seconds_count[2m])
  )
) > 10

clusterd-metrics-missing #

All metrics for a Materialize pod have been missing for 60m — the scrape target is likely down. Labels:
  • audience: platform
  • component: monitoring
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

max by (pod, namespace) (
  up{job=~".*/.*materialize-(environmentd|clusterd)", cluster_environmentd_materialize_cloud_cluster_id!="s5"}
) == 0

critical-metrics-missing #

More than 30% of a critical scrape job’s targets have been down for 30m. Labels:
  • audience: platform
  • component: monitoring
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor.

(
  count by (job, app) (
    up{job=~"cadvisor|kube-state-metrics|node-exporter|alloy-agent|alloy-gateway"} == 0
  )
  / count by (job, app) (
    up{job=~"cadvisor|kube-state-metrics|node-exporter|alloy-agent|alloy-gateway"}
  )
) * 100 > 30

logging-collection-down #

Loki has received no logs for 15m — log collection is down. Labels:
  • audience: platform
  • component: monitoring
  • severity: warning

Installed: by default, wherever it applies. Requires: loki.

(sum(rate(loki_distributor_bytes_received_total[5m])) or vector(0)) == 0

k8s-deployment-unavailable-critical #

A core control-plane deployment is unavailable. Labels:
  • audience: platform
  • component: kubernetes
  • severity: critical

Installed: only when named in rules.selected. Requires: kube-state-metrics.

group by (deployment, namespace) (
  kube_deployment_status_condition{condition=~"Available", status=~"true", deployment=~"${infraCoreWorkloadList}", ${excludeEnvironmentFilter}} == 0
)

k8s-deployment-unavailable-warning #

An important supporting deployment is unavailable. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

group by (deployment, namespace) (
  kube_deployment_status_condition{condition=~"Available", status=~"true", deployment=~"${infraImportantWorkloadList}", ${excludeEnvironmentFilter}} == 0
)

k8s-deployment-unavailable-notice #

A non-essential deployment is unavailable. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics.

group by (deployment, namespace) (
  kube_deployment_status_condition{condition=~"Available", status=~"true", deployment=~"${infraNonessentialWorkloadList}", ${excludeEnvironmentFilter}} == 0
)

k8s-daemonset-saturating-cpu #

Daemonsets are requesting nearly all the CPU reserved for them, squeezing clusterd headroom. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

(
  1.77 - sum(
    max by (container) (
      kube_pod_container_resource_requests{container=~"${infraDaemonsetWorkloadList}", unit="core", resource="cpu"}
    )
  )
) < 0.01

k8s-daemonset-saturating-mem #

Daemonsets are requesting nearly all the memory reserved for them, squeezing clusterd headroom. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

(
  10730942464 - sum(
    max by (container) (
      kube_pod_container_resource_requests{container=~"${infraDaemonsetWorkloadList}", unit="byte", resource="memory"}
    )
  )
) < 5243000

k8s-daemonset-high-cpu #

Daemonsets are approaching the CPU budget reserved for them. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics.

(
  1.77 - sum(
    max by (container) (
      kube_pod_container_resource_requests{container=~"${infraDaemonsetWorkloadList}", unit="core", resource="cpu"}
    )
  )
) < 0.1

container-file-descriptors-critical #

A core container’s open file descriptors are above 70% of its limit and it may be killed. Labels:
  • audience: platform
  • component: kubernetes
  • severity: critical

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (pod, namespace, container) (
    container_file_descriptors{container=~"${infraCoreWorkloadList}"}
  )
  / min by (pod, namespace, container) (
    container_ulimits_soft{ulimit="max_open_files"}
  )
) > 70

container-file-descriptors-warning #

A non-core container’s open file descriptors are above 70% of its limit and it may be killed. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (pod, namespace, container) (
    container_file_descriptors{container!~"${infraCoreWorkloadList}"}
  )
  / min by (pod, namespace, container) (
    container_ulimits_soft{ulimit="max_open_files"}
  )
) > 70

container-file-descriptors-elevated #

A container’s open file descriptors are above 20% of its limit. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (pod, namespace, container) (container_file_descriptors)
  / min by (pod, namespace, container) (
    container_ulimits_soft{ulimit="max_open_files"}
  )
) > 20

infra-pod-high-cpu-ratio #

A vector pod has used more than its full CPU request for 6h and may be throttled. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics, cadvisor.

(
  sum by (pod, namespace, container) (
    rate(container_cpu_usage_seconds_total{container!="", pod=~"vector-.*"}[5m])
  )
  / max by (pod, namespace, container) (
    kube_pod_container_resource_requests{resource="cpu", container!="", pod=~"vector-.*"}
  )
) > 1

infra-pods-high-cpu-ratio #

An infra pod is using more than its full CPU request and may be throttled. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics, cadvisor.

(
  sum by (pod, namespace, container) (
    rate(container_cpu_usage_seconds_total{${excludeMzDeploymentNamespaceFilter}, container!="", pod!~"parca-agent-.*|vector-.*|cilium-egress-.*"}[5m])
  )
  / max by (pod, namespace, container) (
    kube_pod_container_resource_requests{${excludeMzDeploymentNamespaceFilter}, resource="cpu", container!="", pod!~"parca-agent-.*|vector-.*|cilium-egress-.*"}
  )
) > 1

infra-memory-high #

An important infra container is above 80% memory usage. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (namespace, pod, container) (
    container_memory_working_set_bytes{container=~"${infraCoreWorkloadList}"}
  )
  / sum by (namespace, pod, container) (
    container_spec_memory_limit_bytes{container=~"${infraCoreWorkloadList}"} != 0
  )
) > 80

infra-memory-elevated #

A supporting infra container is above 80% memory usage. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (namespace, pod, container) (
    container_memory_working_set_bytes{container=~"${infraImportantWorkloadList}"}
  )
  / sum by (namespace, pod, container) (
    container_spec_memory_limit_bytes{container=~"${infraImportantWorkloadList}"} != 0
  )
) > 80

infra-oomkill-core-systems #

A core infra container has been OOMKilled. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

sum by (container, namespace, pod) (
  kube_pod_container_status_restarts_total{container=~"${infraCoreWorkloadList}"}
  - kube_pod_container_status_restarts_total{container=~"${infraCoreWorkloadList}"} offset 3h > 1
  and ignoring(reason)
  kube_pod_container_status_last_terminated_reason{reason="OOMKilled", container=~"${infraCoreWorkloadList}"}
) > 0

infra-oomkill-important-systems #

An important infra container has been OOMKilled. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

sum by (container, namespace, pod) (
  kube_pod_container_status_restarts_total{container=~"${infraImportantWorkloadList}"}
  - kube_pod_container_status_restarts_total{container=~"${infraImportantWorkloadList}"} offset 3h > 1
  and ignoring(reason)
  kube_pod_container_status_last_terminated_reason{reason="OOMKilled", container=~"${infraImportantWorkloadList}"}
) > 0

infra-oomkill-nonessential-systems #

A non-essential infra container has been OOMKilled. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics.

sum by (container, namespace, pod) (
  kube_pod_container_status_restarts_total{container=~"${infraNonessentialWorkloadList}"}
  - kube_pod_container_status_restarts_total{container=~"${infraNonessentialWorkloadList}"} offset 3h > 1
  and ignoring(reason)
  kube_pod_container_status_last_terminated_reason{reason="OOMKilled", container=~"${infraNonessentialWorkloadList}"}
) > 0

k8s-infra-pod-pending-too-long #

An infra pod has been Pending for 20m — the cluster may be unhealthy or out of capacity. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

max by (namespace, pod) (
  kube_pod_status_phase{${excludeMzDeploymentNamespaceFilter}, phase="Pending"}
) > 0

pod-restart-rate-high #

An important infra container is restarting frequently, which may indicate a crash loop. Labels:
  • audience: platform
  • component: kubernetes
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

avg by (container, namespace) (
  rate(kube_pod_container_status_restarts_total{${excludeMzDeploymentNamespaceFilter}, container!~"${infraNonessentialWorkloadList}"}[10m])
) * 100 > 0

pod-restart-rate-high-nonessential #

A non-essential infra container is restarting frequently, which may indicate a crash loop. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics.

avg by (container, namespace) (
  rate(kube_pod_container_status_restarts_total{container=~"${infraNonessentialWorkloadList}"}[10m])
) * 100 > 0

pods-stuck-in-waiting #

A pod has been stuck in Waiting for over 10m, which can indicate a scheduling or resource problem. Labels:
  • audience: platform
  • component: kubernetes
  • severity: notice

Installed: by default, wherever it applies. Requires: kube-state-metrics.

sum by (namespace, pod) (
  kube_pod_container_status_waiting == 1
  and on (namespace, pod) (time() - kube_pod_start_time > 10 * 60)
) > 0

egress-traffic-missing-metrics #

Egress-gateway node throughput metrics are missing, which may indicate an egress-gateway problem. Labels:
  • audience: platform
  • component: egress-gateway
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor, node-exporter, egress-gateway.

absent(
  sum by (node) (
    rate(node_network_receive_bytes_total{device="eth0"}[2m])
  )
  * on (node) group_left (workload) sum by (node, workload) (
    label_replace(container_network_receive_bytes_total{workload="materialize-egress"}, "node", "$1", "instance", "(.+)")
  ) ^ 0
)

excessive-egress-traffic #

An egress-gateway node has very high traffic, which may indicate traffic is not routing through the internet gateway. Labels:
  • audience: platform
  • component: egress-gateway
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor, node-exporter, egress-gateway.

sum by (node) (
  rate(node_network_receive_bytes_total{device="eth0"}[2m])
)
* on (node) group_left (workload) sum by (node, workload) (
  label_replace(container_network_receive_bytes_total{workload="materialize-egress"}, "node", "$1", "instance", "(.+)")
) ^ 0
> 5.5 * 1000 * 1000 * 1000 / 8

high-egress-traffic #

An egress-gateway node is above 90% of its allowed traffic and may soon exceed capacity. Labels:
  • audience: platform
  • component: egress-gateway
  • severity: notice

Installed: only when named in rules.selected. Requires: cadvisor, node-exporter, egress-gateway.

sum by (node) (
  rate(node_network_receive_bytes_total{device="eth0"}[2m])
)
* on (node) group_left (workload) sum by (node, workload) (
  label_replace(container_network_receive_bytes_total{workload="materialize-egress"}, "node", "$1", "instance", "(.+)")
) ^ 0
> 4.5 * 1000 * 1000 * 1000 / 8

low-egress-traffic #

An egress-gateway node has unusually low traffic, which may indicate the gateway is unhealthy. Labels:
  • audience: platform
  • component: egress-gateway
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor, node-exporter, egress-gateway.

sum by (node) (
  rate(node_network_receive_bytes_total{device="eth0"}[2m])
)
* on (node) group_left (workload) sum by (node, workload) (
  label_replace(container_network_receive_bytes_total{workload="materialize-egress"}, "node", "$1", "instance", "(.+)")
) ^ 0
< 100

no-egress-traffic #

An egress-gateway node has had no traffic for 5m, which may indicate external traffic is blocked. Labels:
  • audience: platform
  • component: egress-gateway
  • severity: warning

Installed: only when named in rules.selected. Requires: node-exporter, kubelet-metrics, egress-gateway.

(
  min by (node) (
    rate(node_network_receive_bytes_total{device="eth0"}[5m:1m])
    or rate(node_network_transmit_bytes_total{device="eth0"}[5m:1m])
  )
  and on (node) kubelet_node_name{workload="materialize-egress"}
) == 0

infra-log-alerts#

Log-derived alerting rules for the monitoring stack.

Each alert matches a line a collector logs when it is failing in a way its own metrics cannot show, because the metrics travel the path that is failing. They are LogQL, evaluated by the Loki ruler against the logs the pipeline collects, and route through the same Alertmanager as the metric alerts.

A match is on the log line’s text, and that text is not a contract. An upstream release can reword a message, and a reworded message makes the rule silent rather than broken. Each alert names the source of the line it matches, so a contributor can check it still exists.

The label contract these read is documented under Logs and Events: namespace, app, container and level are stream labels, and pod is structured metadata, which a rule may still group by.

alloy-gateway-refusing-scrapes #

An alloy-gateway pod is over its memory limiter’s soft limit and is discarding everything it scrapes, so its share of scrape targets is missing from the metrics store. Restart the pod to recover; raise the gateway’s memory limit, and any explicit GOMEMLIMIT, to stop it recurring. Labels:
  • audience: platform
  • component: monitoring
  • severity: warning

Installed: by default, wherever it applies.

Evaluated by: the Loki ruler, as LogQL.

sum by (namespace, pod) (
  count_over_time(
    {app="alloy-gateway", container="alloy", level="ERROR"}
      |= "data refused due to high memory usage"
    [5m]
  )
) > 0

materialize-alerts#

Alerting rules for Materialize.

Each alert carries its firing expression as an inline query, is graded by a severity label (critical > warning > notice), and is tagged with a component label. Per-environment alerts group by namespace and attach the environment name via the mzEnvironmentName template function; the %%{excludeEnvironmentFilter} fragment lets a deployment exclude namespaces from alerting (rules.namespaces.exclude).

Alerts that depend on something a deployment may not contain are gated by capabilities, inferred from the metrics they read or declared with requires, rather than by who operates the deployment. See Authoring Alerts in the internal docs for how these render into rules.

env-uptime-sla #

environmentd is not accepting basic connections and may be unreachable. Labels:
  • audience: platform
  • component: environmentd
  • severity: critical

Installed: only when named in rules.selected. Requires: kube-state-metrics, synthetic-uptime.

avg by (namespace, pod) (
  ${mzSqlPrefix}can_connect{${excludeEnvironmentFilter}}
) <= 0.1
and on (namespace, pod) (
  time() - kube_pod_start_time > 5 * 60
)

env-uptime-slo #

environmentd is not accepting basic connections and may be unreachable. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics, synthetic-uptime.

avg by (namespace, pod) (
  ${mzSqlPrefix}can_connect{${excludeEnvironmentFilter}}
) <= 0.1
and on (namespace, pod) (
  time() - kube_pod_start_time > 5 * 60
)

envd-simplest-query #

environmentd is not responding to a SELECT 1 and may be unhealthy. Labels:
  • audience: platform
  • component: environmentd
  • severity: critical

Installed: only when named in rules.selected. Requires: synthetic-uptime.

avg by (namespace) (
  ${mzSqlPrefix}envd_up{${excludeEnvironmentFilter}}
) <= 0.1

env-query-views-critical #

environmentd is not answering a simple query that reads from object storage. Labels:
  • audience: platform
  • component: environmentd
  • severity: critical

Installed: only when named in rules.selected. Requires: synthetic-uptime.

avg by (namespace) (
  ${mzSqlPrefix}views_query_successful{${excludeEnvironmentFilter}}
) <= 0.1

env-query-views-warning #

environmentd is not answering a simple query that reads from object storage. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: only when named in rules.selected. Requires: synthetic-uptime.

avg by (namespace) (
  ${mzSqlPrefix}views_query_successful{${excludeEnvironmentFilter}}
) <= 0.1

clusterd-not-receiving-commands #

A clusterd has not received commands from environmentd for 5m — the replica may be stalled or disconnected. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize, kube-state-metrics.

max by (namespace, materialize_cloud_organization_name, cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, pod) (
  rate(mz_cluster_server_last_command_received{server_name="compute", pod=~".*0", cluster_environmentd_materialize_cloud_cluster_id!~"s.*[1345]", ${excludeEnvironmentFilter}}[5m])
) == 0
and on (namespace, pod) (
  time() - kube_pod_start_time > 5 * 60
)
and on (namespace, pod) (
  max by (namespace, pod) (increase(kube_pod_container_status_restarts_total{container="clusterd"}[6h:1m])) == 0
)

storage-collection-finalization-stuck #

Storage shards have been stuck finalizing for 8h. Labels:
  • audience: platform
  • component: storage
  • severity: notice

Installed: only when named in rules.selected. Requires: materialize.

max by (namespace, materialize_cloud_organization_name) (
  mz_shard_finalization_outstanding{${excludeEnvironmentFilter}}
) > 0

environmentd-down #

environmentd has not answered its scrape for 5m, so the environment is likely down. Labels:
  • audience: platform
  • component: environmentd
  • severity: critical

Installed: by default, wherever it applies. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, pod) (
  up{job=~".*/.*materialize-environmentd", ${excludeEnvironmentFilter}}
) == 0

environmentd-not-scraped #

No environmentd is being scraped at all. Labels:
  • audience: platform
  • component: monitoring
  • severity: critical

Installed: by default, wherever it applies. Requires: materialize.

absent(up{job=~".*/.*materialize-environmentd"})

auth-errors #

An elevated rate of unexpected authentication errors. Labels:
  • audience: platform
  • component: auth
  • severity: warning

Installed: only when named in rules.selected. Requires: frontegg-auth.

avg by (namespace, status) (
  rate(mz_auth_request_count{status!~"(2|401).*"}[30m])
) > 0.01

auth-refresh-failures #

An elevated rate of failed auth-token refreshes. Labels:
  • audience: platform
  • component: auth
  • severity: warning

Installed: only when named in rules.selected. Requires: frontegg-auth.

avg(
  rate(mz_auth_request_count{status=~"401.*", path="refresh_token"}[15m])
) > 0.002

console-errors #

The web console has returned an error for 2% or more of commands for 30m, across more than one environment. Labels:
  • audience: platform
  • component: console
  • severity: warning

Installed: only when named in rules.selected. Requires: materialize.

(
  sum by (application_name) (
    rate(mz_adapter_commands{application_name="web_console", status="error", ${excludeEnvironmentFilter}}[5m])
  )
  / sum by (application_name) (
    rate(mz_adapter_commands{application_name="web_console", ${excludeEnvironmentFilter}}[5m])
  )
) * 100 > 2
and
count by (application_name) (
  (
    sum by (application_name, namespace) (
      rate(mz_adapter_commands{application_name="web_console", status="error", ${excludeEnvironmentFilter}}[5m])
    )
    / sum by (application_name, namespace) (
      rate(mz_adapter_commands{application_name="web_console", ${excludeEnvironmentFilter}}[5m])
    )
  ) * 100 > 2
) > 1

consensus-unreachable #

Materialize cannot reach its metadata database, so nothing it stores can change and every write is stalled. Labels:
  • audience: platform
  • component: consensus
  • severity: critical

Installed: by default, wherever it applies. Reads: ext:consensus_up, recorded by any of ext_consensus_persist, ext_consensus_cloudsql, ext_consensus_azure_postgres.

min by (flavor, namespace, resource) (ext:consensus_up) == 0

consensus-failures #

Calls to the metadata database are failing with indeterminate errors, and Materialize is retrying its writes. Labels:
  • audience: platform
  • component: consensus
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

sum by (namespace) (rate(mz_persist_consensus_failures[5m])) > 0.1

blob-failures #

Calls to the object store holding Materialize’s data are failing, by operation. Labels:
  • audience: platform
  • component: object-storage
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

sum by (namespace, op) (rate(mz_persist_external_failed_count{op=~"blob_.*"}[5m])) > 0.1

persist-failures #

Failures inside Persist that should be rare are happening frequently. Labels:
  • audience: platform
  • component: persist
  • severity: notice

Installed: only when named in rules.selected. Requires: materialize.

sum by (namespace, metric) (rate(label_replace(mz_persist_state_update_state_slow_path, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_lease_timeout_read, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 3
or sum by (namespace, metric) (rate(label_replace(mz_persist_compaction_noop, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 3
or sum by (namespace, metric) (rate(label_replace(mz_persist_compaction_failed, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_external_blob_delete_noop_count, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_compaction_dropped, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_pushdown_parts_mismatched_stats_count, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_columnar_op_count{op="validation", result="invalid"}, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_schema_cache_fetch_state_count, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1
or sum by (namespace, metric) (rate(label_replace(mz_persist_shard_unconsolidated_snapshot, "metric", "$1", "__name__", "(.*)")[1m:15s])) > 1

envd-terminated #

An environmentd was unexpectedly terminated — this should not happen. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

kube_pod_container_status_last_terminated_exitcode{container="environmentd"} != 166
and on (namespace, pod)
max by (namespace, pod) (increase(kube_pod_container_status_restarts_total{container="environmentd"}[10m:1m])) > 0

environmentd-high-cpu #

An environmentd has been above 80% CPU usage for 90m. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: only when named in rules.selected. Requires: cadvisor.

100 * sum by (namespace, pod) (
  rate(container_cpu_usage_seconds_total{pod=~".*environmentd.+", container!=""}[30m])
)
/ sum by (namespace, pod) (
  container_spec_cpu_quota{pod=~".*environmentd.+", container!=""}
  / container_spec_cpu_period{pod=~".*environmentd.+", container!=""}
) > 80

environmentd-high-memory #

An environmentd has been above 80% memory usage for 30m. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: by default, wherever it applies. Requires: cadvisor.

100 * (
  sum by (namespace, pod, container) (
    container_memory_working_set_bytes{pod=~".*environmentd.+", container!=""}
  )
  / sum by (namespace, pod, container) (
    container_spec_memory_limit_bytes{pod=~".*environmentd.+", container!=""} != 0
  )
) > 80

environmentd-cpu-throttled #

An environmentd container is being CPU throttled more than 50% of the time. Labels:
  • audience: platform
  • component: environmentd
  • severity: notice

Installed: only when named in rules.selected. Requires: cadvisor.

100 * avg by (instance, container, namespace) (
  rate(container_cpu_cfs_throttled_periods_total{container=~"environmentd"}[30m])
  / rate(container_cpu_cfs_periods_total{container=~"environmentd"}[30m])
) > 50

environmentd-memory-elevated #

An environmentd is above 80% memory usage (lower-severity companion to environmentd-high-memory). Labels:
  • audience: platform
  • component: environmentd
  • severity: notice

Installed: only when named in rules.selected. Requires: cadvisor.

100 * (
  sum by (namespace, pod, container) (
    container_memory_working_set_bytes{container!="POD", container!="", container=~"environmentd"}
  )
  / sum by (namespace, pod, container) (
    container_spec_memory_limit_bytes{container!="POD", container!="", container=~"environmentd"} != 0
  )
) > 80

clusterd-error-kill #

A clusterd was terminated with an unexpected exit code — this should not happen. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

max by (namespace, pod) (
  increase(kube_pod_container_status_restarts_total{${mzEnvironmentNamespaceFilter}, container="clusterd"}[1h])
) > 0
and max by (namespace, pod) (kube_pod_container_status_last_terminated_exitcode) != 137
and max by (namespace, pod) (kube_pod_container_status_last_terminated_exitcode) != 135
and max by (namespace, pod) (kube_pod_container_status_last_terminated_exitcode) != 166
and max by (namespace, pod) (kube_pod_container_status_last_terminated_exitcode) != 167

system-cluster-terminated #

A system cluster was unexpectedly terminated — this should not happen. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

kube_pod_container_status_last_terminated_exitcode{pod=~".*cluster-s.*"} != 166
and on (namespace, pod)
max by (namespace, pod) (increase(kube_pod_container_status_restarts_total{pod=~".*cluster-s.*"}[10m:1m])) > 0

system-cluster-high-memory #

A system cluster is above 80% memory usage — this should not happen for system clusters. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: only when named in rules.selected. Requires: memory-limiter.

100 * (
  sum by (namespace, pod) (
    mz_memory_limiter_memory_usage_bytes{pod=~".*cluster-s.+"}
  )
  / sum by (namespace, pod) (
    mz_memory_limiter_memory_limit_bytes{pod=~".*cluster-s.+"} != 0
  )
) > 80

clusterd-expiration-7d #

A cluster replica will expire in less than a week. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

mz_dataflow_replica_expiration_remaining_seconds{pod=~".*cluster.+", ${excludeEnvironmentFilter}} > 0
< 60 * 60 * 24 * 7

swap-cluster-oom #

A swap-enabled cluster was OOMKilled while below 80% swap usage. Labels:
  • audience: platform
  • component: clusterd
  • severity: notice

Installed: only when named in rules.selected. Requires: kube-state-metrics, cadvisor, swap-nodes.

increase(
  (
    max by (namespace, pod) (
      kube_pod_container_status_last_terminated_reason{reason="OOMKilled", ${mzEnvironmentNamespaceFilter}}
    )
    or on (namespace, pod) count by (namespace, pod) (
      container_last_seen{materialize_cloud_swap="true", ${excludeEnvironmentFilter}}
    ) * 0
  )[10m:1m]
) > 0
and on (namespace, pod) (
  max_over_time(
    (
      max by (namespace, pod) (container_memory_swap)
      / max by (namespace, pod) (container_spec_memory_swap_limit_bytes)
    )[10m:1m]
  ) < 0.8
)
and on (namespace, pod) (
  increase(
    max by (namespace, pod) (kube_pod_container_status_restarts_total)[10m:1m]
  ) > 0
)

system-cluster-falling-behind #

A system cluster has averaged over 60s of lag for 15m. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, instance_id) (
  avg_over_time(
    (
      min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
        mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"s.*", ${excludeEnvironmentFilter}}
      ) < 1e18
    )[15m:1m]
  )
) > 60

system-cluster-stale #

A system cluster has averaged over 10m of lag for 15m. Labels:
  • audience: platform
  • component: clusterd
  • severity: critical

Installed: by default, wherever it applies. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, instance_id) (
  avg_over_time(
    (
      min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
        mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"s.*", ${excludeEnvironmentFilter}}
      ) < 1e18
    )[15m:1m]
  )
) > 600

system-cluster-hydration-stuck #

A system cluster has had collections with no hydrated replica for 15m. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

count by (namespace, materialize_cloud_organization_name, instance_id) (
  min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
    mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"s.*", ${excludeEnvironmentFilter}}
  ) >= 1e18
) > 0

system-cluster-memory-near-limit #

A system cluster replica’s heap has reached 90% of the point where the kernel OOM-kills it. Labels:
  • audience: platform
  • component: clusterd
  • severity: critical

Installed: by default, wherever it applies. Requires: materialize.

100 * (
  max by (namespace, materialize_cloud_organization_name, cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, pod) (
    max_over_time(mz_metrics_resource_usage{metric="heap", cluster_environmentd_materialize_cloud_cluster_id=~"s.*", ${excludeEnvironmentFilter}}[15m])
  )
  / on (namespace, pod) group_left ()
  (
    max by (namespace, pod) (mz_metrics_resource_usage{metric="memory_max"})
    + on (namespace, pod)
    max by (namespace, pod) (mz_metrics_resource_usage{metric="swap_max"})
  )
) > 90

environment-pod-pending-critical #

An environment pod has been Pending for 15m — the cluster may be unhealthy or out of capacity. Labels:
  • audience: platform
  • component: environmentd
  • severity: critical

Installed: by default, wherever it applies. Requires: kube-state-metrics.

min by (namespace, pod_base) (
  label_replace(
    kube_pod_status_phase{${mzEnvironmentNamespaceFilter}, phase="Pending", ${excludeEnvironmentFilter}},
    "pod_base", "$1", "pod", "(.*?)(?:-gen-[0-9]+-[0-9]+|([0-9]+)-[0-9]+)?"
  )
) > 0

environment-pod-pending #

An environment pod has been Pending for 15m — the cluster may be unhealthy or out of capacity. Labels:
  • audience: platform
  • component: environmentd
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

max by (namespace, pod) (
  kube_pod_status_phase{${mzEnvironmentNamespaceFilter}, phase="Pending", ${excludeEnvironmentFilter}}
) > 0

cluster-replica-not-ready #

A cluster replica’s pod has been ready for less than half of the last 10m. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

max by (namespace, pod, cluster_environmentd_materialize_cloud_cluster_id) (
  label_replace(
    avg_over_time(kube_pod_status_ready{condition="true", pod=~".*-cluster-[a-z0-9]+-replica-.*", ${mzEnvironmentNamespaceFilter}, ${excludeEnvironmentFilter}}[10m]),
    "cluster_environmentd_materialize_cloud_cluster_id", "$1", "pod", ".*-cluster-([a-z0-9]+)-replica-.*"
  )
) < 0.5
and on (namespace, pod) (time() - kube_pod_created > 20 * 60)

certificate-not-ready #

A certificate has not become ready in 20m, which can prevent an environment from coming up. Labels:
  • audience: platform
  • component: environmentd
  • severity: notice

Installed: only when named in rules.selected. Requires: cert-manager.

max by (name, namespace, condition) (
  certmanager_certificate_ready_status{condition!="True", ${excludeEnvironmentFilter}} == 1
)

console-query-latency #

Web-console query p95 latency has exceeded 10s for 15m. Labels:
  • audience: platform
  • component: console
  • severity: warning

Installed: only when named in rules.selected. Requires: materialize.

sum by (namespace) (
  histogram_quantile(0.95,
    sum by (le, namespace) (
      rate(mz_time_to_first_row_seconds_bucket{instance_id=~"s2", application_name="web_console", ${excludeEnvironmentFilter}}[2m])
    )
  )
) > 10

new-clusterd-restarts #

A previously healthy clusterd is restarting during a release. Labels:
  • audience: platform
  • component: clusterd
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics.

label_replace(
  max by (namespace, pod) (kube_pod_container_status_restarts_total{container="clusterd"}) > 0
  and on (namespace, pod) (
    min by (namespace, pod) (time() - kube_pod_created)
  ) < 60 * 60 * 24,
  "pod_base", "$1", "pod", ".*-(cluster-.*)-gen-([0-9]+)-[0-9]+$"
)
and on (namespace, pod_base) label_replace(
  max by (namespace, pod) (increase(kube_pod_container_status_restarts_total{container="clusterd"}[12h:1m])) == 0
  and on (namespace, pod) (
    min by (namespace, pod) (time() - kube_pod_created)
  ) > 60 * 60 * 24,
  "pod_base", "$1", "pod", ".*-(cluster-.*)-gen-([0-9]+)-[0-9]+$"
)

external-env-uptime #

environmentd is unreachable from outside the network by the external uptime checker. Labels:
  • audience: platform
  • component: external-uptime
  • severity: critical

Installed: only when named in rules.selected. Requires: external-uptime.

avg by (namespace) (
  mz_external_envd_up
) < 1

external-env-uptime-failed #

New external connections to environmentd are failing, per the external uptime checker. Labels:
  • audience: platform
  • component: external-uptime
  • severity: warning

Installed: only when named in rules.selected. Requires: external-uptime.

sum by (connection_type, namespace) (
  rate(mz_external_calls_count{status="failed"}[2m])
) > 0

external-uptime-checker-not-calling #

The external uptime checker has stopped making calls — the checker itself may be down. Labels:
  • audience: platform
  • component: external-uptime
  • severity: warning

Installed: only when named in rules.selected. Requires: external-uptime.

sum by (connection_type, namespace) (
  increase(mz_external_calls_count{status="attempted"}[2m])
) == 0

launchdarkly-stale-sse #

The last LaunchDarkly server-side event is more than 40 minutes old — flag updates may not be reaching environmentd. Labels:
  • audience: platform
  • component: launchdarkly
  • severity: warning

Installed: only when named in rules.selected. Requires: feature-flags.

sum by (namespace) (
  timestamp(mz_parameter_frontend_last_sse_time_seconds{${excludeEnvironmentFilter}})
  - mz_parameter_frontend_last_sse_time_seconds{${excludeEnvironmentFilter}}
) >= 40 * 60

launchdarkly-stale-cse #

The last LaunchDarkly client-side event is more than 40 minutes old — client analytics may be stale. Labels:
  • audience: platform
  • component: launchdarkly
  • severity: notice

Installed: only when named in rules.selected. Requires: feature-flags.

sum by (namespace) (
  timestamp(mz_parameter_frontend_last_cse_time_seconds{${excludeEnvironmentFilter}})
  - mz_parameter_frontend_last_cse_time_seconds{${excludeEnvironmentFilter}}
) >= 40 * 60

materialize-log-alerts#

Log-derived alerting rules for Materialize.

Each alert matches a line Materialize logs when something has gone wrong that its metrics do not show: a process panicking, or an invariant a dataflow relies on for correct results failing. They are LogQL, evaluated by the Loki ruler against the logs the pipeline collects, and route through the same Alertmanager as the metric alerts.

A match is on the log line’s text, and that text is not a contract. Materialize can reword a message in any release, and a reworded message makes the rule silent rather than broken. Each alert names the source of the line it matches, so a contributor can check it still exists.

The window is the alert’s duration. A rule fires on the first evaluation at which a matching line falls inside its range and for is 0s, so the range is how long the alert keeps firing after the last matching line, not how long the condition must persist.

The label contract these read is documented under Logs and Events: namespace, app, container and level are stream labels; pod and the pipeline’s panic fields are structured metadata, which a rule may still group by.

materialize-panic #

A Materialize process panicked. The pod’s logs from just before the restart hold the panic message and backtrace. Labels:
  • audience: platform
  • component: materialize
  • severity: warning

Installed: by default, wherever it applies.

Evaluated by: the Loki ruler, as LogQL.

sum by (namespace, app, pod, panic_location) (
  count_over_time(
    {${mzDeploymentNamespaceFilter}, ${excludeEnvironmentFilter}, level="CRITICAL"}
      | panic_location != ""
    [5m]
  )
) > 0

data-correctness-error #

Materialize logged a dataflow invariant violation, so results from the affected objects may be incorrect. Contact Materialize support. Labels:
  • audience: platform
  • component: compute
  • severity: critical

Installed: by default, wherever it applies.

Evaluated by: the Loki ruler, as LogQL.

sum by (namespace, app, pod) (
  count_over_time(
    {${mzEnvironmentNamespaceFilter}, ${excludeEnvironmentFilter}, level="ERROR"}
      |~ "(?i)(non-positive accumulation|negative accumulation|negative multiplicit(y|ies)|non-positive multiplicity|invalid negative unsigned aggregation|invalid data in source|net-zero records with non-zero accumulation|non-monotonic input to monotonictop1)"
    [15m]
  )
) > 0

persist-filter-pushdown-violation #

Persist filter pushdown skipped data a query needed. Disable it with ALTER SYSTEM SET persist_stats_filter_enabled = false and contact Materialize support. Labels:
  • audience: platform
  • component: persist
  • severity: critical

Installed: by default, wherever it applies.

Evaluated by: the Loki ruler, as LogQL.

sum by (namespace, app, pod) (
  count_over_time(
    {${mzEnvironmentNamespaceFilter}, ${excludeEnvironmentFilter}}
      |= "persist filter pushdown correctness violation"
    [15m]
  )
) > 0

trace-logging-enabled #

A Materialize component has logged at TRACE level for 30m, which multiplies log volume. Reset it with ALTER SYSTEM RESET log_filter. Labels:
  • audience: platform
  • component: materialize
  • severity: notice

Installed: only when named in rules.selected.

Evaluated by: the Loki ruler, as LogQL.

sum by (namespace, app) (
  count_over_time(
    {${mzEnvironmentNamespaceFilter}, ${excludeEnvironmentFilter}, level="TRACE"}
    [10m]
  )
) > 0

materialize-workload-alerts#

Alerting rules for the workloads running on Materialize.

Every alert here carries audience: workload, so a route can send them to the people who own the clusters rather than to whoever runs the deployment. Each is scoped to user clusters (u*); the same condition on a system cluster is a platform alert.

cluster-falling-behind #

A cluster’s collections have averaged over 60s of lag for 15m. Labels:
  • audience: workload
  • component: clusterd
  • severity: warning

Installed: only when named in rules.selected. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, instance_id) (
  avg_over_time(
    (
      min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
        mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"u.*", ${excludeEnvironmentFilter}}
      ) < 1e18
    )[15m:1m]
  )
) > 60

cluster-stale #

A cluster’s collections have averaged over 10m of lag for 15m. Labels:
  • audience: workload
  • component: clusterd
  • severity: critical

Installed: only when named in rules.selected. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, instance_id) (
  avg_over_time(
    (
      min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
        mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"u.*", ${excludeEnvironmentFilter}}
      ) < 1e18
    )[15m:1m]
  )
) > 600

cluster-hydration-stuck #

A cluster has had collections with no hydrated replica for 1h. Labels:
  • audience: workload
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: materialize.

count by (namespace, materialize_cloud_organization_name, instance_id) (
  min by (namespace, materialize_cloud_organization_name, instance_id, collection_id) (
    mz_dataflow_wallclock_lag_seconds{quantile="1", instance_id=~"u.*", ${excludeEnvironmentFilter}}
  ) >= 1e18
) > 0

cluster-memory-near-limit #

A cluster replica’s heap has reached 90% of the point where the kernel OOM-kills it. Labels:
  • audience: workload
  • component: clusterd
  • severity: critical

Installed: by default, wherever it applies. Requires: materialize.

100 * (
  max by (namespace, materialize_cloud_organization_name, cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, pod) (
    max_over_time(mz_metrics_resource_usage{metric="heap", cluster_environmentd_materialize_cloud_cluster_id=~"u.*", ${excludeEnvironmentFilter}}[15m])
  )
  / on (namespace, pod) group_left ()
  (
    max by (namespace, pod) (mz_metrics_resource_usage{metric="memory_max"})
    + on (namespace, pod)
    max by (namespace, pod) (mz_metrics_resource_usage{metric="swap_max"})
  )
) > 90

cluster-memory-high #

A cluster replica’s heap has averaged over 40% of its memory plus swap limits for 15m. Labels:
  • audience: workload
  • component: clusterd
  • severity: warning

Installed: only when named in rules.selected. Requires: materialize.

100 * (
  max by (namespace, materialize_cloud_organization_name, cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, pod) (
    avg_over_time(mz_metrics_resource_usage{metric="heap", cluster_environmentd_materialize_cloud_cluster_id=~"u.*", ${excludeEnvironmentFilter}}[15m])
  )
  / on (namespace, pod) group_left ()
  (
    max by (namespace, pod) (mz_metrics_resource_usage{metric="memory_max"})
    + on (namespace, pod)
    max by (namespace, pod) (mz_metrics_resource_usage{metric="swap_max"})
  )
) > 40

cluster-cpu-high #

A cluster replica has used over 85% of its CPU limit for 15m. Labels:
  • audience: workload
  • component: clusterd
  • severity: warning

Installed: only when named in rules.selected. Requires: kube-state-metrics, cadvisor.

100 * max by (namespace, pod, cluster_environmentd_materialize_cloud_cluster_id) (
  label_replace(
    sum by (namespace, pod) (rate(container_cpu_usage_seconds_total{container="clusterd", pod=~".*-cluster-u[0-9]+-replica-.*", ${excludeEnvironmentFilter}}[15m]))
    / sum by (namespace, pod) (kube_pod_container_resource_limits{container="clusterd", resource="cpu", pod=~".*-cluster-u[0-9]+-replica-.*"}),
    "cluster_environmentd_materialize_cloud_cluster_id", "$1", "pod", ".*-cluster-([a-z0-9]+)-replica-.*"
  )
) > 85

cluster-replica-oomkilled #

A cluster replica was OOM-killed in the last 10m. Labels:
  • audience: workload
  • component: clusterd
  • severity: warning

Installed: by default, wherever it applies. Requires: kube-state-metrics.

max by (namespace, pod, cluster_environmentd_materialize_cloud_cluster_id) (
  label_replace(
    kube_pod_container_status_last_terminated_reason{reason="OOMKilled", container="clusterd", pod=~".*-cluster-u[0-9]+-replica-.*", ${mzEnvironmentNamespaceFilter}, ${excludeEnvironmentFilter}},
    "cluster_environmentd_materialize_cloud_cluster_id", "$1", "pod", ".*-cluster-([a-z0-9]+)-replica-.*"
  )
)
and on (namespace, pod)
max by (namespace, pod) (increase(kube_pod_container_status_restarts_total{container="clusterd"}[10m:1m])) > 0

source-disconnected #

A source has lost sight of its upstream for 5m. Labels:
  • audience: workload
  • component: storage
  • severity: warning

Installed: only when named in rules.selected. Requires: materialize.

max by (namespace, materialize_cloud_organization_name, source_id) (
  mz_source_offset_committed{${excludeEnvironmentFilter}}
)
> max by (namespace, materialize_cloud_organization_name, source_id) (
  mz_source_offset_known{${excludeEnvironmentFilter}}
)