Common Queries#
These queries are used in various dashboards and may be a good starting point for a dashboard developer or for someone looking into particulars of a dashboard.
The underlined values in each query are placeholders, and they are editable. Click one, type your environment’s value, and every other occurrence of that placeholder on the page follows — so copying a query gives you something you can paste straight into Prometheus. Press Enter to commit an edit or Escape to discard it.
You can also set placeholders up front by naming them in this page’s URL:
?mzSqlPrefix=v2_mz_&mzNamespaceList=materialize-prodThe two directions are the same mechanism: editing in place rewrites the URL to match, so once the page reads the way you want, the address bar holds a link you can bookmark or share. Only placeholders you have changed appear there, and a customized value is underlined with a solid rather than a dashed line. Reload without the query string to get the defaults back.
infra-alloy#
Is telemetry collection healthy, and if not, which stage of it broke.
Alloy runs in two roles in this stack. The agent is a DaemonSet with one pod per node. It reads container log files and the node journal, and pushes both to the gateway. The gateway is a Deployment. It receives the agents' logs, collects Kubernetes events, scrapes every metrics target in the cluster, and delivers all of it to the stores the dashboards read.
Every signal in this file travels through the gateway before it can be read. The gateway scrapes Alloy’s own metrics, and Alloy’s own logs reach Loki through the gateway. An empty panel here can therefore mean the gateway is down rather than that nothing happened. When every panel is empty at once, the first check is whether any gateway pod is running.
Scoping#
Three variables, written literally on the same precedent as infra-loki.yaml.
$alloyNamespaceselects where the collectors run.$alloyRoleselectsalloy-agent,alloy-gateway, or both, matched againstapp. Metrics and logs both takeappfrom the pod’sapp.kubernetes.io/namelabel, so one picker serves both engines. Its “All” value is the literal pair rather than.+, which makes it the anchor that keeps every selector to this chart’s collectors.thanos-rulerpublishes the sameprometheus_remote_storage_*family, and an unanchored selector would add its remote-write queue to the gateway’s.$alloyPodselects one collector, most often the agent on one node.
kube-state-metrics and cAdvisor series carry no app label. Those queries
scope with and on (namespace, pod) up{…} against the collectors’ own scrape
targets instead, which applies all three pickers without assuming anything
about pod names.
Queries over up and scrape_* for the targets the gateway scrapes carry
no Alloy scope. Those series describe the target, not the collector, and
nothing on them records which gateway replica produced them. The same holds
for the remote-write senders’ own queues, which are found by their url.
Kubernetes events carry the involved object’s name rather than an app
label. The event queries match name against the role pattern, as
${alloyRole:regex} since it is embedded in a longer pattern, and against the
pod picker. They also match the pre-install validation Jobs, whose names
follow the release rather than the role, as .+-validate-(agent|gateway).
Zero only when something was collected#
Several counts here take the form sum(A) or 0 * sum(B), where B is the
population A is drawn from. Alloy creates an unhealthy-component series only
once a component is unhealthy, so sum(A) alone is empty on a healthy
install. The or arm turns that into a zero, and only when B exists. A
panel reading zero has heard from the collectors; a panel reading nothing
has not.
Deliberate drops#
loki_process_dropped_lines_total counts two different things under one
name. The debug tap in both pipelines forks a copy of the stream, keeps about
1% of it, and drops the rest with reasons sampleDebug, skip self and
skip loki. The gateway also drops Alloy’s own loki.echo output with reason
alloy loki.echo, which stops the tap re-collecting itself. None of those
lines were bound for a store. Every other reason is a guard on the main path,
and a line dropped there is lost. infra.alloy.log_pipeline.guard_drops
excludes the four deliberate reasons by name, so a reason added to a pipeline
later shows up as a loss until it is classified here.
infra.alloy.health.agents_missing #
Nodes the agent DaemonSet is scheduled onto whose agent is not reporting.clamp_min(
sum(
kube_daemonset_status_desired_number_scheduled{
namespace=~"$alloyNamespace",
daemonset="alloy-agent"
}
)
- (
sum(
up{
app="alloy-agent",
namespace=~"$alloyNamespace"
}
)
or vector(0)
),
0
)
infra.alloy.health.unhealthy_components #
Pipeline components that Alloy reports as anything other than healthy, across the selected collectors.sum(
alloy_component_controller_running_components{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
health_type!="healthy"
}
)
or
0 * sum(
alloy_component_controller_running_components{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.health.config_failed #
Collectors whose most recent attempt to load their configuration failed.count(
alloy_config_last_load_successful{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
} == 0
)
or
0 * count(
alloy_config_last_load_successful{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.health.log_lines_lost #
Log lines per second the collectors gave up delivering, across every hop and every reason.sum(
rate(
loki_write_dropped_entries_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.health.samples_lost #
Metric samples per second the gateway failed to deliver and will not retry, across every metrics destination.(
(
sum(
rate(
prometheus_remote_storage_samples_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or vector(0)
)
+ (
sum(
rate(
otelcol_exporter_send_failed_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or vector(0)
)
)
and on ()
(
sum(
rate(
prometheus_remote_storage_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or
sum(
rate(
otelcol_exporter_sent_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.health.pushes_refused #
Push requests per second the gateway answered with an error, across its log-push and remote-write listeners.(
(
sum(
rate(
loki_source_api_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
route!="other",
status_code!~"2.."
}
[5m]
)
)
or vector(0)
)
+ (
sum(
rate(
prometheus_receive_http_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
route!="other",
status_code!~"2.."
}
[5m]
)
)
or vector(0)
)
)
and on ()
(
sum(
rate(
loki_source_api_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or
sum(
rate(
prometheus_receive_http_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.health.up #
Whether the gateway can scrape each collector’s own metrics endpoint.max by (app, pod) (
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.health.restarts #
Container restarts per collector over the selected time range.sum by (namespace, pod) (
increase(
kube_pod_container_status_restarts_total{
namespace=~"$alloyNamespace",
container="alloy"
}
[1h]
)
) > 0
and on (namespace, pod)
max_over_time(
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[1h]
)
infra.alloy.health.versions #
The Alloy version each role is running, and how many collectors run it.count by (app, version) (
alloy_build_info{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.health.lines_delivered #
Log lines per second each collector delivered, split by where it sent them.sum by (app, host) (
rate(
loki_write_sent_entries_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.health.samples_delivered #
Metric samples per second the gateway delivered, split by destination.sum by (component_id) (
rate(
prometheus_remote_storage_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum by (component_id) (
rate(
otelcol_exporter_sent_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.read #
Log lines per second the agents are reading, from container log files and the node journal together.sum(
rate(
loki_source_file_read_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
+ (
sum(
rate(
loki_source_journal_target_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or vector(0)
)
infra.alloy.log_pipeline.stored #
Log lines per second the gateway delivered to its log destinations.sum(
rate(
loki_write_sent_entries_total{
app=~"$alloyRole",
app="alloy-gateway",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.propagation #
How long a log line took from being written to being accepted by the next hop, at the 99th percentile, per collector role.histogram_quantile(
0.99,
sum by (le, app) (
rate(
loki_write_entry_propagation_latency_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.log_pipeline.files #
Container log files each agent currently has open.sum by (pod) (
loki_source_file_files_active_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.log_pipeline.file_lines #
Container log lines per second each agent is reading.sum by (pod) (
rate(
loki_source_file_read_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.journal_lines #
Node journal lines per second each agent is reading.sum by (pod) (
rate(
loki_source_journal_target_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.requests #
Log delivery requests per second, split by destination and the status code the destination returned.sum by (host, status_code) (
rate(
loki_write_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.retries #
Log delivery batches per second a collector had to retry, per destination.sum by (app, host) (
rate(
loki_write_batch_retries_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.dropped #
Log lines per second a collector gave up delivering, split by collector role and reason.sum by (app, reason) (
rate(
loki_write_dropped_entries_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
) > 0
or on (app)
0 * sum by (app) (
rate(
loki_write_sent_entries_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.request_latency #
How long log delivery requests take to be acknowledged, at the 99th percentile, per destination.histogram_quantile(
0.99,
sum by (le, host) (
rate(
loki_write_request_duration_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.log_pipeline.guard_drops #
Log lines per second a pipeline guard discarded before delivery, split by reason.sum by (app, reason) (
rate(
loki_process_dropped_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
reason!~"sampleDebug|skip self|skip loki|alloy loki.echo"
}
[5m]
)
) > 0
or on (app)
0 * sum by (app) (
rate(
loki_process_dropped_lines_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.log_pipeline.truncated #
Container log lines per second the agents truncated while reassembling lines the container runtime had split.sum by (pod) (
rate(
loki_process_cri_lines_truncated_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.targets #
Metrics targets the gateways are scraping, summed across replicas.sum(
prometheus_target_scrape_pool_targets{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.targets_down #
Metrics targets whose most recent scrape failed.count(up == 0)
or
0 * count(up)
infra.alloy.metric_pipeline.samples_sent #
Samples per second the gateway wrote to its remote-write destinations.sum(
rate(
prometheus_remote_storage_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.send_delay #
How far behind collection each remote-write queue is sending, in seconds.clamp_min(
max by (pod, component_id) (
prometheus_remote_storage_highest_timestamp_in_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
- ignoring (url, remote_name) group_right (pod)
prometheus_remote_storage_queue_highest_sent_timestamp_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
),
0
) < 31536000
infra.alloy.metric_pipeline.send_delay.worst #
The furthest behind any remote-write queue is sending, in seconds.max(
clamp_min(
prometheus_remote_storage_highest_timestamp_in_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
- ignoring (url, remote_name) group_right (pod)
prometheus_remote_storage_queue_highest_sent_timestamp_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
},
0
) < 31536000
)
infra.alloy.metric_pipeline.down_targets #
Every metrics target whose most recent scrape failed.max by (job, namespace, pod, instance) (
up == 0
)
infra.alloy.metric_pipeline.targets_by_monitor #
Targets each ServiceMonitor, PodMonitor or scrape job currently matches, summed across gateway replicas.sum by (scrape_job) (
prometheus_target_scrape_pool_targets{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.rejected_scrapes #
Whole scrapes per second the gateway discarded for breaking a configured limit.sum(
rate(
prometheus_target_scrapes_exceeded_sample_limit_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_target_scrapes_exceeded_body_size_limit_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_target_scrape_pool_exceeded_target_limit_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_target_scrape_pool_exceeded_label_limits_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.rejected_samples #
Individual samples per second the gateway discarded from otherwise successful scrapes.sum(
rate(
prometheus_target_scrapes_sample_out_of_order_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_target_scrapes_sample_duplicate_timestamp_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_target_scrapes_sample_out_of_bounds_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.slowest_scrapes #
The ten scrape jobs taking longest to scrape, by their slowest target.topk(
10,
max by (job) (
scrape_duration_seconds
)
)
infra.alloy.metric_pipeline.largest_scrapes #
The ten scrape jobs contributing the most samples per scrape, after relabelling.topk(
10,
sum by (job) (
scrape_samples_post_metric_relabeling
)
)
infra.alloy.metric_pipeline.samples_scraped #
Samples per second the gateway scraped, split by the component that scraped them.sum by (component_id) (
rate(
prometheus_forwarded_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
component_id=~"prometheus\\.(scrape|operator)\\..*"
}
[5m]
)
)
infra.alloy.metric_pipeline.remote_write #
Samples per second each remote-write destination took, rejected, and had retried.sum by (component_id) (
rate(
prometheus_remote_storage_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum by (component_id) (
rate(
prometheus_remote_storage_samples_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum by (component_id) (
rate(
prometheus_remote_storage_samples_retried_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.shards #
Parallel senders each remote-write queue is running, against how many it wants and how many it is allowed.max by (component_id) (
prometheus_remote_storage_shards{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
max by (component_id) (
prometheus_remote_storage_shards_desired{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
max by (component_id) (
prometheus_remote_storage_shards_max{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.send_latency #
How long remote-write batches take to be acknowledged, at the 99th percentile, per destination.histogram_quantile(
0.99,
sum by (le, component_id) (
rate(
prometheus_remote_storage_sent_batch_duration_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.metric_pipeline.active_series #
Series each gateway replica is holding in its write-ahead log.sum by (pod) (
prometheus_remote_write_wal_storage_active_series{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.wal_size #
Bytes each gateway replica’s write-ahead log occupies on disk.sum by (pod) (
prometheus_tsdb_wal_storage_size_bytes{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.wal_errors #
Write-ahead log failures per second: corruptions, failed writes, and failed checkpoints or truncations.sum(
rate(
prometheus_remote_write_wal_corruptions_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_tsdb_wal_writes_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_remote_write_wal_checkpoint_creations_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
prometheus_tsdb_wal_truncations_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.otel_sent #
Metric points and log records per second each OpenTelemetry exporter delivered.sum by (component_id) (
rate(
otelcol_exporter_sent_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum by (component_id) (
rate(
otelcol_exporter_sent_log_records_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.otel_failed #
Metric points and log records per second each OpenTelemetry exporter failed to deliver.sum by (component_id) (
rate(
otelcol_exporter_send_failed_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum by (component_id) (
rate(
otelcol_exporter_send_failed_log_records_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.metric_pipeline.otel_queue #
How full each OpenTelemetry exporter’s sending queue is, as a fraction of its capacity.max by (component_id, data_type) (
otelcol_exporter_queue_size{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
/ otelcol_exporter_queue_capacity{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.metric_pipeline.otel_filtered #
Metric points per second each OpenTelemetry filter removed.sum by (component_id) (
rate(
otelcol_processor_filter_datapoints_filtered_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.log_push.lines #
Log lines per second the gateway accepted on its log-push listener.sum(
rate(
loki_source_api_entries_written{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.log_push.requests #
Log push requests per second the gateway received, split by the status it returned.sum by (route, status_code) (
rate(
loki_source_api_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.log_push.latency #
How long the gateway takes to accept a log push, at the 99th percentile.histogram_quantile(
0.99,
sum by (le) (
rate(
loki_source_api_request_duration_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.ingest.remote_write.received #
Samples per second the gateway accepted over Prometheus remote write.sum(
rate(
prometheus_forwarded_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
component_id=~"prometheus\\.receive_http\\..*"
}
[5m]
)
)
infra.alloy.ingest.remote_write.samples #
Samples per second the gateway accepted over remote write, beside the samples it dropped for carrying invalid labels.sum(
rate(
prometheus_forwarded_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
component_id=~"prometheus\\.receive_http\\..*"
}
[5m]
)
)
sum(
rate(
prometheus_api_remote_write_invalid_labels_samples_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.remote_write.requests #
Remote-write requests per second the gateway answered, split by status code.sum by (route, status_code) (
rate(
prometheus_receive_http_request_duration_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.remote_write.latency #
How long the gateway takes to answer a remote-write request, at the 99th percentile.histogram_quantile(
0.99,
sum by (le) (
rate(
prometheus_receive_http_request_duration_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.ingest.remote_write.senders #
What each remote-write sender pointed at the gateway reports about its own queue: samples sent, failed, and retried.sum by (job) (
rate(
prometheus_remote_storage_samples_total{
url=~".*alloy-gateway.*"
}
[5m]
)
)
sum by (job) (
rate(
prometheus_remote_storage_samples_failed_total{
url=~".*alloy-gateway.*"
}
[5m]
)
)
sum by (job) (
rate(
prometheus_remote_storage_samples_retried_total{
url=~".*alloy-gateway.*"
}
[5m]
)
)
infra.alloy.ingest.otlp.metrics #
Metric points per second the gateway accepted and refused over OTLP.sum(
rate(
otelcol_receiver_accepted_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
otelcol_receiver_refused_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.otlp.logs #
Log records per second the gateway accepted and refused over OTLP.sum(
rate(
otelcol_receiver_accepted_log_records_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
sum(
rate(
otelcol_receiver_refused_log_records_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.ingest.tls_rejections #
TLS handshakes per minute that a collector’s listener refused, usually because the client presented a certificate the listener does not trust.infra.alloy.components.not_healthy #
Components per collector in any state other than healthy.sum by (pod, health_type) (
alloy_component_controller_running_components{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod",
health_type!="healthy"
}
)
or on (pod)
0 * sum by (pod) (
alloy_component_controller_running_components{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.components.running #
Components each collector is running.sum by (pod) (
alloy_component_controller_running_components{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.components.config_by_pod #
The configuration each collector is running, by content hash, whether its last load succeeded, and how long the collector has been up.max by (app, pod, sha256) (
alloy_config_hash{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
* on (namespace, pod) group_left ()
alloy_config_last_load_successful{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
max by (app, pod) (
time()
- alloy_resources_process_start_time_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.components.distinct_configs #
Distinct configurations running per role.count by (app) (
count by (app, sha256) (
alloy_config_hash{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
)
infra.alloy.components.evaluation_rate #
Component evaluations per second per role.sum by (app) (
rate(
alloy_component_evaluation_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.components.evaluation_latency #
How long a component evaluation takes, at the 99th percentile, per role.histogram_quantile(
0.99,
sum by (le, app) (
rate(
alloy_component_evaluation_seconds_bucket{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
)
infra.alloy.components.slow #
Time per second spent in component evaluations that Alloy classed as slow, per component.sum by (app, component_id) (
rate(
alloy_component_evaluation_slow_seconds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or on (app)
0 * sum by (app) (
rate(
alloy_component_evaluation_seconds_count{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.components.queue #
Component evaluations waiting to run, per collector.max by (pod) (
alloy_component_evaluation_queue_size{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.cluster.peers #
How many cluster members each gateway replica can see, beside how many replicas are running.sum by (pod) (
cluster_node_peers{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
count(
cluster_node_info{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.cluster.health_score #
Each gateway replica’s gossip health score, where zero is healthy.max by (pod) (
cluster_node_gossip_health_score{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.cluster.targets #
Metrics targets each gateway replica is scraping.sum by (pod) (
prometheus_target_scrape_pool_targets{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.cluster.transport_failures #
Gossip messages per second each gateway replica failed to send to its peers.sum by (pod) (
rate(
cluster_transport_tx_packets_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
+ sum by (pod) (
rate(
cluster_transport_stream_tx_packets_failed_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.cluster.members #
Each gateway replica’s own cluster state, and whether it considers the cluster ready for traffic.min by (pod, state) (
cluster_ready_for_traffic{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
* on (namespace, pod) group_left (state)
cluster_node_info{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.resources.cpu #
CPU each collector is using, as a fraction of its CPU limit.max by (namespace, pod) (
rate(
container_cpu_usage_seconds_total{
namespace=~"$alloyNamespace",
container="alloy"
}
[5m]
)
)
/ max by (namespace, pod) (
kube_pod_container_resource_limits{
namespace=~"$alloyNamespace",
container="alloy",
resource="cpu"
}
)
and on (namespace, pod)
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
infra.alloy.resources.throttling #
The fraction of scheduling periods in which each collector was held back by its CPU limit.sum by (namespace, pod) (
rate(
container_cpu_cfs_throttled_periods_total{
namespace=~"$alloyNamespace",
container="alloy"
}
[5m]
)
)
/ sum by (namespace, pod) (
rate(
container_cpu_cfs_periods_total{
namespace=~"$alloyNamespace",
container="alloy"
}
[5m]
)
)
and on (namespace, pod)
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
infra.alloy.resources.memory #
Memory each collector is using, as a fraction of its memory limit.max by (namespace, pod) (
container_memory_working_set_bytes{
namespace=~"$alloyNamespace",
container="alloy"
}
)
/ max by (namespace, pod) (
kube_pod_container_resource_limits{
namespace=~"$alloyNamespace",
container="alloy",
resource="memory"
}
)
and on (namespace, pod)
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
infra.alloy.resources.gomemlimit #
Memory the Go runtime holds in each collector, as a fraction of the soft limitGOMEMLIMIT sets for it.max by (pod) (
(
go_memstats_sys_bytes{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
- go_memstats_heap_released_bytes{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
/ go_gc_gomemlimit_bytes{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.resources.memory_limiter #
Metric points per second the gateway’s memory limiter refused.sum by (pod) (
rate(
otelcol_processor_memory_limiter_refused_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
or
0 * sum by (pod) (
rate(
otelcol_processor_memory_limiter_accepted_metric_points_total{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
[5m]
)
)
infra.alloy.resources.goroutines #
Concurrent goroutines in each collector.max by (pod) (
go_goroutines{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.resources.file_descriptors #
Open file descriptors in each collector, as a fraction of its limit.max by (pod) (
process_open_fds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
/ process_max_fds{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
)
infra.alloy.resources.terminations #
Why each collector’s container last stopped, for those that have.max by (namespace, pod, reason) (
kube_pod_container_status_last_terminated_reason{
namespace=~"$alloyNamespace",
container="alloy"
}
) == 1
and on (namespace, pod)
up{
app=~"$alloyRole",
namespace=~"$alloyNamespace",
pod=~"$alloyPod"
}
infra.alloy.events.rate.by_reason #
Kubernetes events about the collectors’ pods, DaemonSet, Deployment and autoscaler, and about the configuration-validation Jobs that run before each install and upgrade, per reason.infra.alloy.events.rate.by_kind #
Kubernetes events about the collectors, per kind of object they concern.infra.alloy.events.warnings #
Kubernetes warning events about the collectors, newest first.infra.alloy.events.stream #
Every Kubernetes event about the collectors, newest first.infra.alloy.logs.stream #
The collectors’ own log feed, newest first, for the selected roles and levels.infra.alloy.logs.warnings.stream #
Warning-and-worse lines from the collectors, newest first.infra.alloy.logs.rate.by_role #
The collectors’ own log lines per second, split by role and level.infra.alloy.logs.errors.by_component #
Warning-and-worse lines per minute from the collectors, split by the pipeline component that logged them.infra.alloy.logs.warnings.rate #
Warning-and-worse lines per minute from the selected collectors.infra-autoscaling#
How the cluster’s nodes follow its pods, whatever does the following.
Three things add nodes to a Materialize deployment, depending on the cloud: Karpenter on EKS, and the managed cluster autoscaler on GKE and AKS. They expose nothing in common, so everything here is read from what every cluster has instead:
- kube-state-metrics says how many pods are waiting for a node, how many nodes there are, and how full each is to the scheduler.
- Kubernetes events say what each autoscaler decided: Karpenter’s
NominatedandLaunched, the cluster autoscaler’sTriggeredScaleUpandNotTriggerScaleUp, the scheduler’sFailedScheduling, and the node lifecycle’sRegisteredNodeandDeletingNode. - The cloud provider pull says whether the cloud could supply the nodes asked for: EC2 status checks and node-group sizes, Compute Engine quota, and on AKS the managed autoscaler’s own gauges. It is opt-in, and its rows render only where a provider is pulled.
Karpenter’s own metrics, which say far more on EKS, are the Karpenter
dashboard’s (infra-karpenter.yaml).
Node pools#
%%{nodePools} gives every node its pool, instance type and zone, from
kube_node_labels. The pool is whichever label the node’s provisioner
sets: a Karpenter NodePool, an EKS managed node group, a GKE node pool or an
AKS agent pool. kube-state-metrics publishes those labels only because the
chart’s metricLabelsAllowlist names them, and publishes no
kube_node_labels at all without it.
Deduplication#
Every kube-state-metrics family here is read through max by its own
identity first, because each extra kube-state-metrics replica reports every
object again. The same shape infra-nodes.yaml uses.
Provider lookback#
Provider samples arrive with the pull, every five minutes by default, so
they are read through last_over_time(...[15m]) and never rated over
$__rate_interval; infra-cloud.yaml explains why.
infra.autoscaling.pods.pending #
Pods in the Pending phase right now, across the cluster.sum(max by (namespace, pod) (kube_pod_status_phase{phase="Pending"}))
infra.autoscaling.pods.unscheduled #
Pods the scheduler has not placed on any node, because no node has room for them or satisfies their constraints.sum(max by (namespace, pod) (kube_pod_status_scheduled{condition="false"}))
infra.autoscaling.nodes.count #
Nodes in the cluster.count(max by (node) (kube_node_info))
infra.autoscaling.nodes.not_ready #
Nodes whose kubelet is not reporting Ready.sum(1 - max by (node) (kube_node_status_condition{condition="Ready", status="true"}))
infra.autoscaling.nodes.added #
Nodes in the cluster now that were not at the start of the selected time range.count(
max by (node) (kube_node_info)
unless max by (node) (kube_node_info offset ${rangeWindow})
)
or 0 * count(max by (node) (kube_node_info))
infra.autoscaling.nodes.removed #
Nodes that left the cluster in the selected time range.infra.autoscaling.hpa.at_max #
Workloads whose HorizontalPodAutoscaler is running as many replicas as it is allowed to.count(
max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_current_replicas)
>= max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_spec_max_replicas)
)
or 0 * count(max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_spec_max_replicas))
infra.autoscaling.pods.unscheduled_by_namespace #
Pods waiting for a node, by namespace.sum by (namespace) (max by (namespace, pod) (kube_pod_status_scheduled{condition="false"})) > 0
infra.autoscaling.pods.pending_by_namespace #
Pods in the Pending phase, by namespace.sum by (namespace) (max by (namespace, pod) (kube_pod_status_phase{phase="Pending"})) > 0
infra.autoscaling.nodes.by_pool #
Nodes in each node pool over time.count by (pool) (${nodePools})
infra.autoscaling.nodes.by_instance_type #
Nodes of each instance type over time.count by (instance_type) (${nodePools})
infra.autoscaling.nodes.by_zone #
Nodes in each availability zone over time.count by (zone) (${nodePools})
infra.autoscaling.capacity.requested #
How much of the cluster’s allocatable CPU and memory is promised to running pods through their requests.sum(
max by (namespace, pod, container, node) (kube_pod_container_resource_requests{resource="cpu", node!=""})
unless on (namespace, pod) (max by (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"}) == 1)
)
/
sum(max by (node) (kube_node_status_allocatable{resource="cpu"}))
sum(
max by (namespace, pod, container, node) (kube_pod_container_resource_requests{resource="memory", node!=""})
unless on (namespace, pod) (max by (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"}) == 1)
)
/
sum(max by (node) (kube_node_status_allocatable{resource="memory"}))
infra.autoscaling.pools.cpu_requested #
The share of each node pool’s allocatable CPU promised to running pods.sum by (pool) (
sum by (node) (
max by (namespace, pod, container, node) (kube_pod_container_resource_requests{resource="cpu", node!=""})
unless on (namespace, pod) (max by (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"}) == 1)
)
* on (node) group_left (pool) ${nodePools}
)
/
sum by (pool) (
max by (node) (kube_node_status_allocatable{resource="cpu"})
* on (node) group_left (pool) ${nodePools}
)
infra.autoscaling.pools.memory_requested #
The share of each node pool’s allocatable memory promised to running pods.sum by (pool) (
sum by (node) (
max by (namespace, pod, container, node) (kube_pod_container_resource_requests{resource="memory", node!=""})
unless on (namespace, pod) (max by (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"}) == 1)
)
* on (node) group_left (pool) ${nodePools}
)
/
sum by (pool) (
max by (node) (kube_node_status_allocatable{resource="memory"})
* on (node) group_left (pool) ${nodePools}
)
infra.autoscaling.pools.not_ready #
Nodes not reporting Ready, by node pool.sum by (pool) (
(1 - max by (node) (kube_node_status_condition{condition="Ready", status="true"}))
* on (node) group_left (pool) ${nodePools}
)
infra.autoscaling.pools.cordoned #
Nodes marked unschedulable, by node pool.sum by (pool) (
max by (node) (kube_node_spec_unschedulable)
* on (node) group_left (pool) ${nodePools}
)
infra.autoscaling.nodes.age #
Every node with its pool, instance type, zone, and how long ago it joined.(time() - max by (node) (kube_node_created))
* on (node) group_left (pool, instance_type, zone) ${nodePools}
infra.autoscaling.pods.unscheduled_list #
Each pod waiting for a node, and how long since it was created.(time() - max by (namespace, pod) (kube_pod_created))
and on (namespace, pod) (max by (namespace, pod) (kube_pod_status_scheduled{condition="false"}) == 1)
infra.autoscaling.events.scheduling_failures #
The scheduler’s explanation for each pod it could not place, newest first.infra.autoscaling.events.scale_up_decisions #
What each autoscaler did about pods waiting for a node, newest first.infra.autoscaling.events.scheduling_failures_rate #
Scheduling failures per minute, by namespace.infra.autoscaling.hpa.replicas #
Every HorizontalPodAutoscaler with its current and desired replicas and the bounds it scales between.max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_current_replicas)
max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_desired_replicas)
max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_spec_min_replicas)
max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_spec_max_replicas)
infra.autoscaling.hpa.share_of_max #
Each HorizontalPodAutoscaler’s replicas as a share of its maximum.max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_current_replicas)
/
max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_spec_max_replicas)
infra.autoscaling.hpa.unable_to_scale #
HorizontalPodAutoscalers that cannot act, by the condition that stopped them.sum by (condition) (
max by (namespace, horizontalpodautoscaler, condition) (
kube_horizontalpodautoscaler_status_condition{condition=~"AbleToScale|ScalingActive", status="false"}
)
)
infra.autoscaling.events.hpa_rate #
HorizontalPodAutoscaler events per minute, by reason.infra.autoscaling.events.hpa_stream #
HorizontalPodAutoscaler events, newest first.infra.autoscaling.cloud.ec2.status_checks #
Cluster nodes failing an EC2 status check, by check.sum(last_over_time(aws_ec2_status_check_failed_system_maximum[15m]))
sum(last_over_time(aws_ec2_status_check_failed_instance_maximum[15m]))
sum(last_over_time(aws_ec2_status_check_failed_attached_ebs_maximum[15m]))
infra.autoscaling.cloud.ec2.node_groups #
Each EKS managed node group’s desired and in-service instances, against its maximum.max by (tag_eks_nodegroup_name) (last_over_time(aws_autoscaling_group_desired_capacity_average[15m]))
max by (tag_eks_nodegroup_name) (last_over_time(aws_autoscaling_group_in_service_instances_average[15m]))
max by (tag_eks_nodegroup_name) (last_over_time(aws_autoscaling_group_max_size_average[15m]))
infra.autoscaling.cloud.ec2.vcpus #
On-Demand vCPUs running in the region, for the whole AWS account.max(last_over_time(aws_usage_resource_count_maximum{dimension_Resource="vCPU"}[15m]))
infra.autoscaling.cloud.gce.cpu_quota #
CPUs in use against the regional quota, for each machine family.max by (vm_family, location) (
last_over_time(stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_cpus_per_vm_family_usage{limit_name=~".*per-project-region"}[15m])
)
/
max by (vm_family, location) (
last_over_time(stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_cpus_per_vm_family_limit{limit_name=~".*per-project-region"}[15m])
)
infra.autoscaling.cloud.gce.ssd_quota #
Local SSD in use against the regional quota, for each machine family.max by (vm_family, location) (
last_over_time(stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_local_ssd_total_storage_per_vm_family_usage{limit_name=~".*per-project-region"}[15m])
)
/
max by (vm_family, location) (
last_over_time(stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_local_ssd_total_storage_per_vm_family_limit{limit_name=~".*per-project-region"}[15m])
)
infra.autoscaling.cloud.gce.refusals #
Requests Compute Engine refused for CPU or local SSD quota, per hour.sum by (vm_family, location) (
increase({__name__=~"stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_cpus_per_vm_family_exceeded(_total)?"}[1h])
)
sum by (vm_family, location) (
increase({__name__=~"stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_local_ssd_total_storage_per_vm_family_exceeded(_total)?"}[1h])
)
infra.autoscaling.cloud.aks.unschedulable #
Pods AKS’s cluster autoscaler cannot place on any existing node, the worst of each five minutes.max by (resourceName) (
last_over_time(azure_microsoft_containerservice_managedclusters_cluster_autoscaler_unschedulable_pods_count_maximum_count[15m])
)
infra.autoscaling.cloud.aks.state #
Whether AKS’s cluster autoscaler considers the cluster safe to scale, whether scale-down is paused, and nodes it would remove.min by (resourceName) (
last_over_time(azure_microsoft_containerservice_managedclusters_cluster_autoscaler_cluster_safe_to_autoscale_minimum_count[15m])
)
max by (resourceName) (
last_over_time(azure_microsoft_containerservice_managedclusters_cluster_autoscaler_scale_down_in_cooldown_maximum_count[15m])
)
max by (resourceName) (
last_over_time(azure_microsoft_containerservice_managedclusters_cluster_autoscaler_unneeded_nodes_count_maximum_count[15m])
)
infra.autoscaling.cloud.aks.vms_down #
Node VMs Azure reports as unavailable, by scale set, at any point in each five minutes.sum by (resourceName) (
1 - min by (resourceName, dimensionVmname) (
last_over_time(azure_microsoft_compute_virtualmachinescalesets_vmavailabilitymetric_minimum_count[15m])
)
)
infra.autoscaling.events.rate_by_reason #
Node and autoscaler events per minute, by reason.infra.autoscaling.events.warnings #
Warning events from the autoscalers and the node lifecycle.infra.autoscaling.events.stream #
Every node and autoscaler event, newest first.infra-cloud#
What the cloud provider publishes about the managed services a Materialize deployment depends on: the metadata database and the object storage buckets.
These are the families the gateway pulls from CloudWatch, Cloud Monitoring and
Azure Monitor when pipeline.metrics.provider.* is enabled — see
packages/alloy-pipelines/gateway-provider.yaml, whose metric sets are the
ones named here. Each family belongs to exactly one provider, so every query
below that covers all three is a list of one expression per provider, and on
any given install two of them return nothing. That is deliberate: it is
cheaper and more honest than a normalization layer, and it keeps each
provider’s own units and resource names visible.
The client’s view of the same dependencies — what Materialize experienced —
is materialize-consensus.yaml and materialize-persist.yaml. The external
dependency design doc records why the two are separate dashboards.
Units#
Each provider speaks its own: CloudWatch and Azure Monitor publish percent
as 0–100, Cloud Monitoring publishes the same quantities as 0–1 (its unit
label reads 10^2.%). Everything here is converted to 0–1 so a panel can
pin its axis. Transaction-ID consumption is converted to a fraction of the
2^31 wraparound limit, which is what Cloud SQL’s own ratio is measured
against.
Lag and the lookback window#
Provider samples are sparse and old. The pull runs every five minutes by
default; CloudWatch and Azure Monitor samples are stamped at the pull,
Cloud SQL’s arrive about three minutes old and GCS’s over ten. A default
five-minute lookback would therefore draw gaps, so every expression reads
its series through last_over_time: 15 minutes for the scrape-stamped
providers and Cloud SQL, 30 for GCS. The cost is that a panel goes on
showing a stopped pull’s last value for that long. rate() is never used
on a provider counter for the same reason: $__rate_interval is sized
from the datasource’s scrape interval, which is far shorter than a pull.
Resource labels#
| Provider | Database | Bucket |
|---|---|---|
| CloudWatch | dimension_DBInstanceIdentifier | dimension_BucketName |
| Cloud Monitoring | database_id (project:instance) | bucket_name |
| Azure Monitor | resourceName | resourceName (the storage account) |
Azure metric names are the exporter’s azure_{type}_{metric}_{aggregation}_{unit},
lowercased, where the type is the metric namespace when one is set — hence
…_storageaccounts_blobservices_… for the blob families.
infra.cloud.db.cpu #
CPU utilization of each metadata database instance, as its cloud provider measures it.max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_cpuutilization_average[15m])
) / 100
max by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_cpu_utilization[15m])
)
max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_cpu_percent_average_percent[15m])
) / 100
infra.cloud.db.connections #
Connections open on each metadata database instance, beside the number Materialize holds.max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_database_connections_maximum[15m])
)
sum by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_postgresql_num_backends[15m])
)
max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_active_connections_maximum_count[15m])
)
sum(mz_persist_postgres_connpool_size) + (sum(mz_ts_oracle_postgres_connpool_size) or vector(0))
infra.cloud.db.transaction_ids #
How much of PostgreSQL’s transaction-ID space each instance has used, as a fraction of the point at which it stops accepting writes.max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_maximum_used_transaction_ids_maximum[15m])
) / 2147483648
max by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_postgresql_transaction_id_utilization[15m])
)
max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_maximum_used_transactionids_maximum_count[15m])
) / 2147483648
infra.cloud.rds.free_storage #
Free storage on each RDS instance.min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_free_storage_space_minimum[15m])
)
infra.cloud.rds.freeable_memory #
Memory on each RDS instance not in use by PostgreSQL or its cache.min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_freeable_memory_minimum[15m])
)
infra.cloud.rds.cpu_credits #
CPU credits left on burstable (db.t*) instances.min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_cpucredit_balance_minimum[15m])
)
infra.cloud.rds.io_latency #
Average time each RDS instance’s disk takes to complete a read and a write.max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_read_latency_average[15m])
)
max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_write_latency_average[15m])
)
infra.cloud.rds.queue_depth #
Disk requests waiting on each RDS instance’s volume.max by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_disk_queue_depth_average[15m])
)
infra.cloud.rds.io_balances #
The burst credit balances that let each RDS volume exceed its baseline performance.min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_burst_balance_minimum[15m])
) / 100
min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_ebsiobalance_percent_minimum[15m])
) / 100
min by (dimension_DBInstanceIdentifier) (
last_over_time(aws_rds_ebsbyte_balance_percent_minimum[15m])
) / 100
infra.cloud.cloudsql.disk #
Fraction of each Cloud SQL instance’s disk in use.max by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_disk_utilization[15m])
)
infra.cloud.cloudsql.memory #
Fraction of each Cloud SQL instance’s memory in use.max by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_memory_utilization[15m])
)
infra.cloud.cloudsql.up #
Whether Cloud SQL considers each instance to be serving.min by (database_id) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_up[15m])
)
infra.cloud.cloudsql.backends_by_database #
Connections on each Cloud SQL instance, by the database they are connected to.sum by (database_id, database) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_postgresql_num_backends[15m])
)
infra.cloud.cloudsql.backends_by_state #
Connections on each Cloud SQL instance, by what they are doing.sum by (database_id, state) (
last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_postgresql_num_backends_by_state[15m])
)
infra.cloud.flexible.storage #
Fraction of each Flexible Server’s storage in use.max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_storage_percent_maximum_percent[15m])
) / 100
infra.cloud.flexible.memory #
Fraction of each Flexible Server’s memory in use.max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_memory_percent_maximum_percent[15m])
) / 100
infra.cloud.flexible.cpu_credits #
CPU credits left on Burstable-tier Flexible Servers.min by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_cpu_credits_remaining_minimum_count[15m])
)
infra.cloud.flexible.alive #
Whether Azure could reach each Flexible Server.min by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_is_db_alive_minimum_count[15m])
)
infra.cloud.flexible.io_consumed #
How much of each Flexible Server’s provisioned disk IOPS and throughput is in use.max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_disk_iops_consumed_percentage_maximum_percent[15m])
) / 100
max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_disk_bandwidth_consumed_percentage_maximum_percent[15m])
) / 100
infra.cloud.flexible.queue_depth #
Disk requests waiting on each Flexible Server.max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_disk_queue_depth_average_count[15m])
)
infra.cloud.flexible.connections_failed #
Connections each Flexible Server refused, per pull interval.max by (resourceName) (
last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_connections_failed_total_count[15m])
)
infra.cloud.bucket.bytes #
Bytes stored in each bucket, as the provider bills them, beside the data Materialize accounts for.max by (dimension_BucketName) (
last_over_time(aws_s3_bucket_size_bytes_average[15m])
)
sum by (bucket_name) (
last_over_time(stackdriver_gcs_bucket_storage_googleapis_com_storage_v_2_total_bytes[30m])
)
max by (resourceName) (
last_over_time(azure_microsoft_storage_storageaccounts_blobservices_blobcapacity_average_bytes[15m])
)
sum(max by (shard) (mz_persist_shard_usage_current_state_batches_bytes))
+ sum(max by (shard) (mz_persist_shard_usage_current_state_rollups_bytes))
+ sum(max by (shard) (mz_persist_shard_usage_referenced_not_current_state_bytes))
+ sum(max by (shard) (mz_persist_shard_usage_not_leaked_not_referenced_bytes))
+ sum(max by (shard) (mz_persist_shard_usage_leaked_bytes))
infra.cloud.bucket.objects #
Objects stored in each bucket.max by (dimension_BucketName) (
last_over_time(aws_s3_number_of_objects_average[15m])
)
sum by (bucket_name) (
last_over_time(stackdriver_gcs_bucket_storage_googleapis_com_storage_v_2_total_count[30m])
)
max by (resourceName) (
last_over_time(azure_microsoft_storage_storageaccounts_blobservices_blobcount_average_count[15m])
)
infra.cloud.gcs.bytes_by_type #
Bytes in each GCS bucket, split into live objects, noncurrent versions and soft-deleted objects.sum by (bucket_name, type) (
last_over_time(stackdriver_gcs_bucket_storage_googleapis_com_storage_v_2_total_bytes[30m])
)
infra.cloud.gcs.objects_by_type #
Objects in each GCS bucket, split into live, noncurrent and soft-deleted.sum by (bucket_name, type) (
last_over_time(stackdriver_gcs_bucket_storage_googleapis_com_storage_v_2_total_count[30m])
)
infra.cloud.blob.availability #
Fraction of requests to each storage account’s blob service that Azure counted as successful.min by (resourceName) (
last_over_time(azure_microsoft_storage_storageaccounts_blobservices_availability_average_percent[15m])
) / 100
infra.cloud.blob.latency #
Average time Azure took to serve each storage account’s successful blob requests, measured at the service and end to end.max by (resourceName) (
last_over_time(azure_microsoft_storage_storageaccounts_blobservices_successserverlatency_average_milliseconds[15m])
) / 1000
max by (resourceName) (
last_over_time(azure_microsoft_storage_storageaccounts_blobservices_successe2elatency_average_milliseconds[15m])
) / 1000
infra.cloud.collection.pulls #
Every provider pull the gateway is running, and whether its last scrape answered.max by (job, instance) (up{job=~"integrations/(cloudwatch|gcp|azure)"})
infra.cloud.collection.resources #
How many of each kind of resource returned data in the last pull.count(max by (dimension_DBInstanceIdentifier) (last_over_time(aws_rds_cpuutilization_average[15m])))
count(max by (dimension_BucketName) (last_over_time(aws_s3_bucket_size_bytes_average[15m])))
count(max by (database_id) (last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_cpu_utilization[15m])))
count(max by (bucket_name) (last_over_time(stackdriver_gcs_bucket_storage_googleapis_com_storage_v_2_total_bytes[30m])))
count(max by (resourceName) (last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_cpu_percent_average_percent[15m])))
count(max by (resourceName) (last_over_time(azure_microsoft_storage_storageaccounts_blobservices_blobcapacity_average_bytes[15m])))
infra.cloud.collection.gcp_errors #
Whether the last Cloud Monitoring pull for each service failed.max by (instance) (
last_over_time(stackdriver_monitoring_last_scrape_error[15m])
)
infra.cloud.collection.gcp_calls #
Cloud Monitoring API calls per minute, by pull — the calls that are billed.sum by (instance) (
rate(stackdriver_monitoring_api_calls_total[30m])
) * 60
infra.cloud.collection.azure_failures #
Azure Monitor API calls that failed, per minute, by HTTP status.sum by (statusCode) (
rate(azurerm_api_request_count{statusCode!~"2.."}[30m])
) * 60
infra.cloud.collection.azure_calls #
Azure API calls per minute, by what they were for — the calls that count against the subscription’s API limits.label_replace(
label_replace(
sum by (resourceProvider, method) (
rate(azurerm_api_request_count[30m])
) * 60,
"resourceProvider", "resource graph", "method", "post"
),
"resourceProvider", "other", "resourceProvider", ""
)
infra-karpenter#
Karpenter’s own account of provisioning and disruption, from the metrics its controller publishes and the events it records.
Karpenter adds nodes on EKS in the self-managed Terraform: it watches for pods the scheduler cannot place, launches an EC2 instance that fits them, and later removes or replaces nodes that are empty, underused, drifted from their NodePool’s spec, or being interrupted by AWS. Everything it does to a node goes through a NodeClaim, one per instance.
The controller runs two replicas and only the leader does any work, so most
families here come from one pod. Counters restart on a leader change; every
query reads them through rate or increase, which tolerates that.
Scrape#
The chart’s ServiceMonitor is enabled by the Terraform’s karpenter module.
It drops the offering price estimates and per-instance-type CPU and memory,
which describe every instance type in the region rather than the cluster,
and keeps offering availability only for the node pools’ instance types.
Karpenter’s generic families (controller_runtime_*, workqueue_*,
client_go_*, aws_sdk_go_*) carry no karpenter_ prefix, so they are
scoped by app="karpenter", which the gateway sets from the pod’s name label.
Names ending in _count#
karpenter_scheduler_unschedulable_pods_count and the
operator_*_status_condition_count families are gauges, not the counts of a
histogram.
Rare events#
NodeClaims are created and removed a few times a day, so their duration
histograms are read over an hour rather than $__rate_interval, which would
leave the panel mostly empty.
infra.karpenter.health.nodepools_not_ready #
NodePools that are not Ready.count(
max by (name) (operator_nodepool_status_condition_count{type="Ready", status!="True"}) == 1
)
or 0 * count(max by (name) (operator_nodepool_status_condition_count{type="Ready"}))
infra.karpenter.health.nodes #
Nodes Karpenter manages, in the selected NodePools.count(
max by (node_name) (karpenter_nodes_allocatable{resource_type="cpu", nodepool!="", nodepool=~"$karpenterNodePool"})
)
infra.karpenter.health.pods_unplaceable #
Pods Karpenter could not find or create a node for in its last scheduling pass.max(karpenter_scheduler_unschedulable_pods_count{controller="provisioner"})
infra.karpenter.health.launch_errors #
Instance launches EC2 refused in the selected time range.sum(increase(karpenter_cloudprovider_errors_total{method="Create"}[1h]))
or 0 * max(karpenter_build_info)
infra.karpenter.health.disrupted #
NodeClaims Karpenter removed or replaced in the selected time range, in the selected NodePools.sum(increase(karpenter_nodeclaims_disrupted_total{nodepool=~"$karpenterNodePool"}[1h]))
or 0 * max(karpenter_build_info)
infra.karpenter.health.synced #
Whether Karpenter’s view of the cluster’s nodes and pods is in step with Kubernetes.max(karpenter_cluster_state_synced)
infra.karpenter.health.version #
The Karpenter version running.max by (version) (karpenter_build_info)
infra.karpenter.nodes.by_nodepool #
Karpenter’s nodes in each NodePool over time.count by (nodepool) (
max by (node_name, nodepool) (karpenter_nodes_allocatable{resource_type="cpu", nodepool!="", nodepool=~"$karpenterNodePool"})
)
infra.karpenter.nodepools.cpu #
CPU each NodePool’s nodes provide, against the limit the NodePool sets.max by (nodepool) (karpenter_nodepools_usage{resource_type="cpu", nodepool=~"$karpenterNodePool"})
max by (nodepool) (karpenter_nodepools_limit{resource_type="cpu", nodepool=~"$karpenterNodePool"})
infra.karpenter.nodepools.memory #
Memory each NodePool’s nodes provide, against the limit the NodePool sets.max by (nodepool) (karpenter_nodepools_usage{resource_type="memory", nodepool=~"$karpenterNodePool"})
max by (nodepool) (karpenter_nodepools_limit{resource_type="memory", nodepool=~"$karpenterNodePool"})
infra.karpenter.utilization #
The share of Karpenter’s nodes’ allocatable CPU, memory and pod slots that pods have requested.max by (resource_type) (karpenter_cluster_utilization_percent{resource_type=~"cpu|memory|pods"}) / 100
infra.karpenter.conditions #
Every condition on each NodePool and EC2NodeClass, and its current status.max by (kind, name, type, status, reason) (
operator_status_condition_count{kind=~"NodePool|EC2NodeClass"}
) == 1
infra.karpenter.provisioning.created #
NodeClaims created per hour, by NodePool and reason.sum by (nodepool, reason) (
rate(karpenter_nodeclaims_created_total{nodepool=~"$karpenterNodePool"}[5m])
) * 3600
infra.karpenter.provisioning.errors #
Errors from EC2 per hour, by error type and operation.sum by (error, method) (
rate(karpenter_cloudprovider_errors_total[5m])
) * 3600
infra.karpenter.provisioning.offerings #
Zones where each of the node pools’ instance types can be launched, by capacity type.sum by (instance_type, capacity_type) (
max by (instance_type, capacity_type, zone) (karpenter_cloudprovider_instance_type_offering_available{capacity_type!="reserved"})
)
infra.karpenter.provisioning.pods_waiting #
Pods waiting for Karpenter: the ones it cannot place, and the queue for its next scheduling pass.max(karpenter_scheduler_unschedulable_pods_count{controller="provisioner"})
max by (controller) (karpenter_scheduler_queue_depth)
infra.karpenter.provisioning.startup #
How long pods that needed a new node took to start running, from when the scheduler gave up on existing nodes.histogram_quantile(0.5, sum by (le) (increase(karpenter_pods_provisioning_startup_duration_seconds_bucket[1h])))
histogram_quantile(0.99, sum by (le) (increase(karpenter_pods_provisioning_startup_duration_seconds_bucket[1h])))
infra.karpenter.provisioning.stages #
How long NodeClaims took to reach each stage of launching, at the 99th percentile over the last hour.histogram_quantile(0.99,
sum by (type, le) (
increase(operator_nodeclaim_status_condition_transition_seconds_bucket{type=~"Launched|Registered|Initialized", karpenter_sh_nodepool=~"$karpenterNodePool"}[1h])
)
)
infra.karpenter.provisioning.scheduling_duration #
How long each of Karpenter’s scheduling passes takes, at the 99th percentile.histogram_quantile(0.99,
sum by (controller, le) (rate(karpenter_scheduler_scheduling_duration_seconds_bucket[5m]))
)
infra.karpenter.provisioning.cloud_latency #
How long Karpenter’s calls to AWS take, by operation, at the 99th percentile.histogram_quantile(0.99,
sum by (method, le) (rate(karpenter_cloudprovider_duration_seconds_bucket[5m]))
)
infra.karpenter.disruption.disrupted #
NodeClaims removed or replaced per hour, by NodePool and reason.sum by (nodepool, reason) (
rate(karpenter_nodeclaims_disrupted_total{nodepool=~"$karpenterNodePool"}[5m])
) * 3600
infra.karpenter.disruption.decisions #
Disruptions Karpenter decided on per hour, by reason and action.sum by (reason, decision) (
rate(karpenter_voluntary_disruption_decisions_total[5m])
) * 3600
infra.karpenter.disruption.eligible #
Nodes Karpenter considers candidates for disruption, by reason.max by (reason) (karpenter_voluntary_disruption_eligible_nodes)
infra.karpenter.disruption.budget #
Nodes each NodePool’s disruption budget allows removing right now, by reason.min by (nodepool, reason) (karpenter_nodepools_allowed_disruptions{nodepool=~"$karpenterNodePool"})
infra.karpenter.disruption.blocked #
Why Karpenter is leaving nodes in place, from itsDisruptionBlocked
and Unconsolidatable events per five minutes.infra.karpenter.disruption.blocked_stream #
Karpenter’s blocked-disruption events, newest first.infra.karpenter.disruption.evictions #
Pod evictions Karpenter requested per minute, by response.sum by (code) (rate(karpenter_nodes_eviction_requests_total[5m])) * 60
infra.karpenter.disruption.termination #
How long NodeClaims took to terminate, from the decision to the instance gone, at the 99th percentile over the last hour.histogram_quantile(0.99,
sum by (le) (increase(operator_nodeclaim_termination_duration_seconds_bucket{karpenter_sh_nodepool=~"$karpenterNodePool"}[1h]))
)
infra.karpenter.disruption.consolidation_timeouts #
Consolidation searches that ran out of time, per hour.sum by (consolidation_type) (
rate(karpenter_voluntary_disruption_consolidation_timeouts_total[5m])
) * 3600
infra.karpenter.nodes.age #
Each of Karpenter’s nodes, with its NodePool, instance type, zone and age.max by (node_name, nodepool, instance_type, zone) (
karpenter_nodes_current_lifetime_seconds{nodepool!="", nodepool=~"$karpenterNodePool"}
)
infra.karpenter.interruption.received #
Notices Karpenter read from its interruption queue per hour, by type.sum by (message_type) (
rate(karpenter_interruption_received_messages_total[5m])
) * 3600
infra.karpenter.interruption.queue_delay #
How long notices waited in the interruption queue before Karpenter read them, at the 99th percentile over the last hour.histogram_quantile(0.99,
sum by (le) (increase(karpenter_interruption_message_queue_duration_seconds_bucket[1h]))
)
infra.karpenter.controller.reconcile_errors #
Reconcile errors per minute, by controller.sum by (controller) (
rate(controller_runtime_reconcile_errors_total{app="karpenter"}[5m])
) * 60 > 0
infra.karpenter.controller.workqueue #
Items waiting in each of Karpenter’s work queues.max by (name) (workqueue_depth{app="karpenter"}) > 0
infra.karpenter.controller.aws_errors #
AWS API calls that did not succeed per minute, by service, action and status.sum by (exported_service, action, code) (
rate(aws_sdk_go_request_total{app="karpenter", code!~"2..|412"}[5m])
) * 60 > 0
infra.karpenter.controller.kube_errors #
Kubernetes API calls that did not succeed per minute, by method and status.sum by (method, code) (
rate(client_go_request_total{app="karpenter", code!~"2.."}[5m])
) * 60 > 0
infra.karpenter.controller.leader #
Which Karpenter replica holds the leader lease.max by (pod) (leader_election_master_status{app="karpenter"})
infra.karpenter.controller.cpu #
CPU each Karpenter replica uses, in cores.sum by (pod) (
rate(container_cpu_usage_seconds_total{container="controller", pod=~"karpenter-.+"}[5m])
)
infra.karpenter.controller.memory #
Memory each Karpenter replica holds, against its limit.max by (pod) (container_memory_working_set_bytes{container="controller", pod=~"karpenter-.+"})
max by (pod) (kube_pod_container_resource_limits{container="controller", pod=~"karpenter-.+", resource="memory"})
infra.karpenter.events.rate_by_reason #
Karpenter’s events per minute, by reason.infra.karpenter.events.stream #
Karpenter’s events, newest first.infra.karpenter.logs.rate #
Karpenter’s log lines per minute, by level.infra.karpenter.logs.problems #
Karpenter’s warning and error log lines, newest first.infra-logs#
Logs from the platform a Materialize deployment runs on: the monitoring stack itself, the Kubernetes system components, and the nodes underneath both.
Separate from materialize-logs.yaml rather than a widening of it, for two
reasons that are structural rather than stylistic.
The selector set differs. These carry component and container filters.
component is what tells one Loki or Thanos process from another — loki
alone splits into canary, querier, ingester, query-frontend,
index-gateway, compactor, distributor and ruler — and container is
the only picker that reaches workloads with no app at all, which on a
representative install is the whole of kube-system. Neither dimension means
anything to a Materialize environment, and adding them to the shared queries
would oblige every dashboard using those to define pickers it has no use for.
The node journal is not reachable from a namespace. Journal lines carry
unit, component, job, level and service_name and no namespace,
app or container, because they come from the node rather than from a pod.
Any selector that requires a namespace excludes them by construction, which is
why they have their own queries here and their own tab on the dashboard.
The Kubernetes-event queries are not duplicated: materialize.events.cluster.*
in materialize-events.yaml is already scoped by the same namespace picker and
carries no Materialize-specific filter, so an infrastructure dashboard uses it
as it stands.
infra.logs.stream #
The log feed for the selected namespaces, apps, components, containers and levels, newest first.infra.logs.warnings.stream #
Warning-and-worse lines from the platform, newest first.infra.logs.rate.by_component #
Log lines per second by application and sub-component — which process of which workload is doing the talking.infra.logs.rate.by_namespace #
Log lines per second by namespace — where in the cluster the volume is.infra.logs.warnings.rate #
Warning-and-worse lines per minute across the platform, as one series.infra.logs.node.stream #
The node journal, newest first —kubelet, containerd, the node
problem detector, and the rest of what systemd runs on each node.infra.logs.node.warnings #
Warning-and-worse lines from the node journal.infra.logs.node.rate.by_unit #
Node journal lines per second by systemd unit.infra-loki#
Is the log store healthy — and if not, which half of it broke.
Every other logs query in this repository reads Loki. These watch it. That makes this file the one place where the monitoring stack is the subject rather than the instrument, and it changes what a reading means: a blank panel here is itself a finding, because the thing that would have reported the problem is the thing that is down.
The audience is whoever operates the monitoring stack. Loki’s vocabulary is not assumed — a stream is one label-set’s worth of lines, a chunk is a stream’s lines batched up for storage, and the write path turns the first into the second before handing it to object storage. Where a panel needs one of those to be legible, its description says so.
Two unrelated loki_ metric families#
The prefix is shared and the producers are not. Reading one for the other is the single easiest mistake to make here, and nothing about the metric name warns you.
- The log store.
job=~"loki/.*", and everything this file’s panels draw. 391 metric names on a reference install. - Alloy’s
loki.*components.job=~"alloy-.*", 43 names, all underloki_write_*,loki_source_{file,api,journal}_*,loki_process_*andloki_relabel_*. These describe the collector — what Alloy read off disk and shipped — not the store that received it.
Only loki_experimental_features_in_use_total is emitted by both.
The two are kept apart by app_instance="loki", which both of this chart’s
Loki ServiceMonitors stamp on every target they collect and nothing else sets.
A job=~"loki/.*" matcher would work too, but it hardcodes the subchart’s
jobPrefix, which is a value an operator may change.
The Alloy-side family belongs to infra-alloy.yaml, which watches the
collectors and is where those metrics are drawn.
Scoping#
Two variables, both written literally rather than reaching through a render
parameter, on the same precedent as $nodeList in node-health.yaml and
$namespaceList in infra-networking.yaml: a meta-monitoring dashboard’s
scope is the monitoring stack’s own, which no Materialize-shaped parameter
describes.
$lokiNamespace— where Loki runs. Every metric query here carries it, so twomaterialize-monitoringreleases sharing a cluster do not read as one store.$lokiComponent— which Loki process. Discovered from the metrics side, off thecontainerlabel, and applied to the log queries ascomponent=~"$lokiComponent"as well.
One picker across both engines is a deliberate call, because the two label sets
agree by construction: container on a metric and component on a log line
are both the Kubernetes container name. The one value that exists on the
metrics side and not the logs side is exporter, the memcached sidecar, which
produces no log lines at all — so selecting it empties the log feeds, which is
the honest answer rather than a bug.
Plaintext exporters#
Three of Loki’s targets never serve TLS, whatever Loki itself is configured to
do: the canary’s own /metrics, and the two memcached exporters. They are
collected by a separate ServiceMonitor for that reason
(templates/scrapers/monitor-loki-plaintext.yaml), and infra.loki.health.up
is the panel that says whether that split is working. It was not, before this
file existed: under profiles/mtls all three were scraped over HTTPS, failed,
and vanished — taking the end-to-end canary with them.
infra.loki.health.canary.missing #
Log lines the canary wrote to Loki and then could not read back, per second.sum(
rate(
loki_canary_missing_entries_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.health.canary.entries #
Log lines per second the canary has written and successfully read back.sum(
rate(
loki_canary_entries_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.health.canary.latency #
How long the canary waited between writing a line and reading it back, at the median and the 99th percentile.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_canary_response_latency_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
histogram_quantile(
0.99,
sum by (le) (
rate(
loki_canary_response_latency_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.health.canary.losses #
Canary lines that went missing, split by when they went missing.sum(
rate(
loki_canary_missing_entries_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
sum(
rate(
loki_canary_spot_check_missing_entries_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.health.up #
Whether each Loki target is being scraped successfully, one series per component.max by (container, service) (
up{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent"
}
)
infra.loki.health.request_errors #
The share of Loki API requests returning a 5xx, as a proportion of all requests.sum(
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent",
status_code=~"5.."
}
[5m]
)
)
/
sum(
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent"
}
[5m]
)
)
infra.loki.health.client_errors #
The share of Loki API requests rejected as the caller’s fault — a 4xx — as a proportion of all requests.sum(
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent",
status_code=~"4.."
}
[5m]
)
)
/
sum(
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent"
}
[5m]
)
)
infra.loki.health.request_failures #
Requests per second that failed without producing an HTTP status, split by component.sum by (container) (
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent",
status_code="error"
}
[5m]
)
)
infra.loki.health.request_latency #
99th-percentile Loki API latency, split by route.histogram_quantile(
0.99,
sum by (le, route) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
container=~"$lokiComponent",
route!~"(?i).*tail.*|/schedulerpb.SchedulerForQuerier/QuerierLoop"
}
[5m]
)
)
)
infra.loki.health.restarts #
Container restarts across the Loki components over the dashboard’s time range.sum by (pod) (
increase(
kube_pod_container_status_restarts_total{
namespace=~"$lokiNamespace",
pod=~"loki-.*"
}
[1h]
)
) > 0
infra.loki.write.lines #
Log lines per second arriving at the distributor — everything the cluster is sending Loki.sum(
rate(
loki_distributor_lines_received_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.write.bytes #
Bytes per second arriving at the distributor, before compression.sum(
rate(
loki_distributor_bytes_received_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.write.discarded #
Log lines per second Loki accepted the connection for and then threw away, split by the reason it gave.sum by (reason) (
rate(
loki_discarded_samples_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.write.discarded_bytes #
Bytes per second discarded, split by reason — the same losses as the line panel, weighted by how much was lost.sum by (reason) (
rate(
loki_discarded_bytes_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.write.streams #
Active streams held in memory by the ingesters. A stream is one distinct label-set — one container’s lines at one level, say.sum by (pod) (
loki_ingester_memory_streams{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.write.chunks_in_memory #
Chunks each ingester is holding before flushing them to object storage.sum by (pod) (
loki_ingester_memory_chunks{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.write.chunks_flushed #
Chunks per second the ingesters have written out to object storage.sum(
rate(
loki_ingester_chunks_flushed_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.write.chunk_utilization #
How full a chunk was when it was flushed, at the median and 10th percentile. 1.0 is a chunk closed because it filled up.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_ingester_chunk_utilization_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
histogram_quantile(
0.10,
sum by (le) (
rate(
loki_ingester_chunk_utilization_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.write.chunk_age #
How old a chunk was when it was flushed, at the median and 99th percentile.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_ingester_chunk_age_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
histogram_quantile(
0.99,
sum by (le) (
rate(
loki_ingester_chunk_age_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.write.wal_disk #
How much of its write-ahead log volume each ingester has used, as a percentage.max by (pod) (
loki_ingester_wal_disk_usage_percent{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.write.wal_disk_full #
Times an ingester could not write to its write-ahead log because the volume was full.sum by (pod) (
increase(
loki_ingester_wal_disk_full_failures_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[1h]
)
)
infra.loki.write.wal_replay #
Whether an ingester is currently replaying its write-ahead log. 1 means replaying.max by (pod) (
loki_ingester_wal_replay_active{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.write.push_latency #
How long a push request took, at the median and 99th percentile.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
route=~"loki_api_v1_push|api_prom_push|/logproto.Pusher/Push"
}
[5m]
)
)
)
histogram_quantile(
0.99,
sum by (le) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
route=~"loki_api_v1_push|api_prom_push|/logproto.Pusher/Push"
}
[5m]
)
)
)
infra.loki.read.query_rate #
Queries per second reaching the query frontend, split by the kind of request.sum by (route) (
rate(
loki_request_duration_seconds_count{
app_instance="loki",
namespace=~"$lokiNamespace",
container="query-frontend",
route=~"loki_api_v1_.*"
}
[5m]
)
)
infra.loki.read.query_latency #
End-to-end query latency at the query frontend, at the median, 90th and 99th percentile.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
container="query-frontend",
route=~"loki_api_v1_.*",
route!~"(?i).*tail.*"
}
[5m]
)
)
)
histogram_quantile(
0.90,
sum by (le) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
container="query-frontend",
route=~"loki_api_v1_.*",
route!~"(?i).*tail.*"
}
[5m]
)
)
)
histogram_quantile(
0.99,
sum by (le) (
rate(
loki_request_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace",
container="query-frontend",
route=~"loki_api_v1_.*",
route!~"(?i).*tail.*"
}
[5m]
)
)
)
infra.loki.read.queue_duration #
How long a query sub-request waited in the scheduler’s queue before a querier picked it up, at the median and 99th percentile.histogram_quantile(
0.50,
sum by (le) (
rate(
loki_query_scheduler_queue_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
histogram_quantile(
0.99,
sum by (le) (
rate(
loki_query_scheduler_queue_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.read.queries_in_flight #
Queries the frontend is currently working on.sum(
loki_query_frontend_queries_in_progress{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.read.bytes_processed #
Bytes per second of stored log data the queriers had to decompress and scan to answer the queries being asked.sum(
rate(
loki_logql_querystats_bytes_processed_per_seconds_sum{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.read.cache_requests #
Cache lookups per second, split by which cache — chunks, query results and the index each have their own.sum by (name) (
rate(
loki_cache_fetched_keys{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.read.cache_hit_rate #
The share of cache lookups that were served from cache, split by which cache.sum by (name) (
rate(
loki_cache_hits{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
/
sum by (name) (
rate(
loki_cache_fetched_keys{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.read.memcached_memory #
How much of its allotted memory each memcached cache is holding, as a proportion of its limit.sum by (service) (
memcached_current_bytes{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
/
sum by (service) (
memcached_limit_bytes{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.store.operations #
Object-storage operations per second, split by kind.sum by (operation) (
rate(
loki_objstore_bucket_operations_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.store.failures #
Object-storage operations per second that failed, split by kind.sum by (operation) (
rate(
loki_objstore_bucket_operation_failures_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.store.latency #
99th-percentile object-storage operation latency, split by kind.histogram_quantile(
0.99,
sum by (le, operation) (
rate(
loki_objstore_bucket_operation_duration_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.store.index_sync #
Index table sync and upload operations per second between the shippers and object storage.sum(
rate(
loki_tsdb_shipper_tables_sync_operation_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
sum(
rate(
loki_tsdb_shipper_tables_upload_operation_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
infra.loki.store.index_wait #
How long a query waited for an index table to be downloaded before it could run, at the 99th percentile.histogram_quantile(
0.99,
sum by (le) (
rate(
loki_tsdb_shipper_query_wait_time_seconds_bucket{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[5m]
)
)
)
infra.loki.store.retention_age #
How long ago the compactor last finished applying retention successfully.time()
-
max(
loki_compactor_apply_retention_last_successful_run_timestamp_seconds{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.store.compaction_age #
How long ago the compactor last finished compacting index tables successfully.time()
-
max(
loki_boltdb_shipper_compact_tables_operation_last_successful_run_timestamp_seconds{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.store.compaction_runs #
Compactor runs over the dashboard’s time range, split by outcome.sum by (status) (
increase(
loki_boltdb_shipper_compact_tables_operation_total{
app_instance="loki",
namespace=~"$lokiNamespace"
}
[1h]
)
)
infra.loki.store.deletes_pending #
Delete requests waiting to be processed, and how old the oldest one is.sum(
loki_compactor_pending_delete_requests_count{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
max(
loki_compactor_oldest_pending_delete_request_age_seconds{
app_instance="loki",
namespace=~"$lokiNamespace"
}
)
infra.loki.logs.stream #
Loki’s own log feed, newest first, for the selected components and levels.infra.loki.logs.warnings.stream #
Warning-and-worse lines from Loki, newest first.infra.loki.logs.rate.by_component #
Loki’s own log lines per second, split by component.infra.loki.logs.warnings.rate #
Warning-and-worse lines per minute from Loki, across the selected components.infra-networking#
How traffic moves through the cluster, and what stops it.
Where infra-nodes asks what one machine is doing, this asks what the
network is doing across all of them: what pods are sending, what Kubernetes
is routing, what the CNI is doing underneath, and which packets a policy
dropped.
Three families meet here, and knowing which is which is most of what makes this file readable:
cAdvisor (
container_network_*) is the only per-pod view. It is always present, because the gateway scrapes every kubelet. Note that its series split in two by whethernamespaceis set: with a namespace they describe a pod, and without one they carryid="/"and describe the node’s root cgroup — every interface on the machine, including the veth of every pod on it. Panels must pick one and say which.Host-network pods are excluded from every rollup here, and must be. cAdvisor reads a container’s counters from its network namespace, and a host-network pod’s namespace is the node’s — so such a pod reports all 31 of a reference node’s interfaces rather than its own one, and its “pod traffic” is the whole machine’s, counting every other pod’s veth. Several DaemonSets do this on every node. Left in, a cluster-wide sum read 27,831 KiB/s against a true 1,700, and the ten busiest pods were all host-network DaemonSets reporting their node.
%%{excludeHostNetworkPods}is theunlessclause that drops them; it appends to asum by (namespace, pod)and cannot be a label matcher, because being host-network is a property of a pod’s whole series set rather than of any one series.kube-state-metrics (
kube_service_*,kube_endpointslice_*,kube_networkpolicy_*) is the declared intent: what Services exist, what is behind them, and which policies were written. It says nothing about whether any of it works.The CNI is the only family that can say what the dataplane actually did — which packet a policy dropped, how many addresses are left, whether the datapath is erroring. It is also the only family that is not guaranteed to be there, which is why every query reading it is on a row the dashboard shows only when that vendor was detected. See “Dataplane detection” below.
node_* is a fourth family and deliberately not redefined here:
node-health.yaml and node-debug.yaml already carry it, every one of their
expressions is scoped by instance=~"$nodeList", and infra-net defines that
variable as a multi-select across the fleet rather than the single address
infra-nodes resolves it to. The same expressions therefore answer for one
node or for all of them depending only on the dashboard that asks.
Dataplane detection#
A cluster’s CNI is not something this repository can know: EKS defaults to the
AWS VPC CNI, GKE to Dataplane V2, AKS to Azure CNI powered by Cilium, and a
bring-your-own cluster to anything at all. Rather than ask the operator, the
CNI monitors in packages/prometheus-scrapers/ stamp every series they
collect with a network_component label, and the dashboard discovers the
answer with label_values(up{network_component=~".+"}, network_component).
That puts the vendor knowledge in the scrape config, where it already had to live, and leaves the queries below reading ordinary vendor metric names.
Two conventions apply throughout:
Scope literally, not by parameter.
namespace=~"$namespaceList"andinstance=~"$nodeList"are written out, the same wayinfra-nodes.yamlwritesnode="$node". A dashboard using these must define the variables; no render parameter supplies them.$nodeListholds node-exporter addresses, inherited from the node query families. Nothing here scopes on anodelabel, which every other family spells as the Kubernetes name – the two cannot be filtered by one picker, and the node-exporter form is the one that buys the 87 vetted expressions innode-health.yamlandnode-debug.yaml.Deduplicate across scrape replicas.
instanceis the scrape target, so a baresumover an HA kube-state-metrics adds each object once per replica. Aggregations keepinstancein the inner step and collapse it with an outermax.
%%{interval} is the rate window, including its brackets.
infra.net.overview.dataplane #
Which networking components this cluster is running, as discovered rather than configured.max by (network_component) (up{network_component=~".+"})
infra.net.overview.throughput #
Total bytes per second in and out of the cluster’s pods.sum (
sum by (namespace, pod) (
rate(container_network_receive_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
) ${excludeHostNetworkPods}
)
sum (
sum by (namespace, pod) (
rate(container_network_transmit_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
) ${excludeHostNetworkPods}
)
infra.net.overview.errors #
Interface errors and dropped packets across every pod in the cluster.sum (
sum by (namespace, pod) (
rate(container_network_receive_errors_total{namespace=~"$namespaceList"}[5m])
+ rate(container_network_transmit_errors_total{namespace=~"$namespaceList"}[5m])
) ${excludeHostNetworkPods}
)
sum (
sum by (namespace, pod) (
rate(container_network_receive_packets_dropped_total{namespace=~"$namespaceList"}[5m])
+ rate(container_network_transmit_packets_dropped_total{namespace=~"$namespaceList"}[5m])
) ${excludeHostNetworkPods}
)
infra.net.overview.top_talkers #
The ten pods moving the most traffic, in and out combined.topk(10,
sum by (namespace, pod) (
rate(container_network_receive_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
+ rate(container_network_transmit_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
) ${excludeHostNetworkPods}
)
infra.net.k8s.throughput.by_namespace #
Pod traffic split by namespace, received and transmitted.sum by (namespace) (
sum by (namespace, pod) (
rate(container_network_receive_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
) ${excludeHostNetworkPods}
)
sum by (namespace) (
sum by (namespace, pod) (
rate(container_network_transmit_bytes_total{namespace=~"$namespaceList", interface!="lo"}[5m])
) ${excludeHostNetworkPods}
)
infra.net.k8s.errors.by_namespace #
Which namespace’s pods are losing packets.sum by (namespace) (
sum by (namespace, pod) (
rate(container_network_receive_errors_total{namespace=~"$namespaceList"}[5m])
+ rate(container_network_transmit_errors_total{namespace=~"$namespaceList"}[5m])
+ rate(container_network_receive_packets_dropped_total{namespace=~"$namespaceList"}[5m])
+ rate(container_network_transmit_packets_dropped_total{namespace=~"$namespaceList"}[5m])
) ${excludeHostNetworkPods}
)
infra.net.k8s.services.by_type #
Services in the cluster, by how they are exposed.max by (type) (
sum by (instance, type) (kube_service_spec_type)
)
infra.net.k8s.endpoints.by_namespace #
How many endpoints back the Services in each namespace.max by (namespace) (
sum by (instance, namespace) (kube_endpointslice_endpoints{namespace=~"$namespaceList"})
)
infra.net.k8s.proxy.sync_latency #
How long kube-proxy takes to turn a Service change into rules on the node.histogram_quantile(0.99,
sum by (le) (
rate(kubeproxy_sync_proxy_rules_duration_seconds_bucket[5m])
)
)
infra.net.k8s.proxy.programming_latency #
End-to-end time from an endpoint changing to the node’s rules reflecting it.histogram_quantile(0.99,
sum by (le) (
rate(kubeproxy_network_programming_duration_seconds_bucket[5m])
)
)
infra.net.cni.aws.ip_utilization #
How much of each node’s pod-IP allowance is in use.max by (node) (awscni_assigned_ip_addresses)
/
max by (node) (awscni_ip_max)
infra.net.cni.aws.addresses #
Addresses assigned to pods on each node, against the pool the CNI is holding and the ceiling it cannot pass.max by (node) (awscni_assigned_ip_addresses)max by (node) (awscni_total_ip_addresses)max by (node) (awscni_ip_max)infra.net.cni.aws.enis #
Network interfaces attached to each node, against the most it can have.max by (node) (awscni_eni_allocated)max by (node) (awscni_eni_max)infra.net.cni.aws.exhaustion #
Nodes that have run out of pod addresses outright.sum (max by (node) (awscni_no_available_ip_addresses))
infra.net.cni.aws.api_latency #
How long the CNI’s calls to the EC2 API are taking, by operation.sum by (api) (rate(awscni_aws_api_latency_ms_sum[5m]))
/
sum by (api) (rate(awscni_aws_api_latency_ms_count[5m]))
infra.net.cni.cilium.endpoints #
Cilium endpoints per node, by state. An endpoint is roughly a pod as the dataplane sees it.sum by (node, endpoint_state) (cilium_endpoint_state)
infra.net.cni.cilium.drops #
Packets the Cilium datapath dropped, by the reason it gives.sum by (reason) (rate(cilium_drop_count_total[5m]))
infra.net.cni.cilium.policy_verdicts #
Policy decisions the dataplane made, allowed against denied.sum by (action) (rate(cilium_policy_verdict_total[5m]))
infra.net.cni.cilium.bpf_map_pressure #
How full Cilium’s BPF maps are, as a fraction of their capacity.max by (node, map_name) (cilium_bpf_map_pressure)
infra.net.cni.cilium.endpoint_regeneration #
How long Cilium takes to reprogram an endpoint’s datapath, at the 99th percentile.histogram_quantile(0.99,
sum by (le, scope) (
rate(cilium_endpoint_regeneration_time_stats_seconds_bucket[5m])
)
)
infra.net.cni.cilium.unreachable #
Nodes and health endpoints Cilium’s own connectivity probes cannot reach.sum (max by (node) (cilium_unreachable_nodes))sum (max by (node) (cilium_unreachable_health_endpoints))infra.net.cni.cilium.hubble_flows #
Network flows Hubble observed, by verdict.sum by (verdict) (rate(hubble_flows_processed_total[5m]))
infra.net.security.policies.by_namespace #
How many NetworkPolicy objects each namespace has.max by (namespace) (
count by (instance, namespace) (kube_networkpolicy_spec_ingress_rules{namespace=~"$namespaceList"})
)
infra.net.security.policies.uncovered #
Namespaces that run pods and have no NetworkPolicy.count by (namespace) (kube_pod_info{namespace=~"$namespaceList"})
unless on (namespace)
count by (namespace) (kube_networkpolicy_spec_ingress_rules)
infra.net.security.policies.rules #
Each policy’s ingress and egress rule counts, one row per policy.max by (namespace, networkpolicy) (kube_networkpolicy_spec_ingress_rules{namespace=~"$namespaceList"})
max by (namespace, networkpolicy) (kube_networkpolicy_spec_egress_rules{namespace=~"$namespaceList"})
infra.net.security.drops.aws #
Packets the AWS network policy agent dropped for violating a policy.sum by (node) (rate(network_policy_drop_count_total[5m]))
infra.net.security.drops.cilium #
Packets Cilium dropped specifically because a policy denied them.sum by (node) (rate(cilium_drop_count_total{reason=~"Policy denied.*"}[5m]))
infra.net.cloud.load_balancers #
Every Service in the cluster that has a cloud load balancer, and the address it was given.max by (namespace, service, ip, hostname) (kube_service_status_load_balancer_ingress)
infra-nodes#
What kubectl describe node would tell you, for whoever cannot run it.
These answer the questions an operator asks about one machine: what is it,
how big is it, how much of it is already promised, is Kubernetes willing to
put work on it, and what has it been saying. The measurements of what the
machine is actually doing live in node-health.yaml and node-debug.yaml,
which read node-exporter; this file is the Kubernetes side, plus the node’s
own journal and the events filed against it.
Two identifier conventions meet here, and the difference is the thing to know:
kube-state-metrics names a node
node="<kubernetes name>". Every query in this file scopes withnode="$node"written literally, the same way the node-exporter families writeinstance=~"$nodeList"literally. A dashboard using either must define the variable; no render parameter supplies it.node-exporter names the same machine
instance="<ip>:9100". The join isnode_uname_info, whosenodenameis the Kubernetes name — which is what the$nodeListvariable resolves through. Nothing in this file needs the join, because nothing in this file reads node-exporter.
Loki knows the node a third way: node is structured metadata on journal
lines, not a stream label, so it is filtered in the pipeline (| node=...)
rather than in the selector. Node events are Kubernetes events whose involved
object is the node itself, which is kind="Node" with the node’s name.
Two conventions apply to every query here that aggregates:
Deduplicate across kube-state-metrics replicas.
instanceis the scrape target, so a baresumorcountover an HA deployment adds each object once per replica. Every aggregation keepsinstancein the inner step and collapses it with an outermax, which is the shapematerialize-kubernetes.yamlestablished. Queries that only evermaxare already safe, sincemaxacross identical replicas is idempotent.Terminal pods do not count against the node. A
SucceededorFailedpod has released its CPU and memory, and the scheduler no longer counts it against the pod limit — kube-state-metrics agrees, and stops reportingkube_pod_container_resource_*for it. Butkube_pod_infokeeps reporting it until garbage collection, so anything counting pods rather than their resources has to subtract them explicitly or it overstates how full the node is. Completed Jobs are the common case.
%%{interval} is the rate window, including its brackets.
infra.nodes.info.kubelet #
The kubelet version this node runs, which is the version Kubernetes itself is on here.max by (node, kubelet_version) (kube_node_info{node="$node"})
infra.nodes.info.os #
The node’s operating system image.max by (node, os_image) (kube_node_info{node="$node"})
infra.nodes.info.kernel #
The node’s kernel version.max by (node, kernel_version) (kube_node_info{node="$node"})
infra.nodes.info.runtime #
The container runtime that starts and stops containers on this node.max by (node, container_runtime_version) (kube_node_info{node="$node"})
infra.nodes.info.address #
The address the cluster reaches this node on.max by (node, internal_ip) (kube_node_info{node="$node"})
infra.nodes.created #
Wall-clock time the node joined the cluster.max by (node) (kube_node_created{node="$node"}) * 1000
infra.nodes.capacity.cpu #
Cores the node reports to Kubernetes.max by (node) (kube_node_status_capacity{node="$node", resource="cpu"})
infra.nodes.capacity.memory #
Bytes of RAM the node reports to Kubernetes.max by (node) (kube_node_status_capacity{node="$node", resource="memory"})
infra.nodes.capacity.pods #
The most pods Kubernetes will place on this node.max by (node) (kube_node_status_capacity{node="$node", resource="pods"})
infra.nodes.capacity.ephemeral_storage #
Bytes of node-local disk available to pods for scratch space.max by (node) (kube_node_status_capacity{node="$node", resource="ephemeral_storage"})
infra.nodes.allocation.cpu #
Fraction of the node’s schedulable CPU already promised to pods through their requests.max by (node) (
sum by (node, instance) (
kube_pod_container_resource_requests{node="$node", resource="cpu"}
)
)
/
max by (node) (kube_node_status_allocatable{node="$node", resource="cpu"})
infra.nodes.allocation.memory #
Fraction of the node’s schedulable memory already promised to pods through their requests.max by (node) (
sum by (node, instance) (
kube_pod_container_resource_requests{node="$node", resource="memory"}
)
)
/
max by (node) (kube_node_status_allocatable{node="$node", resource="memory"})
infra.nodes.allocation.pods #
Fraction of the node’s pod slots in use.max by (node) (
count by (node, instance) (
kube_pod_info{node="$node"}
unless on (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"} == 1)
)
)
/
max by (node) (kube_node_status_allocatable{node="$node", resource="pods"})
infra.nodes.pods.by_namespace #
What is actually running on this node, grouped by namespace.max by (namespace) (
count by (namespace, instance) (
kube_pod_info{node="$node"}
unless on (namespace, pod) (kube_pod_status_phase{phase=~"Succeeded|Failed"} == 1)
)
)
infra.nodes.condition.ready #
Whether Kubernetes considers the node healthy enough to run work.max by (node) (kube_node_status_condition{node="$node", condition="Ready", status="true"})
infra.nodes.conditions #
The node’s pressure and availability conditions — memory, disk, PIDs and network — each 1 when the condition is active.max by (node, condition) (
kube_node_status_condition{
node="$node",
condition=~"MemoryPressure|DiskPressure|PIDPressure|NetworkUnavailable",
status="true"
}
)
infra.nodes.unschedulable #
Whether the node has been cordoned against new work.max by (node) (kube_node_spec_unschedulable{node="$node"})
infra.nodes.taints #
The taints on this node, which restrict what may be scheduled.max by (node, key, value, effect) (kube_node_spec_taint{node="$node"})
infra.nodes.pods.by_phase #
Pods on this node, counted by lifecycle phase.max by (phase) (
count by (phase, instance) (
kube_pod_status_phase == 1
and on (namespace, pod) kube_pod_info{node="$node"}
)
)
infra.nodes.pods.not_ready #
Pods on this node that are not reporting Ready.max by (namespace, pod) (
kube_pod_status_ready{condition="true"}
and on (namespace, pod) kube_pod_info{node="$node"}
) == 0
unless on (namespace, pod) (kube_pod_status_phase{phase="Succeeded"} == 1)
infra.nodes.pods.restarts #
Container restarts for pods on this node.max by (namespace, pod) (
sum by (namespace, pod, instance) (
kube_pod_container_status_restarts_total
and on (namespace, pod) kube_pod_info{node="$node"}
)
)
infra.nodes.pods.budgets #
What each pod on this node reserved and what it is capped at — CPU and memory, requests beside limits.max by (namespace, pod) (
sum by (namespace, pod, instance) (
kube_pod_container_resource_requests{node="$node", resource="cpu"}
)
)
max by (namespace, pod) (
sum by (namespace, pod, instance) (
kube_pod_container_resource_limits{node="$node", resource="cpu"}
)
)
max by (namespace, pod) (
sum by (namespace, pod, instance) (
kube_pod_container_resource_requests{node="$node", resource="memory"}
)
)
max by (namespace, pod) (
sum by (namespace, pod, instance) (
kube_pod_container_resource_limits{node="$node", resource="memory"}
)
)
infra.nodes.journal.rate.by_unit #
Journal lines per second from this node, split by systemd unit.infra.nodes.journal.warnings #
Warning-and-worse journal lines from this node.infra.nodes.journal.stream #
The systemd journal from this node, newest first.infra.nodes.events.rate.by_reason #
Kubernetes events filed against this node, by reason.infra.nodes.events.stream #
Kubernetes events filed against this node, newest first.materialize-clusters#
Inventory and sizing of a Materialize deployment’s clusters and replicas. Adapted from the Overview dashboard’s “Cluster Objects / Replicas” tab.
materialize.clusters.count #
How many clusters exist, split into the Materialize-managed system clusters (mz_catalog_server, mz_system, mz_probe, …) that every environment has and the user clusters you created. The gap between the two is your own footprint.count(
group by (compute_cluster_id) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*"}
)
)
count(
group by (compute_cluster_id) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*", compute_cluster_id=~"^s.*"}
)
)
count_not_null(
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}} by {compute_cluster_id}
)
count_not_null(
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}, compute_cluster_id:s*} by {compute_cluster_id}
)
materialize.clusters.replicas.count #
How many replicas back the selected clusters, and how many of those are redundancy beyond the first. Every cluster needs one replica to run; anything above that is capacity or availability headroom you’ve opted into, so a non-zero “additional” count is the quick check that HA is actually configured where you expect it.count(
group by (compute_cluster_id, compute_replica_id) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*", compute_replica_id=~".*"}
)
)
count(
group by (compute_cluster_id, compute_replica_id) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*", compute_replica_id=~".*", compute_replica_name!="r1"}
)
)
count_not_null(
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}, compute_replica_id:${mzReplicaList}} by {compute_cluster_id,compute_replica_id}
)
default_zero(
count_not_null(
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}, compute_replica_id:${mzReplicaList}, !compute_replica_name:r1} by {compute_cluster_id,compute_replica_id}
)
)
materialize.clusters.replicas.sizes #
The replica fleet grouped by configured size. Most deployments settle on a handful of sizes; a long tail of one-off sizes usually means an experiment or a half-finished migration. The total agrees with the replica count.count by (size) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*", compute_replica_id=~".*"}
)
sum:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}, compute_replica_id:${mzReplicaList}} by {size}
materialize.clusters.info #
A reference row per (cluster, replica): ids, names, size, version, and scheduling metadata. The “what does my fleet actually look like” lookup — most useful for grabbing a cluster or replica id to scope the rest of a dashboard to.mz_compute_cluster_status{materialize_cloud_organization_name=~".*", compute_cluster_id=~".*", compute_replica_id=~".*"}
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}, compute_cluster_id:${mzClusterList}, compute_replica_id:${mzReplicaList}}
by {compute_cluster_id,compute_cluster_name,compute_replica_id,compute_replica_name,size,mz_version}
materialize-compute#
Queries for the compute side of a Materialize deployment — the indexes, materialized views, and subscribes that run as dataflows on cluster replicas, plus their freshness, hydration, and resource footprint.
materialize.compute.materialized_views.count #
Materialized views Materialize is actively maintaining. Each one is a query whose result is kept continuously up to date, so this tracks roughly how much standing compute the environment carries.max(mz_mzd_views_count{materialize_cloud_organization_name=~".*"})
max:${mzSqlPrefix}mzd_views_count{${mzEnvironmentFilter}}
materialize.compute.indexes.count #
Indexes in the catalog. An index is an in-memory arrangement that makes reads against its relation effectively instant, in exchange for memory — so growth here is a leading indicator of cluster memory growth.max(sum by (instance) (mz_indexes_count{materialize_cloud_organization_name=~".*"}))
default_zero(sum:${mzSqlPrefix}indexes_count{${mzEnvironmentFilter}} by {instance})
materialize.compute.views.count #
Non-materialized views — query templates evaluated on demand. They cost nothing until something reads them, so this is a catalog-shape signal rather than a load one.max(mz_views_count{materialize_cloud_organization_name=~".*"})
max:${mzSqlPrefix}views_count{${mzEnvironmentFilter}}
materialize.compute.subscribes.active #
Live SUBSCRIBE sessions — long-running queries that stream updates to a client as data changes. A handful ofsystem subscribes are
Materialize’s own internal probes; a persistently climbing user count
usually means a client is opening subscribes and never closing them.sum by (session_type) (mz_active_subscribes{materialize_cloud_organization_name=~".*"})
sum:mz_active_subscribes{${mzEnvironmentFilter}} by {session_type}
materialize.compute.indexes.by_type #
Indexes split by the kind of relation they sit on. Workloads normally lean heavily on indexes over views (the standard “keep a query’s result hot” pattern); a large share of indexes on base tables is unusual and usually worth a second look.sum by (relation_type) (mz_indexes_count{materialize_cloud_organization_name=~".*"})
sum:${mzSqlPrefix}indexes_count{${mzEnvironmentFilter}} by {relation_type}
materialize.compute.hydration.currently_hydrating #
Objects still rebuilding their in-memory state — a live hydration-queue proxy. After a restart, replica creation, or some DDL, a dataflow has to rebuild from persisted storage before it can serve, and until it does it produces no results.count(
max by (instance_id, collection_id) (
mz_dataflow_wallclock_lag_seconds{materialize_cloud_organization_name=~".*", instance_id=~".*", instance_id!="", quantile="1"} > 1e15
)
)
default_zero(
count_not_null(
max:mz_dataflow_wallclock_lag_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}, quantile:1} by {instance_id,collection_id}
)
)
materialize.compute.hydration.queue_size #
Objects waiting in each replica’s hydration queue. environmentd schedules hydration in batches; a backlog means work is arriving faster than the replica can rebuild it.sum by (instance_id, replica_id) (
mz_compute_controller_hydration_queue_size{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*"}
) > 0
sum:mz_compute_controller_hydration_queue_size{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}}
by {instance_id,replica_id}
materialize.compute.hydration.slowest_collections #
The 15 objects that took longest to finish hydrating. Hydration time scales with the size of the state being rebuilt, so large materialized views and indexes naturally top the list.topk(15,
mz_compute_hydration_time_seconds{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*", hydrated="1"}
)
top(
max:${mzSqlPrefix}compute_hydration_time_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}, hydrated:1} by {instance_id,collection_id},
15, 'max', 'desc'
)
materialize.compute.freshness.lag_by_cluster #
How far behind real time each cluster’s least fresh object is — the worst-case freshness across every index, materialized view, and source on the cluster.max by (instance_id) (
mz_dataflow_wallclock_lag_seconds{materialize_cloud_organization_name=~".*", instance_id=~".*", instance_id!="", quantile="1"} < 1e9
)
max:mz_dataflow_wallclock_lag_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}, quantile:1} by {instance_id}
materialize.compute.freshness.lag_total_by_cluster #
The freshness of every object on each cluster, added together — one number for how far behind the cluster is in total.sum by (instance_id) (
max by (instance_id, collection_id) (
mz_dataflow_wallclock_lag_seconds{materialize_cloud_organization_name=~".*", instance_id=~".*", instance_id!="", quantile="1"} < 1e9
)
)
sum:mz_dataflow_wallclock_lag_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}, quantile:1} by {instance_id}
materialize.compute.freshness.top_collections #
The 15 objects whose results are furthest behind real time — the per-object breakdown behind Worst Freshness by Cluster, labeled by object name.topk(15,
max by (instance_id, collection_id) (
mz_dataflow_wallclock_lag_seconds{materialize_cloud_organization_name=~".*", instance_id=~".*", instance_id!="", replica_id=~".*", quantile="1"} < 1e9
)
)
top(
max:mz_dataflow_wallclock_lag_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}, quantile:1} by {instance_id,collection_id},
15, 'max', 'desc'
)
materialize.compute.dataflows.count #
Active dataflows on each replica. Every index, materialized view, and live SUBSCRIBE runs as one or more dataflows, so this count rises with DDL and subscribe activity.max by (cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id) (
mz_compute_replica_history_dataflow_count{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}
)
max:mz_compute_replica_history_dataflow_count{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
by {cluster_environmentd_materialize_cloud_cluster_id,cluster_environmentd_materialize_cloud_replica_id}
materialize.compute.dataflows.count_by_worker #
The dataflow count broken out per worker. Workers in a replica run in lockstep and should see exactly the same dataflows, so their series should overlap perfectly.max by (cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, worker_id) (
mz_compute_replica_history_dataflow_count{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}
)
max:mz_compute_replica_history_dataflow_count{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
by {cluster_environmentd_materialize_cloud_cluster_id,cluster_environmentd_materialize_cloud_replica_id,worker_id}
materialize.compute.dataflows.elapsed_rate #
CPU-cores busy inside dataflows, per cluster — the whole of dataflow work: arrangement maintenance, query evaluation, and hydration. Capped by cluster size (a 400cc cluster can’t exceed 400 cores).sum by (instance_id) (
max without (job) (rate(
mz_dataflow_elapsed_seconds_total{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*"}[5m]
))
)
sum:${mzSqlPrefix}dataflow_elapsed_seconds_total{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}}
by {instance_id}.as_rate()
materialize.compute.arrangements.maintenance_rate #
CPU-cores spent maintaining arrangements — the in-memory indexed snapshots behind every index and materialized view — summed across a replica’s workers, so an N-worker replica can reach N.sum by (cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id) (
max without (job) (rate(
mz_arrangement_maintenance_seconds_total{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]
))
)
sum:mz_arrangement_maintenance_seconds_total{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
by {cluster_environmentd_materialize_cloud_cluster_id,cluster_environmentd_materialize_cloud_replica_id}.as_rate()
materialize.compute.arrangements.maintenance_rate_by_worker #
The same maintenance CPU, split per worker — each worker tops out at 1.0.sum by (cluster_environmentd_materialize_cloud_cluster_id, cluster_environmentd_materialize_cloud_replica_id, worker_id) (
max without (job) (rate(
mz_arrangement_maintenance_seconds_total{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]
))
)
sum:mz_arrangement_maintenance_seconds_total{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
by {cluster_environmentd_materialize_cloud_cluster_id,cluster_environmentd_materialize_cloud_replica_id,worker_id}.as_rate()
materialize.compute.arrangements.records.system #
Row counts of arrangements for Materialize’s internal system objects (object id starts withs). These back the catalog and internal
probes, not user data, so they shouldn’t grow with your workload —
unexpected growth here can point at a Materialize bug.max by (collection_id) (
mz_arrangement_record_count{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*", collection_id=~"s.*"}
)
max:${mzSqlPrefix}arrangement_record_count{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}, collection_id:s*}
by {collection_id}
materialize.compute.arrangements.records.user #
Row counts of arrangements for your compute objects (object id starts withu) — the row count of every user index and materialized view, and
the primary driver of cluster memory. Growth on an object tracks the
size of its underlying data.max by (collection_id) (
mz_arrangement_record_count{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*", collection_id=~"u.*"}
)
max:${mzSqlPrefix}arrangement_record_count{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}, collection_id:u*}
by {collection_id}
materialize.compute.arrangements.records.transient #
Row counts of transient (object idt*) and uncategorized (none)
arrangements — short-lived intermediates from query optimization and
dataflow execution. Normally small and ephemeral.max by (collection_id) (
mz_arrangement_record_count{materialize_cloud_organization_name=~".*", instance_id=~".*", replica_id=~".*", collection_id=~"t.*|none"}
)
max:${mzSqlPrefix}arrangement_record_count{${mzEnvironmentFilter}, instance_id:${mzClusterList}, replica_id:${mzReplicaList}, (collection_id:t* OR collection_id:none)}
by {collection_id}
materialize-connections#
Sessions, query activity, and SQL control-plane (adapter) traffic for a Materialize deployment. Adapted from the Overview dashboard’s “Connections / Activity” tab.
materialize.connections.sessions.active #
Open SQL sessions, split intosystem (Materialize’s internal probing —
a few are always present) and user (client connections).sum by (session_type) (mz_active_sessions{materialize_cloud_organization_name=~".*"})
sum:mz_active_sessions{${mzEnvironmentFilter}} by {session_type}
materialize.connections.queries.rate #
Queries per second by session type —user tracks your client traffic,
system is the steady single-digit baseline of internal health checks.
Bursty is normal.sum by (session_type) (rate(mz_query_total{materialize_cloud_organization_name=~".*"}[5m]))
sum:mz_query_total{${mzEnvironmentFilter}} by {session_type}.as_rate()
materialize.connections.adapter.command_rate #
Commands per second through the adapter — the SQL protocol layer (parse, execute, prepare, fetch). Normally runs higher than the query rate, since one query is several commands.sum(rate(mz_adapter_commands{materialize_cloud_organization_name=~".*"}[5m]))
sum:mz_adapter_commands{${mzEnvironmentFilter}}.as_rate()
materialize.connections.queries.distribution #
The mix of query kinds over the selected window — a workload-shape signal, not a rate. Heavyset_variable/fetch traffic is normal
(that’s how Postgres clients manage session state); heavy
insert/update/delete on something you think of as read-mostly is
worth a look.sum by (statement_type) (increase(mz_query_total{materialize_cloud_organization_name=~".*"}[1h])) > 0
sum:mz_query_total{${mzEnvironmentFilter}} by {statement_type}.as_count()
materialize.connections.queries.rate_by_statement #
Query rate broken down by statement type and session type, fully time-resolved — the moving picture behind the distribution donut. A spike inselect/user is the thing to line up against peek latency
to confirm the system kept pace.sum by (statement_type, session_type) (rate(mz_query_total{materialize_cloud_organization_name=~".*"}[5m])) > 0
sum:mz_query_total{${mzEnvironmentFilter}} by {statement_type,session_type}.as_rate()
materialize.connections.peek_latency.p50 #
Median read-query latency — the typical time to look up the current state of an arrangement, which is the operation behind everySELECT
against an index. Your “what does a normal query feel like” number.histogram_quantile(0.50,
sum by (le, instance_id) (
rate(mz_compute_peek_duration_seconds_bucket{materialize_cloud_organization_name=~".*", instance_id=~".*"}[5m])
)
)
p50:mz_compute_peek_duration_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}}
by {instance_id}
materialize.connections.peek_latency.p90 #
90th-percentile read-query latency — how slow the slowest 10% of queries feel. Catches the contention bursts and cold paths that the median hides.histogram_quantile(0.90,
sum by (le, instance_id) (
rate(mz_compute_peek_duration_seconds_bucket{materialize_cloud_organization_name=~".*", instance_id=~".*"}[5m])
)
)
p90:mz_compute_peek_duration_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}}
by {instance_id}
materialize.connections.peek_latency.p99 #
Tail read-query latency — the slowest 1% of queries, the ones users complain about.histogram_quantile(0.99,
sum by (le, instance_id) (
rate(mz_compute_peek_duration_seconds_bucket{materialize_cloud_organization_name=~".*", instance_id=~".*"}[5m])
)
)
p99:mz_compute_peek_duration_seconds{${mzEnvironmentFilter}, instance_id:${mzClusterList}}
by {instance_id}
materialize.connections.adapter.commands_by_application #
SQL control-plane command totals per clientapplication_name over the
window, so you can see which clients drive the adapter and which are
failing. Most clients set application_name in their connection string;
those that don’t bucket as unrecognized/unspecified (normal).sum by (application_name, status) (increase(mz_adapter_commands{materialize_cloud_organization_name=~".*"}[1h]))
sum:mz_adapter_commands{${mzEnvironmentFilter}} by {application_name,status}.as_count()
materialize-consensus#
The metadata (consensus) database, from the vantage point of the Materialize processes that use it.
Persist is Materialize’s storage layer. Every durable collection — a table, a source, a materialized view, the catalog itself — is a persist shard: data files in object storage, plus a small record of the shard’s current state in the metadata database. Every change to a shard commits a new version of that record with a compare-and-set, so the database sees a steady stream of small conditional writes (about 80 a second on an idle environment) and very little else. The timestamp oracle is the second client: environmentd keeps the read and write timestamps of each timeline in the same database.
Everything here is the client’s measurement, which is the SLI: it is what
a Materialize process experienced, including the network, DNS, TLS and its
own connection pool. It needs no cloud credentials and is identical on every
database flavor. What the database was doing at the time — CPU, storage,
connection ceilings — is the provider’s view, on infra-cloud. The external
dependency design doc records why the two are separate.
Label families#
op on mz_persist_external_* names the operation: consensus_cas (the
commit), consensus_head, consensus_scan, consensus_truncate and
consensus_list_keys. The blob_* values of the same label are object
storage and belong to materialize-persist.yaml. The latency histogram
mz_persist_external_op_latency covers consensus_cas only on this side;
the other operations have a mean from _seconds / _started_count.
op on mz_persist_retry_* is a retry loop, named after the work rather
than the call. The five listed in materialize.consensus.retries.by_operation
are the ones that wrap a metadata-database call; next_listen_batch and
snapshot also retry constantly, but they are polling for new data rather
than recovering from a failure and must never be drawn as retries. The same
names, as they appear in the logs, back materialize.consensus.logs.retries.
The process breakdown is app (environmentd or clusterd) plus the
long-form cluster id, which only clusterd carries. mzClusterName turns
the id into the cluster’s name and leaves environmentd’s blank.
Two things that are easy to get wrong#
mz_persist_postgres_connpool_availablegoes negative by the number of callers waiting for a connection (deadpool’sStatus.available), and is sampled at each acquire.clamp_min(-available, 0)is the queue.- Pool size is connections the process currently holds, which is what the
database counts against
max_connections. The pool’s configured maximum is not published.
Every mean here is a ratio of two rates wrapped in (…) >= 0. An operation
nobody called in the window divides zero by zero, and the resulting NaN
would otherwise be drawn — and would put the panel’s empty-state text into
its legend, which Grafana applies per field.
materialize.consensus.health.failed_ops #
Metadata database calls that failed, per second, across every Materialize process in this environment.sum(rate(mz_persist_external_failed_count{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m]))
materialize.consensus.health.connection_errors #
Failed attempts to open a connection to the metadata database, per second.sum(rate(mz_persist_postgres_connpool_connection_errors{materialize_cloud_organization_name=~".*"}[5m]))
+ (sum(rate(mz_ts_oracle_postgres_connpool_connection_errors{materialize_cloud_organization_name=~".*"}[5m])) or vector(0))
materialize.consensus.health.commit_latency_p99 #
How long the slowest 1% of commits to the metadata database take.histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m])
)
)
materialize.consensus.health.waiting #
Calls currently queued for a free database connection, across every process.sum(clamp_min(-mz_persist_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0))
+ (sum(clamp_min(-mz_ts_oracle_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0)) or vector(0))
materialize.consensus.health.connections_held #
Connections to the metadata database this environment currently holds open.sum(mz_persist_postgres_connpool_size{materialize_cloud_organization_name=~".*"})
+ (sum(mz_ts_oracle_postgres_connpool_size{materialize_cloud_organization_name=~".*"}) or vector(0))
materialize.consensus.health.commits #
Commits to the metadata database per second — the load this environment puts on it.sum(rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m]))
materialize.consensus.latency.commit #
Commit latency to the metadata database: the median and the slowest 1%.histogram_quantile(0.50,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m])
)
)
histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m])
)
)
materialize.consensus.latency.round_trip #
The network round trip from each Materialize process to the metadata database, as that process last measured it.max by (app, cluster_environmentd_materialize_cloud_cluster_id) (
mz_persist_external_rtt_latency{materialize_cloud_organization_name=~".*", external="consensus"}
)
materialize.consensus.failures.by_operation #
Failed metadata database calls per second, by operation.sum by (op) (
rate(mz_persist_external_failed_count{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m])
)
sum by (op) (
rate(mz_ts_oracle_failed_count{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.consensus.failures.connection_errors #
Failed connection attempts per second, by the process that made them.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_postgres_connpool_connection_errors{materialize_cloud_organization_name=~".*"}[5m])
)
sum(rate(mz_ts_oracle_postgres_connpool_connection_errors{materialize_cloud_organization_name=~".*"}[5m]))
materialize.consensus.logs.retries #
What Materialize logged while retrying calls to the metadata database, with the database’s own error text.materialize.consensus.ops.by_type #
Metadata database calls per second, by operation.sum by (op) (
rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m])
)
materialize.consensus.ops.commits_by_process #
Commits per second, by the process making them.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m])
)
materialize.consensus.ops.bytes #
Bytes per second exchanged with the metadata database, by operation.sum by (op) (
rate(mz_persist_external_bytes_count{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m])
)
materialize.consensus.ops.collections #
Durable objects in this environment — each one a stream of commits to the metadata database.count(group by (shard) (mz_persist_shard_usage_current_state_batches_bytes{materialize_cloud_organization_name=~".*"}))
materialize.consensus.latency.mean_by_operation #
Average time per metadata database call, by operation.(
sum by (op) (
rate(mz_persist_external_seconds{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m])
)
/
sum by (op) (
rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op=~"consensus_.*"}[5m])
)
) >= 0
materialize.consensus.latency.commit_by_process #
Slowest 1% of commits, by the process making them.histogram_quantile(0.99,
sum by (le, app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m])
)
)
materialize.consensus.contention.conflicts #
Commits that lost a race with another writer and were retried, per second, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_cmd_cas_mismatch_count{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.consensus.retries.by_operation #
Metadata database calls retried after a failure, per second, by the work that was retrying.sum by (op) (
rate(mz_persist_retry_retries_count{materialize_cloud_organization_name=~".*", op=~"fetch_state::scan|apply_unbatched_cmd::cas|gc::truncate|maybe_init::cas|consensus::open"}[5m])
)
sum by (op) (
rate(mz_ts_oracle_retry_retries_count{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.consensus.oracle.ops #
Timestamp oracle calls per second, by operation.sum by (op) (
rate(mz_ts_oracle_started_count{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.consensus.oracle.latency #
Average time per timestamp oracle call, by operation.(
sum by (op) (
rate(mz_ts_oracle_seconds{materialize_cloud_organization_name=~".*"}[5m])
)
/
sum by (op) (
rate(mz_ts_oracle_started_count{materialize_cloud_organization_name=~".*"}[5m])
)
) >= 0
materialize.consensus.pool.size #
Connections to the metadata database held open, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
mz_persist_postgres_connpool_size{materialize_cloud_organization_name=~".*"}
)
sum(mz_ts_oracle_postgres_connpool_size{materialize_cloud_organization_name=~".*"})
materialize.consensus.pool.in_use #
Connections actively lent out to a call, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
mz_persist_postgres_connpool_size{materialize_cloud_organization_name=~".*"}
- clamp_min(mz_persist_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0)
)
sum(
mz_ts_oracle_postgres_connpool_size{materialize_cloud_organization_name=~".*"}
- clamp_min(mz_ts_oracle_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0)
)
materialize.consensus.pool.waiting #
Calls queued for a connection, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
clamp_min(-mz_persist_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0)
)
sum(clamp_min(-mz_ts_oracle_postgres_connpool_available{materialize_cloud_organization_name=~".*"}, 0))
materialize.consensus.pool.acquire_wait #
Average time a call waits to be handed a connection, by process.(
sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_postgres_connpool_acquire_seconds{materialize_cloud_organization_name=~".*"}[5m])
)
/
sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_postgres_connpool_acquires{materialize_cloud_organization_name=~".*"}[5m])
)
) >= 0
(
sum(rate(mz_ts_oracle_postgres_connpool_acquire_seconds{materialize_cloud_organization_name=~".*"}[5m]))
/
sum(rate(mz_ts_oracle_postgres_connpool_acquires{materialize_cloud_organization_name=~".*"}[5m]))
) >= 0
materialize.consensus.pool.churn #
New connections opened per second, and how many of those replaced a connection that had reached its age limit.sum(rate(mz_persist_postgres_connpool_connections_created{materialize_cloud_organization_name=~".*"}[5m]))
+ (sum(rate(mz_ts_oracle_postgres_connpool_connections_created{materialize_cloud_organization_name=~".*"}[5m])) or vector(0))
sum(rate(mz_persist_postgres_connpool_ttl_reconnections{materialize_cloud_organization_name=~".*"}[5m]))
+ (sum(rate(mz_ts_oracle_postgres_connpool_ttl_reconnections{materialize_cloud_organization_name=~".*"}[5m])) or vector(0))
materialize.consensus.state.versions #
State versions written to the metadata database per second, beside versions deleted.sum(rate(mz_persist_external_succeeded_count{materialize_cloud_organization_name=~".*", op="consensus_cas"}[5m]))
sum(rate(mz_persist_external_consensus_truncated_count{materialize_cloud_organization_name=~".*"}[5m]))
materialize.consensus.state.live #
State versions currently stored in the metadata database, summed across every object — roughly the number of rows persist keeps there.sum(max by (shard) (mz_persist_shard_gc_live_diffs{materialize_cloud_organization_name=~".*"}))
materialize.consensus.state.held #
The ten objects holding the most old state versions, which cleanup cannot delete until whatever is reading them moves on.topk(10, max by (shard) (mz_persist_shard_seqnos_held{materialize_cloud_organization_name=~".*"}))
materialize.consensus.gc.runs #
Cleanup passes per second: started, finished, and skipped because another process had already done the work.sum(rate(mz_persist_gc_started{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_gc_finished{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_gc_noop{materialize_cloud_organization_name=~".*"}[5m]))
materialize.consensus.gc.lease_timeouts #
Readers whose lease on an object expired, per second.sum(rate(mz_persist_lease_timeout_read{materialize_cloud_organization_name=~".*"}[5m]))
materialize-events#
Kubernetes events from the namespaces a Materialize deployment occupies: the operator’s own namespace and the environments’ namespace.
Events are logs, not metrics. They arrive through
loki.source.kubernetes_events in the monitoring gateway, which reads them
from the Kubernetes API and forwards them to Loki, so every query here is
LogQL against the logs datasource rather than PromQL. The gateway’s processor
lifts reason, name, kind, count, node and reportingcontroller out
of each event into structured metadata, which is what these queries filter and
group on; the event type becomes the level stream label, Normal as
INFO and Warning as WARN.
Two scopes live here. The deployment and operator queries below are
rollout-scoped: they answer “is this upgrade going through”, and the
env-upgrade dashboard is their consumer. The cluster queries at the end are
the general-purpose view, scoped by the same Loki-discovered namespace picker a
logs dashboard uses, and are deliberately separate definitions rather than the
same ones widened — the rollout queries carry filters (generation, reporting
controller) that a general event browser has no business inheriting.
Kubernetes keeps events for about an hour. Loki keeps them for as long as the deployment’s retention says, which is what makes a rollout that finished yesterday still explainable.
An event’s namespace is the involved object’s, not the reporter’s. The
operator runs in its own namespace and reconciles resources in the
environments’ namespace, and every event it publishes is filed against the
resource — so the operator’s own events are found in the environment
namespace, where nothing else about them suggests they would be. That is why
the queries below scope to both namespaces and pick the operator out by
reportingcontroller, which is the reporter’s identity and the only field
that actually says an event came from orchestratord.
Only the deployment-wide feeds filter by generation. The operator’s own
events are filed against the Materialize, Balancer and Console resources,
which carry no generation at all, so %%{mzGenerationEventFilter} could only
ever be a no-op on them. The filter itself keeps generation-less objects on
purpose — on a representative deployment only 6 of 70 event names carry a
generation, and dropping the other 64 would take the whole rollout narrative
with them.
materialize.events.deployment.stream #
Every Kubernetes event from the operator and environment namespaces, newest first — the unfiltered record of what the cluster did.materialize.events.deployment.warnings #
Kubernetes events the reporting component flagged as warnings — a pod that will not schedule, an image that will not pull, a container failing its probes.materialize.events.deployment.rate.by_reason #
How often each kind of event is being reported, by reason. The shape of a rollout:Pulled, Created and Started rise together as pods are
replaced, and fall back to nothing when it finishes.materialize.events.deployment.warning.rate #
Warning events per interval across both namespaces, as one series — the at-a-glance answer to whether anything is complaining right now.materialize.events.operator.lifecycle #
Every phase the Materialize resource moved through, as the operator reported it:Applying, ReadyToPromote, WaitingForApproval,
Promoting, Applied, and the two that end a rollout badly,
RolloutTimeout and FailedDeploy.materialize.events.operator.lifecycle.rate #
Lifecycle transitions over time, by phase — where a rollout got to, and when.materialize.events.operator.reconciliation.failures #
Why the operator could not reconcile a resource. The event carries the error’s whole cause chain, which is usually the actionable half — an admission webhook that is down, a secret that does not exist yet, a license key that will not parse.materialize.events.operator.reconciliation.failures.rate #
Reconciliation failures over time, by the kind of resource that failed.materialize.events.cluster.stream #
Every Kubernetes event in the selected namespaces, newest first.materialize.events.cluster.warnings #
Kubernetes events the reporting component flagged as warnings, across the selected namespaces.materialize.events.cluster.rate.by_reason #
How often each kind of event is being reported, by reason — the shape of what the cluster is doing.materialize.events.cluster.rate.by_namespace #
Event rate by namespace — where in the cluster things are happening.materialize-generations#
What each deployment generation of an environment is doing, during and after a blue/green rollout.
A rollout stands a new generation of environmentd and its replicas up beside
the old one, lets it rehydrate from persisted storage, and promotes it only once
it has caught up. Both generations are live and scraped at the same time, so
every ordinary panel sums them together — which is exactly the wrong thing while
the question is whether one of them is ready yet.
The generation is not a label. orchestratord records it as a Kubernetes
annotation, which neither kube-state-metrics nor cAdvisor surfaces. Where it
does reach a query is the object name, in two shapes:
…-environmentd-<generation>-<ordinal> and, for a replica,
…-gen-<generation>-<ordinal>. %%{mzGenerationFilter} selects on those, and
%%{mzGenerationPattern} is the same shape as a capture, for the
label_replace that lifts the number into a generation label panels can group
by. Both live in the render context, so they cannot drift apart.
materialize.generations.active #
How many deployment generations are currently running — one between rollouts, two while one is in flight.count(
count by (generation) (
label_replace(
mz_compute_commands_total{
materialize_cloud_organization_name=~".*", ${mzGenerationFilter}
},
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
)
materialize.generations.version #
The Materialize version each generation is running — what the rollout is actually changing.group by (generation, mz_version) (
label_replace(
max_over_time(
mz_compute_cluster_status{
materialize_cloud_organization_name=~".*", ${mzGenerationFilter}
}
[1h]
),
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
materialize.generations.pods #
Pods belonging to each generation — its environmentd and the cluster replicas standing behind it.count by (generation) (
group by (pod, generation) (
label_replace(
container_memory_working_set_bytes{
container!="POD", container!="", ${mzGenerationFilter}
},
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
)
materialize.generations.hydrating #
Objects still rebuilding their in-memory state, split by generation — the panel that answers whether a new generation is ready to promote.sum by (generation) (
max by (generation, instance_id, collection_id) (
label_replace(
mz_dataflow_wallclock_lag_seconds{
materialize_cloud_organization_name=~".*",
${mzGenerationFilter},
instance_id=~".*",
instance_id!="",
quantile="1"
},
"generation", "$1", "pod", "${mzGenerationPattern}"
) > bool 1e15
)
)
materialize.generations.collections #
Objects each generation is tracking — the denominator for hydration, and the shape of a new generation building out its dataflows.count by (generation) (
max by (generation, instance_id, collection_id) (
label_replace(
mz_dataflow_wallclock_lag_seconds{
materialize_cloud_organization_name=~".*",
${mzGenerationFilter},
instance_id=~".*",
instance_id!="",
quantile="1"
},
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
)
materialize.generations.lag.max #
The worst freshness in each generation — how far behind real time its least fresh object is.max by (generation) (
label_replace(
mz_dataflow_wallclock_lag_seconds{
materialize_cloud_organization_name=~".*",
${mzGenerationFilter},
instance_id=~".*",
instance_id!="",
quantile="1"
},
"generation", "$1", "pod", "${mzGenerationPattern}"
) < 1e9
)
materialize.generations.lag.total #
Every hydrated object’s freshness in each generation, added together — how far behind the generation is in total.sum by (generation) (
max by (generation, instance_id, collection_id) (
label_replace(
mz_dataflow_wallclock_lag_seconds{
materialize_cloud_organization_name=~".*",
${mzGenerationFilter},
instance_id=~".*",
instance_id!="",
quantile="1"
},
"generation", "$1", "pod", "${mzGenerationPattern}"
) < 1e9
)
)
materialize.generations.lag.total_by_cluster #
Total freshness split by generation and cluster — which cluster in which generation is furthest behind.sum by (generation, instance_id) (
max by (generation, instance_id, collection_id) (
label_replace(
mz_dataflow_wallclock_lag_seconds{
materialize_cloud_organization_name=~".*",
${mzGenerationFilter},
instance_id=~".*",
instance_id!="",
quantile="1"
},
"generation", "$1", "pod", "${mzGenerationPattern}"
) < 1e9
)
)
materialize.generations.cpu #
CPU used by each generation’s pods — what a rollout costs while both sides are up.sum by (generation) (
label_replace(
rate(
container_cpu_usage_seconds_total{
container!="POD", container!="", ${mzGenerationFilter}
}
[5m]
),
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
materialize.generations.memory #
Memory used by each generation’s pods.sum by (generation) (
label_replace(
container_memory_working_set_bytes{
container!="POD", container!="", ${mzGenerationFilter}
},
"generation", "$1", "pod", "${mzGenerationPattern}"
)
)
materialize-health#
Common queries for checking the health of a Materialize deployment.
materialize.scraper.mzmon.environmentd #
environmentd metrics are reaching the gateway. If this goes to 0 or disappears, every environmentd-backed panel and alert for the environment is blind — the environment may be perfectly healthy and you simply can’t see it, so rule this out first.up{
job="monitoring/mzmon-materialize-environmentd",
namespace=~"materialize-environment"
} == 1
avg:up{job:monitoring/mzmon-materialize-environmentd, ${mzEnvironmentNamespaceFilter}}
materialize.scraper.mzmon.clusterd #
clusterd (compute replica) metrics are reaching the gateway. When this drops, per-cluster compute signals — arrangements, peeks, dataflows — go dark even though the replicas may still be serving.up{
job="monitoring/mzmon-materialize-clusterd",
namespace=~"materialize-environment"
} == 1
avg:up{job:monitoring/mzmon-materialize-clusterd, ${mzEnvironmentNamespaceFilter}}
materialize.scraper.mzmon.orchestratord #
The Materialize operator (orchestratord) is being scraped. Losing it blinds you to cluster and replica lifecycle — creation, resize, and rollout progress — not to the running workloads themselves.up{
job="materialize/mzmon-materialize-operator",
namespace=~"materialize"
} == 1
avg:up{job:materialize/mzmon-materialize-operator, ${mzOperatorNamespaceFilter}}
materialize.health.clusters.status.percentage #
The share of the environment’s clusters currently reporting ready — the at-a-glance health headline for the whole environment.count(
mz_compute_cluster_status{
materialize_cloud_organization_name=~".*"
} == 1
) / count(
mz_compute_cluster_status{
materialize_cloud_organization_name=~".*"
}
) * 100
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}} * 100
materialize.health.environment.availability.percentage #
An SLO-style snapshot: how much of the selected window the environment’s clusters were ready. Sustained dips are the signal that something restarted or went down while you weren’t watching.avg by (materialize_cloud_organization_namespace) (
avg_over_time(
mz_compute_cluster_status{
materialize_cloud_organization_name=~".*"
}[1h]
) * 100
)
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}}
by {materialize_cloud_organization_namespace}.rollup(avg, ${range}) * 100
materialize.info.version #
The version of Materialize running in the environment. A single version is the steady state; multiple values appear briefly during a rolling upgrade.group by (mz_version) (
mz_compute_cluster_status{materialize_cloud_organization_name=~".*"}
)
avg:${mzSqlPrefix}compute_cluster_status{${mzEnvironmentFilter}} by {mz_version}
materialize.info.max_lag #
The worst freshness seen anywhere in the environment over the selected window — how far the most-behind object’s output trailed real time. A top-level freshness pointer.max(
max_over_time(
(
mz_dataflow_wallclock_lag_seconds{materialize_cloud_organization_name=~".*", instance_id!="", quantile="1"} < 1e9
)[${rangeWindow}:1m]
)
)
max:mz_dataflow_wallclock_lag_seconds{${mzEnvironmentFilter}, quantile:1}.rollup(max, 3600)
materialize-kubernetes#
Kubernetes-side view of a Materialize deployment: capacity, workload readiness, per-pod resource usage, and networking. Adapted from the Overview dashboard’s “Kubernetes Workloads” tab and the k8s-sourced Summary panels.
These read kube-state-metrics and cAdvisor (via the kubelet), NOT Materialize
metrics — the same meta-monitoring surface other targets (e.g. Loki health)
will draw on. %%{cAdvisorFilter} is the container-scoping filter fragment
(namespace + drop empty/pause series); %%{mzNamespaceList} is the raw
namespace selector used by the kube_* and container_network_* metrics.
The percent-of-limit panels have an absolute-units sibling for deployments whose metrics source (e.g. GKE’s managed cAdvisor/KSM) doesn’t expose resource limits; the dashboard picks whichever fits the environment.
materialize.kubernetes.cpu.capacity #
Total CPU cores configured across the environment’s containers (sum of cAdvisor CPU limits), excluding the monitoring exporter — i.e. the CPU available to the actual workload.sum by (container) (
container_spec_cpu_quota{container!="POD", container!="", container!="new-promsql-exporter"}
/ container_spec_cpu_period{container!="POD", container!="", container!="new-promsql-exporter"}
)
sum:container_spec_cpu_quota{${cAdvisorFilter}, !container:new-promsql-exporter} by {container} / 100000
materialize.kubernetes.memory.capacity #
Total memory configured across the environment’s containers (sum of cAdvisor memory limits), excluding the monitoring exporter. Memory is Materialize’s dominant constraint — in-memory arrangements live in here.sum by (container) (
container_spec_memory_limit_bytes{container!="POD", container!="", container!="new-promsql-exporter"}
)
sum:container_spec_memory_limit_bytes{${cAdvisorFilter}, !container:new-promsql-exporter} by {container}
materialize.kubernetes.cpu.capacity.all_containers #
Total CPU cores configured across every container in the environment, including the monitoring exporter — the Kubernetes view of what is provisioned rather than what is available to the workload.sum by (container) (
container_spec_cpu_quota{container!="POD", container!=""}
/ container_spec_cpu_period{container!="POD", container!=""}
)
sum:container_spec_cpu_quota{${cAdvisorFilter}} by {container} / 100000
materialize.kubernetes.memory.capacity.all_containers #
Total memory configured across every container in the environment, including the monitoring exporter — the Kubernetes view of what is provisioned rather than what is available to the workload.sum by (container) (
container_spec_memory_limit_bytes{container!="POD", container!=""}
)
sum:container_spec_memory_limit_bytes{${cAdvisorFilter}} by {container}
materialize.kubernetes.cpu.usage.percent #
Current CPU usage per container type as a fraction of its limit, averaged over the last 5 minutes — shows the worst-loaded container types.sum by (namespace, container) (
rate(container_cpu_usage_seconds_total{container!="POD", container!=""}[5m])
) / sum by (namespace, container) (
kube_pod_container_resource_limits{resource="cpu", namespace=~"materialize-environment"}
)
sum:container_cpu_usage_seconds_total{${cAdvisorFilter}} by {namespace,container}.as_rate()
/ sum:kube_pod_container_resource_limits{resource:cpu, namespace:${mzNamespaceList}} by {namespace,container}
materialize.kubernetes.cpu.usage.absolute #
Current CPU usage per container type in cores (rate over 5 minutes), for deployments whose metrics source doesn’t expose CPU limits — read it against the replica sizes you configured.sum by (namespace, container) (
rate(container_cpu_usage_seconds_total{container!="POD", container!=""}[5m])
)
sum:container_cpu_usage_seconds_total{${cAdvisorFilter}} by {namespace,container}.as_rate()
materialize.kubernetes.memory.usage.percent #
Current memory usage per container type as a fraction of its limit — shows the worst-loaded container types.sum by (namespace, container) (
avg by (namespace, pod, container) (
container_memory_working_set_bytes{container!="POD", container!="", container!="new-promsql-exporter"}
)
) / sum by (namespace, container) (
avg by (namespace, pod, container) (
container_spec_memory_limit_bytes{container!="POD", container!="", container!="new-promsql-exporter"}
)
)
sum:container_memory_working_set_bytes{${cAdvisorFilter}, !container:new-promsql-exporter} by {namespace,container}
/ sum:container_spec_memory_limit_bytes{${cAdvisorFilter}, !container:new-promsql-exporter} by {namespace,container}
materialize.kubernetes.memory.usage.absolute #
Current memory (working set) per container type in bytes, for deployments whose metrics source doesn’t expose memory limits.sum by (namespace, container) (
container_memory_working_set_bytes{container!="POD", container!="", container!="new-promsql-exporter"}
)
sum:container_memory_working_set_bytes{${cAdvisorFilter}, !container:new-promsql-exporter} by {namespace,container}
materialize.kubernetes.last_restart #
Seconds since the most recent container restart in the environment.time()
- topk(1,
container_start_time_seconds{container!="POD", container!="", container!="new-promsql-exporter"}
)
top(
max:container_start_time_seconds{${cAdvisorFilter}, !container:new-promsql-exporter} by {namespace,pod,container},
1, 'max', 'desc'
)
materialize.kubernetes.pods.readiness #
Pods in the Materialize namespace grouped by phase (Running, Pending, Failed, …).max by (phase, namespace) (
sum by (phase, namespace, instance) (
kube_pod_status_phase{namespace=~"materialize-environment"}
)
)
sum:kube_pod_status_phase{namespace:${mzNamespaceList}} by {phase,namespace}
materialize.kubernetes.statefulsets.ready #
StatefulSet replicas reporting Ready. environmentd and the cluster pods are StatefulSets.max by (namespace) (
sum by (namespace, instance) (
kube_statefulset_status_replicas_ready{namespace=~"materialize-environment"}
)
)
sum:kube_statefulset_status_replicas_ready{namespace:${mzNamespaceList}} by {namespace}
materialize.kubernetes.deployments.readiness #
Deployment replica health — Ready vs Unavailable. Deployments back stateless services (e.g. the promsql exporter).max by (namespace) (
sum by (namespace, instance) (
kube_deployment_status_replicas_ready{namespace=~"materialize-environment"}
)
)
max by (namespace) (
sum by (namespace, instance) (
kube_deployment_status_replicas_unavailable{namespace=~"materialize-environment"}
)
)
sum:kube_deployment_status_replicas_ready{namespace:${mzNamespaceList}} by {namespace}
sum:kube_deployment_status_replicas_unavailable{namespace:${mzNamespaceList}} by {namespace}
materialize.kubernetes.pods.cpu_usage #
CPU utilization per pod as a fraction of the pod’s limit. Split so the cluster/replica selectors filter the cluster pods while envd/balancer/ exporter stay visible.sum by (namespace, pod, container) (
rate(container_cpu_usage_seconds_total{container!="POD", container!="", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
) / sum by (namespace, pod, container) (
kube_pod_container_resource_limits{resource="cpu", namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}
)
sum by (namespace, pod, container) (
rate(container_cpu_usage_seconds_total{container!="POD", container!="", pod!~".*-cluster-.*-replica-.*"}[5m])
) / sum by (namespace, pod, container) (
kube_pod_container_resource_limits{resource="cpu", namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}
)
sum:container_cpu_usage_seconds_total{${cAdvisorFilter}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod,container}.as_rate()
/ sum:kube_pod_container_resource_limits{resource:cpu, namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod,container}
sum:container_cpu_usage_seconds_total{${cAdvisorFilter}, !pod:*-cluster-*-replica-*} by {namespace,pod,container}.as_rate()
/ sum:kube_pod_container_resource_limits{resource:cpu, namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod,container}
materialize.kubernetes.pods.memory_usage #
Memory usage per pod as a fraction of the pod’s limit (working-set basis), same cluster/non-cluster split as pod CPU.avg by (namespace, pod, container) (
container_memory_working_set_bytes{container!="POD", container!="", container!="new-promsql-exporter", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}
) / avg by (namespace, pod, container) (
container_spec_memory_limit_bytes{container!="POD", container!="", container!="new-promsql-exporter", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}
)
avg by (namespace, pod, container) (
container_memory_working_set_bytes{container!="POD", container!="", container!="new-promsql-exporter", pod!~".*-cluster-.*-replica-.*"}
) / avg by (namespace, pod, container) (
container_spec_memory_limit_bytes{container!="POD", container!="", container!="new-promsql-exporter", pod!~".*-cluster-.*-replica-.*"}
)
avg:container_memory_working_set_bytes{${cAdvisorFilter}, !container:new-promsql-exporter, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod,container}
/ avg:container_spec_memory_limit_bytes{${cAdvisorFilter}, !container:new-promsql-exporter, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod,container}
avg:container_memory_working_set_bytes{${cAdvisorFilter}, !container:new-promsql-exporter, !pod:*-cluster-*-replica-*} by {namespace,pod,container}
/ avg:container_spec_memory_limit_bytes{${cAdvisorFilter}, !container:new-promsql-exporter, !pod:*-cluster-*-replica-*} by {namespace,pod,container}
materialize.kubernetes.pods.network_rx #
Network bytes/sec received per pod. For cluster pods, Rx tracks ingest from upstream and inter-pod replication; for envd/balancer it’s client SQL traffic. Surges alongside hydration are normal catchup.sum by (namespace, pod) (
rate(container_network_receive_bytes_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_receive_bytes_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum:container_network_receive_bytes_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_receive_bytes_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
materialize.kubernetes.pods.network_tx #
Network bytes/sec transmitted per pod. For cluster pods Tx covers sink output, inter-pod replication, and query results returning to envd; for envd it’s client query responses.sum by (namespace, pod) (
rate(container_network_transmit_bytes_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_transmit_bytes_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum:container_network_transmit_bytes_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_transmit_bytes_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
materialize.kubernetes.pods.network_errors #
Network rx + tx errors per pod per second (counted at the NIC/kernel level).sum by (namespace, pod) (
rate(container_network_receive_errors_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_receive_errors_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_transmit_errors_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_transmit_errors_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum:container_network_receive_errors_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_receive_errors_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
sum:container_network_transmit_errors_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_transmit_errors_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
materialize.kubernetes.pods.network_drops #
Network packets dropped (rx + tx) per pod per second — when kernel buffers fill faster than the app reads (rx) or egress rate-limiting kicks in (tx).sum by (namespace, pod) (
rate(container_network_receive_packets_dropped_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_receive_packets_dropped_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_transmit_packets_dropped_total{namespace=~"materialize-environment", pod=~".*-cluster-${mzClusterListRegex}-replica-${mzReplicaListRegex}-.*"}[5m])
)
sum by (namespace, pod) (
rate(container_network_transmit_packets_dropped_total{namespace=~"materialize-environment", pod!~".*-cluster-.*-replica-.*"}[5m])
)
sum:container_network_receive_packets_dropped_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_receive_packets_dropped_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
sum:container_network_transmit_packets_dropped_total{namespace:${mzNamespaceList}, pod:*-cluster-${mzClusterList}-replica-${mzReplicaList}-*} by {namespace,pod}.as_rate()
sum:container_network_transmit_packets_dropped_total{namespace:${mzNamespaceList}, !pod:*-cluster-*-replica-*} by {namespace,pod}.as_rate()
materialize-logs#
Logs, as collected by the monitoring stack and stored in Loki.
Every query here is LogQL, and the scope is Loki-discovered end to end: namespace, app and level come from Loki’s own label values rather than from the metrics pipeline. That is deliberate. Reading logs is frequently how an operator works out why the metrics pipeline is broken, and a logs dashboard that derived its scope from Prometheus would go blind at exactly the moment it is most needed.
The label contract the agent and gateway produce is documented under
Logs and Events. What matters here:
namespace, app, level and container are stream labels and belong in the
selector; pod, node, organization_name and the rest are structured
metadata and are filtered after a |. Narrowing the selector before the line
filters is the single biggest speedup.
Every selector here carries %%{mzLogJobFilter}, and it is not only a filter.
LogQL rejects a stream selector whose every matcher can match the empty string —
“queries require at least one regexp or equality matcher that does not have an
empty-compatible value” — and a dashboard built from =~ pickers is exactly
that shape. The job picker’s “All” is .+ rather than the discovered values, so
it always contributes a non-empty matcher and the selector parses whatever the
other pickers are set to. Without it, “All” everywhere is a query error rather
than a wide result.
The event queries in materialize-events.yaml need no such anchor: they pin
job="loki.source.kubernetes_events", which is already a non-empty equality
matcher, and a second job matcher would AND with it and zero the panel.
level is normalized by the pipeline where it can be and falls back to
UNKNOWN, so the levels present are a property of the workloads running rather
than a fixed vocabulary — which is why the dashboard discovers them instead of
hard-coding a list.
materialize.logs.stream #
The log feed for the selected namespaces, apps and levels, newest first.materialize.logs.rate.by_app #
Log lines per second by application — which component is doing the talking.materialize.logs.rate.by_level #
Log lines per second by severity — the shape of how much of the volume is something going wrong.materialize.logs.rate.total #
Total log lines per second reaching Loki for the current selection, averaged over each interval.materialize.logs.warnings.rate #
Warning-and-worse log lines per minute, as one series — the at-a-glance answer to whether anything is complaining.materialize.logs.warnings.stream #
The warning-and-worse feed, newest first — what the components are actually complaining about.materialize-operator#
How the Materialize operator’s reconciliation loop is behaving.
orchestratord watches the Materialize, Balancer and Console resources
and drives each toward the state its spec asks for. One trip through that
work is a pass, and a pass moves through named steps. Both are counted by
outcome, and both are timed, which is what lets a stuck rollout say not just
that it is stuck but which phase it is stuck in.
These are the operator’s own metrics, scraped from its pods in the operator
namespace, so they are scoped by %%{mzOperatorNamespaceFilter} and by
nothing else. They carry no organization label, so the environment picker
does not narrow them: one operator reconciles every environment in the
cluster, and its loop is a single shared thing rather than a per-environment
one.
Only the replica holding the leadership lease reconciles. The others export the same metric families sitting at zero, which is why every query here sums across replicas rather than picking one out.
materialize.operator.reconciling.replicas #
How many operator replicas hold the leadership lease and are therefore reconciling. This should be exactly one.sum(
orchestratord_is_leader{namespace=~"materialize"}
)
materialize.operator.environments.needing_update #
How many environments in this cluster are still running an outdated pod template — the count an upgrade is working to bring to zero.sum(
environmentd_needs_update{namespace=~"materialize"}
)
materialize.operator.reconciliation.rate #
Reconciliation passes per second across every controller — whether the loop is turning at all.sum(
rate(
orchestratord_reconciliations_total{namespace=~"materialize"}
[5m]
)
)
materialize.operator.reconciliation.failures.total #
Reconciliation passes that returned an error over the selected time range.sum(
increase(
orchestratord_reconciliations_total{
namespace=~"materialize", outcome="failed"
}
[1h]
)
)
materialize.operator.reconciliation.outcomes #
What reconciliation passes concluded, by outcome. The shape of a rollout:waiting climbs while the new generation’s pods come up, then
gives way to applied when they are ready.sum by (outcome) (
rate(
orchestratord_reconciliations_total{namespace=~"materialize"}
[5m]
)
)
materialize.operator.reconciliation.failures.by_controller #
Failing passes broken out by which controller failed and which of its entry points was running — the first question after “something is failing”.sum by (controller, event_type) (
rate(
orchestratord_reconciliations_total{
namespace=~"materialize", outcome="failed"
}
[5m]
)
)
materialize.operator.reconciliation.duration #
How long one reconciliation pass takes, at the 50th, 90th and 99th percentiles.histogram_quantile(0.5, sum by (le) (
rate(
orchestratord_reconciliation_duration_seconds_bucket{
namespace=~"materialize"
}
[5m]
)
))
histogram_quantile(0.9, sum by (le) (
rate(
orchestratord_reconciliation_duration_seconds_bucket{
namespace=~"materialize"
}
[5m]
)
))
histogram_quantile(0.99, sum by (le) (
rate(
orchestratord_reconciliation_duration_seconds_bucket{
namespace=~"materialize"
}
[5m]
)
))
materialize.operator.reconciliation.step.duration.p99 #
The slowest phase of a reconciliation pass, at the 99th percentile per step — where the time in a pass actually goes.histogram_quantile(0.99, sum by (le, step) (
rate(
orchestratord_reconciliation_step_duration_seconds_bucket{
namespace=~"materialize"
}
[5m]
)
))
materialize.operator.reconciliation.steps.rate #
Which phases of reconciliation are running, and how often. A rollout moves through these in order, so the set that is active says where the operator has got to.sum by (step) (
rate(
orchestratord_reconciliation_steps_total{namespace=~"materialize"}
[5m]
)
)
materialize.operator.reconciliation.steps.incomplete #
Steps that did not complete, by step and by how they ended — the panel that turns “reconciliation is failing” into “reconciliation is failing here”.sum by (step, outcome) (
rate(
orchestratord_reconciliation_steps_total{
namespace=~"materialize", outcome=~"failed|abandoned"
}
[5m]
)
)
materialize-persist#
Object storage (the persist bucket), from the vantage point of the Materialize processes that use it.
Persist is Materialize’s storage layer. Every durable collection — a table, a
source, a materialized view, the catalog itself — is a persist shard: data
files in object storage, plus a small record of the shard’s current state in
the metadata database. That second half is materialize-consensus.yaml; this
file is the first. Writers upload new data as immutable parts, compaction
merges small parts into larger ones and deletes what it replaced, and readers
fetch parts — through an in-memory cache — to hydrate and to serve queries.
Everything here is the client’s measurement: what a Materialize process
experienced, including the network between it and the store. It is identical
on S3, GCS (which Materialize reaches through GCS’s S3-compatible API), Azure
Blob and any S3-compatible on-premise store, and needs no cloud credentials.
What the bucket itself reports — its size, its object count, reclaimable
waste — is on infra-cloud.
Label families#
op on mz_persist_external_* is the operation: blob_get, blob_set,
blob_delete, blob_list_keys and restore. The consensus_* values of
the same label are the metadata database. mz_persist_external_op_latency is
a histogram for blob_get and blob_set only; the rest have a mean from
_seconds / _started_count.
mz_persist_s3_* is the S3 client beneath persist, used on S3, GCS and
S3-compatible stores but not on Azure Blob, which Materialize reaches with
Azure’s own client. The counters are registered on every install, though, so
on Azure they exist and read zero; queries on them gate on the client making
calls. It is the only family carrying the store’s error code (created on
first error, so absent until then), and its op values are raw API calls
(get_part, set_single, delete_object, …) rather than persist
operations.
op on mz_persist_retry_* is a retry loop named after the work. The ones
in materialize.persist.retries.by_operation wrap object-storage calls; the
consensus file owns the rest, and next_listen_batch and snapshot are
polling loops that must never be drawn as retries.
The mz_persist_shard_usage_* families are per shard, and each process
reports every shard it has open, so any total over them takes a max by (shard) first. environmentd has every shard open, so the deduplicated total
is the environment’s.
Every mean and share here is a ratio of two rates wrapped in (…) >= 0, for
the reason materialize-consensus.yaml gives: a window with no calls divides
zero by zero.
materialize.persist.health.failed_ops #
Object storage calls that failed, per second, across every Materialize process in this environment.sum(rate(mz_persist_external_failed_count{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m]))
materialize.persist.health.read_latency_p99 #
How long the slowest 1% of reads from object storage take.histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_get"}[5m])
)
)
materialize.persist.health.write_latency_p99 #
How long the slowest 1% of writes to object storage take.histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_set"}[5m])
)
)
materialize.persist.health.write_stalls #
Writes that had to wait because too many uploads were already in flight, per second.sum(rate(mz_persist_user_write_stall_count{materialize_cloud_organization_name=~".*"}[5m]))
+ sum(rate(mz_persist_compaction_write_stall_count{materialize_cloud_organization_name=~".*"}[5m]))
materialize.persist.health.compaction_failures #
Compactions that failed or timed out, per second.sum(rate(mz_persist_compaction_failed{materialize_cloud_organization_name=~".*"}[5m]))
+ sum(rate(mz_persist_compaction_timed_out{materialize_cloud_organization_name=~".*"}[5m]))
materialize.persist.health.stored #
Data this environment’s current state refers to in object storage.sum(max by (shard) (mz_persist_shard_usage_current_state_batches_bytes{materialize_cloud_organization_name=~".*"}))
+ sum(max by (shard) (mz_persist_shard_usage_current_state_rollups_bytes{materialize_cloud_organization_name=~".*"}))
materialize.persist.latency.read_write #
Read and write latency to object storage: the median and the slowest 1% of each.histogram_quantile(0.50,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_get"}[5m])
)
)
histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_get"}[5m])
)
)
histogram_quantile(0.50,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_set"}[5m])
)
)
histogram_quantile(0.99,
sum by (le) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_set"}[5m])
)
)
materialize.persist.latency.round_trip #
The round trip from each Materialize process to object storage, as that process last measured it.max by (app, cluster_environmentd_materialize_cloud_cluster_id) (
mz_persist_external_rtt_latency{materialize_cloud_organization_name=~".*", external="blob"}
)
materialize.persist.failures.by_operation #
Failed object storage calls per second, by operation.sum by (op) (
rate(mz_persist_external_failed_count{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m])
)
materialize.persist.failures.store_errors #
Errors the object store returned, per second, by API call and error code.sum by (op, code) (
rate(mz_persist_s3_errors{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.persist.failures.timeouts #
Object storage requests that timed out, per second, by where they timed out.sum(rate(mz_persist_s3_connect_timeouts{materialize_cloud_organization_name=~".*"}[5m]))
and on () (sum(rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])) > 0)
sum(rate(mz_persist_s3_read_timeouts{materialize_cloud_organization_name=~".*"}[5m]))
and on () (sum(rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])) > 0)
sum(rate(mz_persist_s3_operation_attempt_timeouts{materialize_cloud_organization_name=~".*"}[5m]))
and on () (sum(rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])) > 0)
sum(rate(mz_persist_s3_operation_timeouts{materialize_cloud_organization_name=~".*"}[5m]))
and on () (sum(rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])) > 0)
materialize.persist.logs.retries #
What Materialize logged while retrying calls to object storage, with the store’s own error text.materialize.persist.ops.by_type #
Object storage calls per second, by operation.sum by (op) (
rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m])
)
materialize.persist.ops.throughput #
Bytes per second read from and written to object storage, by operation.sum by (op) (
rate(mz_persist_external_bytes_count{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m])
)
materialize.persist.ops.requests #
Requests sent to the object store’s API per second, by call — the count a cloud provider bills.sum by (op) (
rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])
)
and on () (sum(rate(mz_persist_s3_operations{materialize_cloud_organization_name=~".*"}[5m])) > 0)
materialize.persist.latency.mean_by_operation #
Average time per object storage call, by operation.(
sum by (op) (
rate(mz_persist_external_seconds{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m])
)
/
sum by (op) (
rate(mz_persist_external_started_count{materialize_cloud_organization_name=~".*", op=~"blob_.*"}[5m])
)
) >= 0
materialize.persist.latency.read_by_process #
Slowest 1% of reads from object storage, by the process making them.histogram_quantile(0.99,
sum by (le, app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_external_op_latency_bucket{materialize_cloud_organization_name=~".*", op="blob_get"}[5m])
)
)
materialize.persist.retries.by_operation #
Object storage calls retried after a failure, per second, by the work that was retrying.sum by (op) (
rate(mz_persist_retry_retries_count{materialize_cloud_organization_name=~".*", op=~"batch::set|batch::delete|fetch_batch::get|rollup::get|rollup::set|rollup::delete|hollow_run::get|hollow_run::set|blob::open|compaction_noop::delete|storage_usage::shard_size"}[5m])
)
materialize.persist.cache.hits #
Share of bytes read that the in-memory cache served, rather than object storage.(
sum(rate(mz_persist_blob_cache_hits_bytes{materialize_cloud_organization_name=~".*"}[5m]))
/
(
sum(rate(mz_persist_blob_cache_hits_bytes{materialize_cloud_organization_name=~".*"}[5m]))
+ sum(rate(mz_persist_external_bytes_count{materialize_cloud_organization_name=~".*", op="blob_get"}[5m]))
)
) >= 0
materialize.persist.hedging #
Reads that were slow enough to send a second, duplicate request, and how many of those the duplicate won, per second.sum(rate(mz_persist_blob_hedges_fired{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_blob_hedges_won{materialize_cloud_organization_name=~".*"}[5m]))
materialize.persist.compaction.outcomes #
What happened to compaction requests, per second.sum(rate(mz_persist_compaction_requested{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_started{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_applied{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_failed{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_timed_out{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_dropped{materialize_cloud_organization_name=~".*"}[5m]))
materialize.persist.compaction.queue_wait #
Average time a compaction request waits in the queue before it starts.(
sum(rate(mz_persist_compaction_queued_seconds{materialize_cloud_organization_name=~".*"}[5m]))
/
sum(rate(mz_persist_compaction_started{materialize_cloud_organization_name=~".*"}[5m]))
) >= 0
materialize.persist.compaction.busy #
Compactions running at once, on average, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_compaction_seconds{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.persist.writes.amplification #
Bytes per second of new data written to object storage, beside the bytes compaction rewrites.sum(rate(mz_persist_user_bytes{materialize_cloud_organization_name=~".*"}[5m]))
sum(rate(mz_persist_compaction_bytes{materialize_cloud_organization_name=~".*"}[5m]))
materialize.persist.writes.stalls_by_process #
Writes that waited for in-flight uploads to finish, per second, by process.sum by (app, cluster_environmentd_materialize_cloud_cluster_id) (
rate(mz_persist_user_write_stall_count{materialize_cloud_organization_name=~".*"}[5m])
+ rate(mz_persist_compaction_write_stall_count{materialize_cloud_organization_name=~".*"}[5m])
)
materialize.persist.usage.by_state #
What this environment’s data in object storage is for.sum(max by (shard) (mz_persist_shard_usage_current_state_batches_bytes{materialize_cloud_organization_name=~".*"}))
sum(max by (shard) (mz_persist_shard_usage_current_state_rollups_bytes{materialize_cloud_organization_name=~".*"}))
sum(max by (shard) (mz_persist_shard_usage_referenced_not_current_state_bytes{materialize_cloud_organization_name=~".*"}))
sum(max by (shard) (mz_persist_shard_usage_not_leaked_not_referenced_bytes{materialize_cloud_organization_name=~".*"}))
sum(max by (shard) (mz_persist_shard_usage_leaked_bytes{materialize_cloud_organization_name=~".*"}))
materialize.persist.usage.leaked #
Data in object storage that nothing refers to and that cleanup did not remove.sum(max by (shard) (mz_persist_shard_usage_leaked_bytes{materialize_cloud_organization_name=~".*"}))
materialize.persist.usage.largest #
The ten objects using the most object storage, now.topk(10,
max by (shard) (mz_persist_shard_usage_current_state_batches_bytes{materialize_cloud_organization_name=~".*"})
+ max by (shard) (mz_persist_shard_usage_current_state_rollups_bytes{materialize_cloud_organization_name=~".*"})
)
materialize.persist.usage.held #
The ten objects keeping the most data only because a reader still needs an older version of it.topk(10, max by (shard) (mz_persist_shard_usage_referenced_not_current_state_bytes{materialize_cloud_organization_name=~".*"}))
materialize-storage#
Sources and sinks for a Materialize deployment — catalog shape, throughput, lag, and upstream/downstream health. Adapted from the Overview dashboard’s “Sources and Sinks” tab.
The clusterd-side throughput/lag/error metrics (mz_source_* / mz_sink_) carry
the long-form cluster_environmentd_materialize_cloud_ id labels. These queries
assume one Prometheus job per clusterd endpoint; if the same endpoint is
scraped by several jobs, a plain sum-rate reads N× — dedupe the job at the
deployment (fix the scrape config, or wrap the inner rate in max without (job)) rather than baking it into the canonical query.
materialize.storage.sources.count #
Active sources in the catalog — each is a continuous ingestion connection from an external system (Kafka, Postgres, MySQL, S3, …), so this is roughly how many upstream feeds the environment maintains. Counts distinct source objects (the hidden per-source_progress
subsources are excluded), matching mz_sources.count(group by (id) (mz_storage_objects{materialize_cloud_organization_name=~".*", type="source"}))
default_zero(count_not_null(avg:${mzSqlPrefix}storage_objects{${mzEnvironmentFilter}, type:source} by {id}))
materialize.storage.sinks.count #
Active sinks in the catalog — each emits the results of a materialized view or query to an external system (Kafka, Iceberg, …). Counts distinct sink objects (excluding_progress subsources), matching mz_sinks.count(group by (id) (mz_storage_objects{materialize_cloud_organization_name=~".*", type="sink"}))
default_zero(count_not_null(avg:${mzSqlPrefix}storage_objects{${mzEnvironmentFilter}, type:sink} by {id}))
materialize.storage.tables.count #
User-created tables in the catalog. Tables are write-once-read-many;INSERTs feed dataflows downstream. Mostly a catalog-shape signal — for
actual ingest activity look at source throughput.max(sum by (instance) (mz_tables_count{materialize_cloud_organization_name=~".*"}))
default_zero(sum:${mzSqlPrefix}tables_count{${mzEnvironmentFilter}} by {instance})
materialize.storage.sources.by_type #
Sources by connector type (kafka / postgres / mysql / …) — what flavors of upstream feed make up the ingest workload. Most environments concentrate on one or two.count by (object_type) (
group by (id, object_type) (
mz_storage_objects{materialize_cloud_organization_name=~".*", type="source"}
)
) > 0
sum:${mzSqlPrefix}storage_objects{${mzEnvironmentFilter}, type:source} by {object_type}
materialize.storage.sources.catalog #
A catalog of sources — one row per source (by name) with its connector type, envelope, and the cluster it ingests on. The metric-side “what sources do I have” reference.group by (id, object_type, connection_type, envelope_type, cluster_id) (
mz_storage_objects{materialize_cloud_organization_name=~".*", type="source"}
)
avg:${mzSqlPrefix}storage_objects{${mzEnvironmentFilter}, type:source}
by {id,object_type,connection_type,envelope_type,cluster_id}
materialize.storage.sources.bytes_received #
Inbound throughput per primary source — bytes/second pulled from upstream. Subsources (e.g. per-table Postgres replication) roll up to their primary, so each line is one logical source.sum by (parent_source_id) (
max without (job) (rate(mz_source_bytes_received{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))
) > 0
sum:mz_source_bytes_received{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {parent_source_id}.as_rate()
materialize.storage.sources.ingestion_by_replica #
Messages ingested per second, split per source AND replica. Replicas read their upstream independently and should track together.sum by (parent_source_id, cluster_environmentd_materialize_cloud_replica_id) (
max without (job) (rate(mz_source_messages_received{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))
)
sum:mz_source_messages_received{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {parent_source_id,cluster_environmentd_materialize_cloud_replica_id}.as_rate()
materialize.storage.sources.upstream_errors #
Per-source upstream health, with two complementary signals — both nominal at 0, so an empty panel is healthy and any series means a source needs attention.sum by (source_id) (
max without (job) (rate(mz_source_offset_commit_failures{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))
) > 0
(
max by (source_id) (mz_source_offset_committed{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"})
> bool max by (source_id) (mz_source_offset_known{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"})
) > 0
sum:mz_source_offset_commit_failures{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {source_id}.as_rate()
max:mz_source_offset_committed{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {source_id}
- max:mz_source_offset_known{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {source_id}
materialize.storage.sinks.by_type #
Sinks by (type, envelope) — e.g.kafka / upsert, kafka / debezium,
iceberg / upsert. The envelope is how Materialize encodes changes:
upsert writes the latest value per key, debezium writes change
events with old+new values.count by (object_type, envelope_type) (
group by (id, object_type, envelope_type) (
mz_storage_objects{materialize_cloud_organization_name=~".*", type="sink"}
)
) > 0
sum:${mzSqlPrefix}storage_objects{${mzEnvironmentFilter}, type:sink} by {object_type,envelope_type}
materialize.storage.sinks.throughput #
Outbound throughput per sink — bytes/second successfully committed to the downstream system (Kafka broker, Iceberg catalog, …).sum by (sink_id) (
max without (job) (rate(mz_sink_bytes_committed{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))
) > 0
sum:mz_sink_bytes_committed{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
materialize.storage.sinks.lag #
Bytes staged for a sink but not yet committed downstream — an in-flight queue depth in bytes. Oscillates around a small value in normal operation as commits happen periodically.clamp_min(
sum by (sink_id) (max without (job) (mz_sink_bytes_staged{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}))
- sum by (sink_id) (max without (job) (mz_sink_bytes_committed{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"})),
0
)
clamp_min(
sum:mz_sink_bytes_staged{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}
- sum:mz_sink_bytes_committed{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id},
0
)
materialize.storage.sinks.iceberg.commit_latency #
Iceberg commit-duration percentiles (p50/p90/p99) — how long eachCOMMIT against the Iceberg catalog takes (write a snapshot manifest,
ask the catalog to atomically swap it in).histogram_quantile(0.50, sum by (le) (max without (job) (rate(mz_sink_iceberg_commit_duration_seconds_bucket{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))))
histogram_quantile(0.90, sum by (le) (max without (job) (rate(mz_sink_iceberg_commit_duration_seconds_bucket{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))))
histogram_quantile(0.99, sum by (le) (max without (job) (rate(mz_sink_iceberg_commit_duration_seconds_bucket{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m]))))
p50:mz_sink_iceberg_commit_duration_seconds{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
p90:mz_sink_iceberg_commit_duration_seconds{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
p99:mz_sink_iceberg_commit_duration_seconds{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}}
materialize.storage.sinks.iceberg.commit_failures #
Per-sink rate of failed and conflicting Iceberg commits. Conflicts (concurrent-writer races on the snapshot pointer) are recoverable — Materialize retries — but a high rate means something else is writing the same Iceberg table; failures are commit-side errors (network, auth, schema).sum by (sink_id) (max without (job) (rate(mz_sink_iceberg_commit_failures{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum by (sink_id) (max without (job) (rate(mz_sink_iceberg_commit_conflicts{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum:mz_sink_iceberg_commit_failures{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
sum:mz_sink_iceberg_commit_conflicts{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
materialize.storage.sinks.iceberg.file_rate #
Per-sink rate of files and snapshots written to Iceberg. Each commit produces one snapshot with data files (new rows) and delete files (tombstones for upserts). The data:delete ratio reflects your workload — pure-insert sinks produce ~0 deletes; upsert-heavy ones roughly 1:1.sum by (sink_id) (max without (job) (rate(mz_sink_iceberg_data_files_written{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum by (sink_id) (max without (job) (rate(mz_sink_iceberg_delete_files_written{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum by (sink_id) (max without (job) (rate(mz_sink_iceberg_snapshots_committed{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum:mz_sink_iceberg_data_files_written{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
sum:mz_sink_iceberg_delete_files_written{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
sum:mz_sink_iceberg_snapshots_committed{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
materialize.storage.sinks.kafka.tx_errors #
Per-sink rate of TX errors from the librdkafka client — each is one failed produce-request against the broker.sum by (sink_id) (max without (job) (rate(mz_sink_rdkafka_txerrs{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum:mz_sink_rdkafka_txerrs{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
materialize.storage.sinks.kafka.output_buffer #
Messages sitting in the librdkafka output buffer, waiting to be sent to the broker. Normal buffer fluctuates briefly as messages flow through.sum by (sink_id) (max without (job) (mz_sink_rdkafka_outbuf_msg_cnt{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}))
sum:mz_sink_rdkafka_outbuf_msg_cnt{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}
materialize.storage.sinks.kafka.connect_rate #
Connect and disconnect events per sink against the Kafka broker. Healthy connections are persistent — a couple of connects at startup and zero disconnects afterward.sum by (sink_id) (max without (job) (rate(mz_sink_rdkafka_connects{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum by (sink_id) (max without (job) (rate(mz_sink_rdkafka_disconnects{materialize_cloud_organization_name=~".*", cluster_environmentd_materialize_cloud_cluster_id=~".*", cluster_environmentd_materialize_cloud_replica_id=~".*"}[5m])))
sum:mz_sink_rdkafka_connects{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
sum:mz_sink_rdkafka_disconnects{${mzEnvironmentFilter}, cluster_environmentd_materialize_cloud_cluster_id:${mzClusterList}, cluster_environmentd_materialize_cloud_replica_id:${mzReplicaList}} by {sink_id}.as_rate()
materialize-workload-alerts#
Alerting rules for the workloads running on Materialize.
Every alert here carries audience: workload, so a route can send them to
the people who own the clusters rather than to whoever runs the deployment.
Each is scoped to user clusters (u*); the same condition on a system
cluster is a platform alert.
materialize.clusters.names #
Each cluster’s name, by namespace and cluster id; the map the alerts join through to name a cluster.group by (namespace, cluster_id, name) (mz_cluster_info{materialize_cloud_organization_name=~".*"})
max:mz_cluster_info{${mzEnvironmentFilter}} by {namespace,cluster_id,name}
node-debug#
The breakdowns you reach for once node-health.yaml has told you a node is in
trouble: which mode the CPU is in, where the memory went, which device is
slow, and where the packets are being lost. Adapted from the Node Exporter
Full dashboard (https://grafana.com/grafana/dashboards/1860, revision 45).
Split from node-health.yaml on tier rather than on subject. These sit at
recommended, so a deployment collecting only the essential tier still gets
the health surface and pays nothing for the detail. Nothing here should back
an alert — if something here is worth paging on, it belongs in
node-health.yaml instead.
The conventions are the same as node-health.yaml: instance=~"$nodeList" rather
than =, every query wrapped in max by (instance, ...) (or min where low
is the bad direction) so a second scrape job cannot double-count, inner
aggregations carrying by (instance, job) so the outer wrapper is what
collapses job, and %%{interval} as the rate window.
Only collectors this chart’s allowlist enables are referenced. Notable
omissions, because the dashboard has panels for them and they will render
empty: node_processes_* (processes collector), node_interrupts_total
(interrupts), node_tcp_connection_states (tcpstat), node_systemd_*
(systemd), and node_textfile_scrape_error (textfile).
node.debug.cpu.by_mode #
CPU time by mode — system, user, iowait, and the interrupt modes — averaged across cores.max by (instance) (
avg by (instance, job) (
rate(node_cpu_seconds_total{mode="system", instance=~"$nodeList"}[5m])
)
)
max by (instance) (
avg by (instance, job) (
rate(node_cpu_seconds_total{mode="user", instance=~"$nodeList"}[5m])
)
)
max by (instance) (
avg by (instance, job) (
rate(node_cpu_seconds_total{mode="iowait", instance=~"$nodeList"}[5m])
)
)
max by (instance) (
avg by (instance, job) (
sum without (mode) (
rate(node_cpu_seconds_total{mode=~".*irq", instance=~"$nodeList"}[5m])
)
)
)
max by (instance) (
avg by (instance, job) (
rate(node_cpu_seconds_total{mode="steal", instance=~"$nodeList"}[5m])
)
)
avg:node_cpu_seconds_total{instance:$nodeList, mode:system} by {instance}.as_rate()
avg:node_cpu_seconds_total{instance:$nodeList, mode:user} by {instance}.as_rate()
avg:node_cpu_seconds_total{instance:$nodeList, mode:iowait} by {instance}.as_rate()
avg:node_cpu_seconds_total{instance:$nodeList, mode:*irq} by {instance}.as_rate()
avg:node_cpu_seconds_total{instance:$nodeList, mode:steal} by {instance}.as_rate()
node.debug.cpu.per_core #
Non-idle CPU time per core, so a single saturated core is visible.1 - max by (instance, cpu) (
rate(node_cpu_seconds_total{mode="idle", instance=~"$nodeList"}[5m])
)
1 - max:node_cpu_seconds_total{instance:$nodeList, mode:idle} by {instance,cpu}.as_rate()
node.debug.schedstat.waiting #
Time tasks spent runnable but not running, per core, from/proc/schedstat.max by (instance, cpu) (
rate(node_schedstat_waiting_seconds_total{instance=~"$nodeList"}[5m])
)
max:node_schedstat_waiting_seconds_total{instance:$nodeList} by {instance,cpu}.as_rate()
node.debug.context_switches #
Context switches and hardware interrupts per second.max by (instance) (
rate(node_context_switches_total{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_intr_total{instance=~"$nodeList"}[5m])
)
max:node_context_switches_total{instance:$nodeList} by {instance}.as_rate()
max:node_intr_total{instance:$nodeList} by {instance}.as_rate()
node.debug.memory.breakdown #
Where RAM went: total, used by processes, reclaimable cache, and free.max by (instance) (node_memory_MemTotal_bytes{instance=~"$nodeList"})
max by (instance) (
node_memory_MemTotal_bytes{instance=~"$nodeList"}
- node_memory_MemFree_bytes{instance=~"$nodeList"}
- node_memory_Cached_bytes{instance=~"$nodeList"}
- node_memory_Buffers_bytes{instance=~"$nodeList"}
- node_memory_SReclaimable_bytes{instance=~"$nodeList"}
)
max by (instance) (
node_memory_Cached_bytes{instance=~"$nodeList"}
+ node_memory_Buffers_bytes{instance=~"$nodeList"}
+ node_memory_SReclaimable_bytes{instance=~"$nodeList"}
)
min by (instance) (node_memory_MemFree_bytes{instance=~"$nodeList"})
max:node_memory_MemTotal_bytes{instance:$nodeList} by {instance}
max:node_memory_MemTotal_bytes{instance:$nodeList} by {instance}
- max:node_memory_MemFree_bytes{instance:$nodeList} by {instance}
- max:node_memory_Cached_bytes{instance:$nodeList} by {instance}
- max:node_memory_Buffers_bytes{instance:$nodeList} by {instance}
- max:node_memory_SReclaimable_bytes{instance:$nodeList} by {instance}
max:node_memory_Cached_bytes{instance:$nodeList} by {instance}
+ max:node_memory_Buffers_bytes{instance:$nodeList} by {instance}
+ max:node_memory_SReclaimable_bytes{instance:$nodeList} by {instance}
min:node_memory_MemFree_bytes{instance:$nodeList} by {instance}
node.debug.memory.kernel #
Kernel-side memory: slab total, reclaimable and unreclaimable slab, and committed address space.max by (instance) (node_memory_Slab_bytes{instance=~"$nodeList"})
max by (instance) (node_memory_SReclaimable_bytes{instance=~"$nodeList"})
max by (instance) (node_memory_SUnreclaim_bytes{instance=~"$nodeList"})
max by (instance) (node_memory_Committed_AS_bytes{instance=~"$nodeList"})
max:node_memory_Slab_bytes{instance:$nodeList} by {instance}
max:node_memory_SReclaimable_bytes{instance:$nodeList} by {instance}
max:node_memory_SUnreclaim_bytes{instance:$nodeList} by {instance}
max:node_memory_Committed_AS_bytes{instance:$nodeList} by {instance}
node.debug.memory.page_faults #
Total and major page faults per second.max by (instance) (
rate(node_vmstat_pgfault{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_vmstat_pgmajfault{instance=~"$nodeList"}[5m])
)
max:node_vmstat_pgfault{instance:$nodeList} by {instance}.as_rate()
max:node_vmstat_pgmajfault{instance:$nodeList} by {instance}.as_rate()
node.debug.memory.reclaim #
Pages scanned and reclaimed per second, split by who did the reclaiming:kswapd (background) or direct (an allocating thread).max by (instance) (
rate(node_vmstat_pgscan_kswapd{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_vmstat_pgscan_direct{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_vmstat_pgsteal_kswapd{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_vmstat_pgsteal_direct{instance=~"$nodeList"}[5m])
)
max:node_vmstat_pgscan_kswapd{instance:$nodeList} by {instance}.as_rate()
max:node_vmstat_pgscan_direct{instance:$nodeList} by {instance}.as_rate()
max:node_vmstat_pgsteal_kswapd{instance:$nodeList} by {instance}.as_rate()
max:node_vmstat_pgsteal_direct{instance:$nodeList} by {instance}.as_rate()
node.debug.disk.iops #
Completed reads and writes per second, per device.max by (instance, device) (
rate(node_disk_reads_completed_total{instance=~"$nodeList"}[5m])
)
max by (instance, device) (
rate(node_disk_writes_completed_total{instance=~"$nodeList"}[5m])
)
max:node_disk_reads_completed_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_disk_writes_completed_total{instance:$nodeList} by {instance,device}.as_rate()
node.debug.disk.throughput #
Bytes read and written per second, per device.max by (instance, device) (
rate(node_disk_read_bytes_total{instance=~"$nodeList"}[5m])
)
max by (instance, device) (
rate(node_disk_written_bytes_total{instance=~"$nodeList"}[5m])
)
max:node_disk_read_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_disk_written_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
node.debug.disk.latency #
Average time per completed read and per completed write, per device.max by (instance, device) (
(rate(node_disk_reads_completed_total{instance=~"$nodeList"}[5m]) > bool 0)
*
(
rate(node_disk_read_time_seconds_total{instance=~"$nodeList"}[5m])
/ rate(node_disk_reads_completed_total{instance=~"$nodeList"}[5m])
)
)
max by (instance, device) (
(rate(node_disk_writes_completed_total{instance=~"$nodeList"}[5m]) > bool 0)
*
(
rate(node_disk_write_time_seconds_total{instance=~"$nodeList"}[5m])
/ rate(node_disk_writes_completed_total{instance=~"$nodeList"}[5m])
)
)
max:node_disk_read_time_seconds_total{instance:$nodeList} by {instance,device}.as_rate()
/ max:node_disk_reads_completed_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_disk_write_time_seconds_total{instance:$nodeList} by {instance,device}.as_rate()
/ max:node_disk_writes_completed_total{instance:$nodeList} by {instance,device}.as_rate()
node.debug.disk.queue_depth #
Average I/O queue depth per device.max by (instance, device) (
rate(node_disk_io_time_weighted_seconds_total{instance=~"$nodeList"}[5m])
)
max:node_disk_io_time_weighted_seconds_total{instance:$nodeList} by {instance,device}.as_rate()
node.debug.filesystem.inodes.available.ratio #
Fraction of inodes still free, per mountpoint.min by (instance, mountpoint) (
node_filesystem_files_free{instance=~"$nodeList", fstype!="rootfs"}
)
/
max by (instance, mountpoint) (
node_filesystem_files{instance=~"$nodeList", fstype!="rootfs"}
)
min:node_filesystem_files_free{instance:$nodeList, !fstype:rootfs} by {instance,mountpoint}
/ max:node_filesystem_files{instance:$nodeList, !fstype:rootfs} by {instance,mountpoint}
node.debug.network.throughput #
Bytes received and transmitted per second, per interface.max by (instance, device) (
rate(node_network_receive_bytes_total{instance=~"$nodeList"}[5m])
)
max by (instance, device) (
rate(node_network_transmit_bytes_total{instance=~"$nodeList"}[5m])
)
max:node_network_receive_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_network_transmit_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
node.debug.network.saturation #
Receive and transmit throughput as a fraction of the interface’s reported link speed.max by (instance, device) (
(node_network_speed_bytes{instance=~"$nodeList"} > bool 0)
*
(
rate(node_network_receive_bytes_total{instance=~"$nodeList"}[5m])
/ node_network_speed_bytes{instance=~"$nodeList"}
)
)
max by (instance, device) (
(node_network_speed_bytes{instance=~"$nodeList"} > bool 0)
*
(
rate(node_network_transmit_bytes_total{instance=~"$nodeList"}[5m])
/ node_network_speed_bytes{instance=~"$nodeList"}
)
)
max:node_network_receive_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
/ max:node_network_speed_bytes{instance:$nodeList} by {instance,device}
max:node_network_transmit_bytes_total{instance:$nodeList} by {instance,device}.as_rate()
/ max:node_network_speed_bytes{instance:$nodeList} by {instance,device}
node.debug.network.operstate #
Whether each interface is operationally up, and whether it has carrier.min by (instance, device) (
node_network_up{instance=~"$nodeList"}
)
min by (instance, device) (
node_network_carrier{instance=~"$nodeList"}
)
min:node_network_up{instance:$nodeList} by {instance,device}
min:node_network_carrier{instance:$nodeList} by {instance,device}
node.debug.softnet.processed #
Packets processed by the network softirq path, per CPU.max by (instance, cpu) (
rate(node_softnet_processed_total{instance=~"$nodeList"}[5m])
)
max:node_softnet_processed_total{instance:$nodeList} by {instance,cpu}.as_rate()
node.debug.softnet.dropped #
Packets dropped in the network softirq path because the backlog queue was full, per CPU.max by (instance, cpu) (
rate(node_softnet_dropped_total{instance=~"$nodeList"}[5m])
)
max:node_softnet_dropped_total{instance:$nodeList} by {instance,cpu}.as_rate()
node.debug.softnet.squeezed #
Times the softirq handler exhausted its budget with work still queued, per CPU.max by (instance, cpu) (
rate(node_softnet_times_squeezed_total{instance=~"$nodeList"}[5m])
)
max:node_softnet_times_squeezed_total{instance:$nodeList} by {instance,cpu}.as_rate()
node.debug.sockets.tcp #
TCP sockets by state: in use, allocated, orphaned, and TIME_WAIT.max by (instance) (node_sockstat_TCP_inuse{instance=~"$nodeList"})
max by (instance) (node_sockstat_TCP_alloc{instance=~"$nodeList"})
max by (instance) (node_sockstat_TCP_orphan{instance=~"$nodeList"})
max by (instance) (node_sockstat_TCP_tw{instance=~"$nodeList"})
max:node_sockstat_TCP_inuse{instance:$nodeList} by {instance}
max:node_sockstat_TCP_alloc{instance:$nodeList} by {instance}
max:node_sockstat_TCP_orphan{instance:$nodeList} by {instance}
max:node_sockstat_TCP_tw{instance:$nodeList} by {instance}
node.debug.sockets.memory #
Kernel socket buffer memory held by TCP and UDP.max by (instance) (node_sockstat_TCP_mem_bytes{instance=~"$nodeList"})
max by (instance) (node_sockstat_UDP_mem_bytes{instance=~"$nodeList"})
max:node_sockstat_TCP_mem_bytes{instance:$nodeList} by {instance}
max:node_sockstat_UDP_mem_bytes{instance:$nodeList} by {instance}
node.debug.tcp.retransmits #
TCP segment retransmits and SYN retransmits per second, against total segments out.max by (instance) (
rate(node_netstat_Tcp_RetransSegs{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_TcpExt_TCPSynRetrans{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_Tcp_OutSegs{instance=~"$nodeList"}[5m])
)
max:node_netstat_Tcp_RetransSegs{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_TcpExt_TCPSynRetrans{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_Tcp_OutSegs{instance:$nodeList} by {instance}.as_rate()
node.debug.tcp.errors #
TCP listen-queue overflows, listen drops, receive-queue drops, and timeouts per second.max by (instance) (
rate(node_netstat_TcpExt_ListenOverflows{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_TcpExt_ListenDrops{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_TcpExt_TCPRcvQDrop{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_TcpExt_TCPTimeouts{instance=~"$nodeList"}[5m])
)
max:node_netstat_TcpExt_ListenOverflows{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_TcpExt_ListenDrops{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_TcpExt_TCPRcvQDrop{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_TcpExt_TCPTimeouts{instance:$nodeList} by {instance}.as_rate()
node.debug.udp.errors #
UDP receive errors, receive-buffer errors, and packets to no listening port, per second.max by (instance) (
rate(node_netstat_Udp_InErrors{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_Udp_RcvbufErrors{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_netstat_Udp_NoPorts{instance=~"$nodeList"}[5m])
)
max:node_netstat_Udp_InErrors{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_Udp_RcvbufErrors{instance:$nodeList} by {instance}.as_rate()
max:node_netstat_Udp_NoPorts{instance:$nodeList} by {instance}.as_rate()
node.debug.udp.queues #
Bytes queued in UDP receive and transmit buffers.max by (instance) (
node_udp_queues{ip="v4", queue="rx", instance=~"$nodeList"}
)
max by (instance) (
node_udp_queues{ip="v4", queue="tx", instance=~"$nodeList"}
)
max:node_udp_queues{instance:$nodeList, ip:v4, queue:rx} by {instance}
max:node_udp_queues{instance:$nodeList, ip:v4, queue:tx} by {instance}
node.debug.arp.entries #
ARP table entries per interface.max by (instance, device) (
node_arp_entries{instance=~"$nodeList"}
)
max:node_arp_entries{instance:$nodeList} by {instance,device}
node.debug.time.sync_status #
Whether the kernel clock is synchronized (1) or NTP has given up (0).min by (instance) (
node_timex_sync_status{instance=~"$nodeList"}
)
min:node_timex_sync_status{instance:$nodeList} by {instance}
node.debug.time.drift #
Estimated clock offset, maximum error, and estimated error, in seconds.max by (instance) (node_timex_offset_seconds{instance=~"$nodeList"})
max by (instance) (node_timex_maxerror_seconds{instance=~"$nodeList"})
max by (instance) (node_timex_estimated_error_seconds{instance=~"$nodeList"})
max:node_timex_offset_seconds{instance:$nodeList} by {instance}
max:node_timex_maxerror_seconds{instance:$nodeList} by {instance}
max:node_timex_estimated_error_seconds{instance:$nodeList} by {instance}
node.debug.entropy.available #
Available entropy, against the pool size.min by (instance) (node_entropy_available_bits{instance=~"$nodeList"})
max by (instance) (node_entropy_pool_size_bits{instance=~"$nodeList"})
min:node_entropy_available_bits{instance:$nodeList} by {instance}
max:node_entropy_pool_size_bits{instance:$nodeList} by {instance}
node.debug.exporter.scrape_duration #
How long each node-exporter collector took on the last scrape.max by (instance, collector) (
node_scrape_collector_duration_seconds{instance=~"$nodeList"}
)
max:node_scrape_collector_duration_seconds{instance:$nodeList} by {instance,collector}
node-health#
Node-level health: the at-a-glance answers to “is this machine in trouble”, adapted from the Node Exporter Full dashboard (https://grafana.com/grafana/dashboards/1860, revision 45) and narrowed to what actually alerts.
These read node-exporter, NOT Materialize metrics. Only collectors this
chart’s allowlist enables are referenced — see the Node Exporter section of
the chart’s values reference for the list and the reasoning. The deeper
breakdowns live in node-debug.yaml at the recommended tier.
Three conventions apply to every query here:
instance=~"$nodeList", neverinstance="$nodeList". A regex match makes the selector work unchanged whether the dashboard variable resolves to one node or many, so the same query backs a single-node view and a fleet view.Every query is wrapped in
max by (instance, ...), which dropsjob. A node should only ever be scraped by one job. If a second one appears — a pre-existing node-exporter alongside ours, or a migration with both running — the same series arrives twice under differentjoblabels, and everysum()silently doubles while every binary operation between two metrics loses its match. Aggregatingjobaway makes both failure modes impossible rather than merely unlikely.maxis the default;minwhere low is the bad direction (available memory, free space, collector success), so the aggregate always reports the worst case rather than hiding it behind a healthy duplicate.Where an inner aggregation is needed (averaging across CPUs, summing across devices), it carries
by (instance, job)and the outermax/mincollapsesjobafterwards. Aggregating both in one step would blend two jobs' readings into one number instead of picking one.
%%{interval} is the rate window, including its brackets.
node.cpu.utilization #
Fraction of CPU time the node spent doing anything other than idling, averaged across its cores.1 - max by (instance) (
avg by (instance, job) (
rate(node_cpu_seconds_total{mode="idle", instance=~"$nodeList"}[5m])
)
)
1 - avg:node_cpu_seconds_total{mode:idle, instance:$nodeList} by {instance}.as_rate()
node.load.normalized #
One-minute load average divided by the node’s core count, so it is comparable across instance sizes.max by (instance) (node_load1{instance=~"$nodeList"})
/
max by (instance) (
count by (instance, job) (
count by (instance, job, cpu) (node_cpu_seconds_total{instance=~"$nodeList"})
)
)
max:node_load1{instance:$nodeList} by {instance}
/ count_not_null(avg:node_cpu_seconds_total{mode:idle, instance:$nodeList} by {cpu})
node.cpu.pressure #
PSI: the fraction of wall time at least one task was stalled waiting for CPU.max by (instance) (
rate(node_pressure_cpu_waiting_seconds_total{instance=~"$nodeList"}[5m])
)
max:node_pressure_cpu_waiting_seconds_total{instance:$nodeList} by {instance}.as_rate()
node.memory.available.ratio #
Fraction of RAM the kernel estimates is available for new allocations without swapping, fromMemAvailable.min by (instance) (
node_memory_MemAvailable_bytes{instance=~"$nodeList"}
)
/
max by (instance) (
node_memory_MemTotal_bytes{instance=~"$nodeList"}
)
min:node_memory_MemAvailable_bytes{instance:$nodeList} by {instance}
/ max:node_memory_MemTotal_bytes{instance:$nodeList} by {instance}
node.memory.pressure #
PSI: the fraction of wall time at least one task was stalled on memory (waiting), and the fraction where every task was (stalled).max by (instance) (
rate(node_pressure_memory_waiting_seconds_total{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_pressure_memory_stalled_seconds_total{instance=~"$nodeList"}[5m])
)
max:node_pressure_memory_waiting_seconds_total{instance:$nodeList} by {instance}.as_rate()
max:node_pressure_memory_stalled_seconds_total{instance:$nodeList} by {instance}.as_rate()
node.swap.used.ratio #
Fraction of configured swap in use. Zero when the node has no swap configured, rather than returning no data.(
(
max by (instance) (node_memory_SwapTotal_bytes{instance=~"$nodeList"})
-
min by (instance) (node_memory_SwapFree_bytes{instance=~"$nodeList"})
)
/
max by (instance) (node_memory_SwapTotal_bytes{instance=~"$nodeList"})
)
and
(max by (instance) (node_memory_SwapTotal_bytes{instance=~"$nodeList"}) > 0)
(
max:node_memory_SwapTotal_bytes{instance:$nodeList} by {instance}
- min:node_memory_SwapFree_bytes{instance:$nodeList} by {instance}
)
/ max:node_memory_SwapTotal_bytes{instance:$nodeList} by {instance}
node.swap.activity #
Pages swapped in and out per second, fromvmstat.max by (instance) (
rate(node_vmstat_pswpin{instance=~"$nodeList"}[5m])
)
max by (instance) (
rate(node_vmstat_pswpout{instance=~"$nodeList"}[5m])
)
max:node_vmstat_pswpin{instance:$nodeList} by {instance}.as_rate()
max:node_vmstat_pswpout{instance:$nodeList} by {instance}.as_rate()
node.memory.oom_kills #
Rate of OOM-killer invocations on the node.max by (instance) (
rate(node_vmstat_oom_kill{instance=~"$nodeList"}[5m])
)
max:node_vmstat_oom_kill{instance:$nodeList} by {instance}.as_rate()
node.filesystem.available.ratio #
Fraction of each mounted filesystem still available, per mountpoint.min by (instance, mountpoint) (
node_filesystem_avail_bytes{instance=~"$nodeList", fstype!="rootfs"}
)
/
max by (instance, mountpoint) (
node_filesystem_size_bytes{instance=~"$nodeList", fstype!="rootfs"}
)
min:node_filesystem_avail_bytes{instance:$nodeList, !fstype:rootfs} by {instance,mountpoint}
/ max:node_filesystem_size_bytes{instance:$nodeList, !fstype:rootfs} by {instance,mountpoint}
node.filesystem.readonly #
Whether a filesystem has been remounted read-only, per mountpoint.max by (instance, mountpoint) (
node_filesystem_readonly{instance=~"$nodeList", fstype!="rootfs"}
)
max:node_filesystem_readonly{instance:$nodeList, !fstype:rootfs} by {instance,mountpoint}
node.disk.io_utilization #
Fraction of wall time each block device had at least one I/O in flight.max by (instance, device) (
rate(node_disk_io_time_seconds_total{instance=~"$nodeList"}[5m])
)
max:node_disk_io_time_seconds_total{instance:$nodeList} by {instance,device}.as_rate()
node.filefd.utilization #
Allocated file descriptors as a fraction of the system-wide maximum.max by (instance) (node_filefd_allocated{instance=~"$nodeList"})
/
max by (instance) (node_filefd_maximum{instance=~"$nodeList"})
max:node_filefd_allocated{instance:$nodeList} by {instance}
/ max:node_filefd_maximum{instance:$nodeList} by {instance}
node.network.errors #
Receive and transmit error rates per interface.max by (instance, device) (
rate(node_network_receive_errs_total{instance=~"$nodeList"}[5m])
)
max by (instance, device) (
rate(node_network_transmit_errs_total{instance=~"$nodeList"}[5m])
)
max:node_network_receive_errs_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_network_transmit_errs_total{instance:$nodeList} by {instance,device}.as_rate()
node.network.drops #
Receive and transmit packet drop rates per interface.max by (instance, device) (
rate(node_network_receive_drop_total{instance=~"$nodeList"}[5m])
)
max by (instance, device) (
rate(node_network_transmit_drop_total{instance=~"$nodeList"}[5m])
)
max:node_network_receive_drop_total{instance:$nodeList} by {instance,device}.as_rate()
max:node_network_transmit_drop_total{instance:$nodeList} by {instance,device}.as_rate()
node.network.rx.total #
Bytes per second received across every interface on the node.sum by (instance) (
max by (instance, device) (
rate(node_network_receive_bytes_total{instance=~"$nodeList"}[5m])
)
)
node.network.tx.total #
Bytes per second transmitted across every interface on the node.sum by (instance) (
max by (instance, device) (
rate(node_network_transmit_bytes_total{instance=~"$nodeList"}[5m])
)
)
node.conntrack.utilization #
Netfilter connection-tracking table occupancy as a fraction of its limit.max by (instance) (node_nf_conntrack_entries{instance=~"$nodeList"})
/
max by (instance) (node_nf_conntrack_entries_limit{instance=~"$nodeList"})
max:node_nf_conntrack_entries{instance:$nodeList} by {instance}
/ max:node_nf_conntrack_entries_limit{instance:$nodeList} by {instance}
node.uptime #
Seconds since the node booted.min by (instance) (
node_time_seconds{instance=~"$nodeList"} - node_boot_time_seconds{instance=~"$nodeList"}
)
min:node_time_seconds{instance:$nodeList} by {instance}
- max:node_boot_time_seconds{instance:$nodeList} by {instance}
node.collector.success #
Whether each enabled node-exporter collector returned data on the last scrape, per collector.min by (instance, collector) (
node_scrape_collector_success{instance=~"$nodeList"}
)
min:node_scrape_collector_success{instance:$nodeList} by {instance,collector}