Roadmap#

The goal of materialize-monitoring is first-class, composable observability for self-managed Materialize: logs, metrics, events and alerts. It is best-practices-by-default for customers who want a one-stop shop, and it does not force this stack on customers who already run their own. The same stack is consumed through Helm, through the Terraform module in this repository, and through the per-cloud Terraform wrappers downstream. Every component can be turned off in favour of one a customer already runs.

This page lists the work still to do, by project and milestone. What has shipped is summarized once, in What 1.0 ships, and itemized in the changelog. Linear holds the status of each item; this page holds the reasoning that does not fit in a ticket.

Projects#

TagProjectScope
ADOPTAdopt materialize-monitoring in CloudMaterialize Cloud’s own monitoring onto this chart, through Pulumi
CLOMMetrics for CloudMaterialize-side metrics: the tenant-scoped read path, the label vocabulary, and what Materialize emits
BYOCBYOC ObservabilityThe observability control plane, for BYOC and for self-managed call-home
BacklogObservability BacklogEverything else, with no committed date. Each milestone is a candidate project of its own

Two projects are closed. First-class Observability Infrastructure (FCO) built the platform and closed on August 31. Operational Observability (OO) hardened it, answered the operational questions, stamped 1.0, and closed on October 6.

Releases track a monthly cadence aligned to the 15th, and milestone dates are soft targets. A milestone is cited as its project tag and its Linear name, such as ADOPT · Loki cutover.

StatusMeaning
🔨In progress
⬜Planned
⛓️Blocked on upstream instrumentation; see Metrics contract

What 1.0 ships#

materialize-monitoring v1.0.0 stamps the chart and the Terraform module together, since the module’s Git tag is the chart version. From 1.0, a breaking change to the committed surface owes a deprecation cycle under the stability guarantees.

AreaAt 1.0
DeliveryThe umbrella chart, the optional CRDs chart and the dashboards chart, published to GHCR as OCI artifacts. The common Terraform module lives here and is wrapped per cloud in materialize-terraform-self-managed, where observability is on by default. Profiles cover sizing, scheduling, storage class, registries, Grafana ingress and persistence, and mTLS
CollectionAn Alloy agent for pod logs and the node journal, and an Alloy gateway for metrics. The gateway scrapes Materialize, kube-state-metrics, node-exporter, every kubelet’s cAdvisor, the AWS VPC CNI, Cilium and kube-proxy. It also pulls provider metrics from CloudWatch, Cloud Monitoring and Azure Monitor. Rust panics are merged into one entry at ingest
DestinationsThe bundled Loki and Thanos, any number of Prometheus remote-write destinations, OTLP, Datadog and Google Cloud’s Telemetry API, each filtered by importance tier
Dashboardsenv-top, env-logs, env-upgrade, env-persist and env-consensus for Materialize. infra-logs, infra-nodes, infra-net, infra-autoscaling, infra-karpenter, infra-cloud, infra-loki and infra-alloy for the platform. All are rendered from Rust against the query registry
AlertingThanos and Loki rulers notifying a highly available Alertmanager, with severity-to-receiver presets, capability-tagged rules, notification templates that link to Grafana, and a Terraform surface. Thirty-five alerts are in the default set, four of them log-derived, beside the ext:consensus_* recording rules. Alert and recording-rule names are a committed surface
SecurityA NetworkPolicy on every component, opt-in mTLS through cert-manager in three phases, and Trivy scanning of the rendered chart and the published images
QualificationTier-0 Terraform render checks, kind tiers 1 and 2, and a Rust E2E suite run on kind and on real cloud installs
ReleasePer-component SemVer streams, automated release PRs, a changelog carrying each PR’s release notes, a deprecation policy, and a committed-surface check

Four designs are fully shipped and now record why the stack is the way it is: Loki production infrastructure, Terraform modules, TLS authentication, and the stability and deprecation policy. The designs with work left are under Designs in flight.

Cloud adoption#

Tracked in Adopt materialize-monitoring in Cloud (ADOPT). Materialize Cloud moves its own monitoring onto this chart, on the same values.yaml surface self-managed customers consume. Cloud deploys through Pulumi, so a Pulumi component wraps the chart the way the Terraform module does (CLO-182). Cloud-only rules and configuration layer through values, with no privileged path (DEP-125). Cloud running the chart customers run is the strongest available test of the customer contract.

The migration is ordered by risk. Stateless components go first, because each can be replaced in place and rolled back. Each stateful component then gets a cutover of its own.

MilestoneWhat movesItems
Stateless servicesAlloy (agent and gateway), kube-state-metrics and node-exporter replace Cloud’s own. Cloud parity needs generic scrape discovery and the OTLP agent→gateway transport with a node-local WAL🔨 CLO-182 · ⬜ DEP-193 · ⬜ DEP-189 · ⬜ DEP-223
Thanos beside AMPThanos stands up beside AMP, and the gateway fans out to both as components move. AMP keeps the v2_mz_* families, so what reads it today is unaffected; Thanos takes canonical mz_* only. AMP scales down as its consumers move⬜ DEP-383
Loki cutoverCloud’s Loki moves onto the chart’s Loki on its existing storage backend, so no log data is copied. Log tenancy moves to tenantMap: byEnvironment, which the read path needs⬜ CLO-296
AlertingCloud’s rules run on the chart’s rulers and route through its Alertmanager to Cloud’s receivers. The Loki rules Cloud clicked into Grafana are already ported by shape in materialize-log-alerts.yaml⬜ DEP-125

Grafana is not migrated. Cloud’s Grafana lives elsewhere and stays there; it gains datasources for Thanos and Loki.

The Thanos milestone has two chart gaps to close first, both in DEP-383. The destination map filters by importance tier alone, so nothing can send v2_mz_* to AMP and keep it out of Thanos. materialize.deploymentMode: cloud also makes every rule read v2_mz_, which a Thanos holding only mz_* cannot answer. Separately, a destination below the all tier drops every recorded series and ALERTS until DEP-368 gives them a tier.

Metrics for Cloud#

Tracked in Metrics for Cloud (CLOM): what Materialize emits, what it is called, and how it is read.

Tenant-scoped read path#

Designed in A Tenant-Scoped Query API and tracked as DEP-269. Every consumer of this stack so far is a Grafana the chart deploys, reading backends over a ClusterIP Service. The read path serves the Materialize console and a customer’s own Grafana through a tenant-scoped proxy instead. The proxy is a Rust service in this workspace, tested against prom-label-proxy as an oracle.

ItemStatus
DEP-272 Per-family tenancy classification, generated from the query registry⬜
DEP-273 query-proxy chart component: JWT verification, issuer values, label enforcement, fail-closed defaults⬜
DEP-274 Separate grants for logs and metrics, and for the classes within each⬜
DEP-275 Exposure controls: the query frontend, per-tenant limits, NetworkPolicy narrowing, a public-exposure guard⬜
DEP-276 The proxy’s read-audit stream and a jti denylist⬜
DEP-277 A producer for the audit log class, which values declare and nothing reads⬜
DEP-278 A published query manifest that Console consumes by query ID⬜
DEP-279 A query-proxy profile, Terraform variables, and a minimum sizing envelope⬜
A customer-facing page: obtaining a token, the two Grafana datasource shapes, and what a tenant can read⬜

The read path is where this stack stops being optional. Console reads environment metrics from SQL today, which keeps a short window of history and is unavailable exactly when the environment is. Reading PromQL and LogQL instead makes a deployment without those endpoints a deployment with a broken Console. The design resolves that by mandating an interface rather than an implementation. The interface is a PromQL and a LogQL endpoint carrying the documented label contract, behind the proxy. Thanos and Loki are the bundled implementation, and a customer already running an equivalent points the proxy at theirs. This supersedes the earlier plan to expose a Prometheus endpoint for customers to scrape.

Metric label standardization#

One canonical label vocabulary for Materialize’s metrics, a reserved-label list, and an authoritative environment key. The key is environment_id from Materialize.spec.environmentId, with environment_name on mz_environment_info. Designed in A Label Vocabulary for Materialize Metrics, in review as #445. The work lands in two passes, ordered so that consumers never observe the second. This repository’s monitors first fill the canonical labels where they are empty, and Materialize then adopts the names upstream behind a label lint.

WhereItems, all ⬜
DesignDEP-346 review and acceptance
This repositoryDEP-349 labels.yaml and generated relabel rules · DEP-351 environment_id on every Materialize monitor · DEP-350 the GMP flavor drops target relabelings · DEP-354 dashboards, rules and pickers onto environment_id · DEP-355 the registry onto canonical names · DEP-356 a customer-facing label contract page · DEP-361 mz_environment_info through kube-state-metrics · DEP-360 retire the aliases after 30 days
UpstreamDEP-348 the orchestratord pod label and mz_environment_info · DEP-347 clusterd’s cluster="compute" · DEP-352 the label lint · DEP-353 /metrics/public replica labels · DEP-357 cluster_id · DEP-358 entity-qualified _info labels · DEP-359 the cluster and honeycomb const labels · DEP-362 shard, persist-name and version renames · DEP-363 an optional spec.environmentName · DEP-364 wallclock lag as a histogram · DEP-365 environment_id on reconciliation metrics

Metrics contract (upstream dependency)#

Tracked as DEP-207. The metric and label contract is the public API for everything here. Several dashboards depend on instrumentation that lives in MaterializeInc/materialize. About 39 signal families exist only through the SQL-on-scrape exporter path, which Cloud names v2_mz_*. v2_mz_ is not a newer variant; it names the exporter path, and the target is canonical mz_. environmentd has to emit those families natively before that path can retire. mz_object_info and the _info family are delivered, so name enrichment is unblocked everywhere.

AskUnblocksStatus
Native source and sink status. No sink status metric exists at allThe Sources and Sinks drilldowns⛓️
Native hydration and frontier / freshness signalsThe Hydration and Freshness drilldowns⛓️
Label-family harmonizationJoins across families. Now carried by Metric label standardization⬜
Latency histograms for consensus_scan, consensus_truncate, blob_delete and the timestamp oracle, which publish only a meanTail latency on env-persist and env-consensus⬜
balancerd and console metrics on self-managed. None reach Thanos todayDashboards for the two components users connect through⬜

balancerd does publish connection metrics, whose semantics CLO-282 questions. Its half of the last row may therefore be a scrape gap rather than an upstream one.

Materialize-side emission#

ItemStatus
CLO-55 Reconciliation timing, stalls, errors and successes for the k8s controllers, and an optional reconcile timeout⬜
CLO-299 The controller crate logs every failed reconcile at error level, transient Kubernetes errors included⬜
CLO-282 What mz_balancer_connection_status{status="success"} means for a refused connection⬜
CLO-256 Reduce less-essential environmentd metrics on very large environments⬜
CLO-255 Compress the metrics endpoint’s responses⬜
DEP-221 Move the scrapers into the materialize-operator chart, which should own them⬜

BYOC#

Tracked in BYOC Observability (BYOC). The observability control plane is where telemetry crosses from a customer-run cluster into Materialize’s. One channel serves two segments: BYOC environments, which Materialize operates, and self-managed installs, which opt in. Designed in Observability for Bring-Your-Own-Cloud.

Observability control plane#

ItemStatus
DEP-124 The dual-destination pipeline pattern: customer-local plus control plane⬜
DEP-220 Sanitize telemetry before control-plane egress⬜
DEP-236 A trust bundle for a private CA: the roots a corporate proxy or an on-premise object store needs, merged with the internal issuer’s⬜
DEP-287 oauth2.tls on the destination schema: certificate-authenticated token exchange, shared with call-home⬜
DEP-224 A Day 1 dependencies dashboard: are the Materialize and observability prerequisites satisfied⬜
BYOC · Observability implementationNot yet scoped

A reduced copy of telemetry crosses into the control plane, and the customer’s full-fidelity copy always stays with them. Metrics cross selected by importance tier, which already works per destination. Logs cross as an allowlisted, redacted, level-filtered subset, and log egress is opt-out at a stated cost in support quality. Redaction attaches to the destination rather than the pipeline, so the reduced copy is a fork of the customer’s stream and never a downgrade of it. The gateway pair enforces the boundary, rather than ad-hoc network configuration.

Secret filtering#

Designed in Secret Filtering and tracked as DEP-306. Credentials are the one exception to redaction at the destination. They are filtered at ingest, before anything is stored, because a credential in the customer’s own Loki is their exposure. Each control-plane branch runs a second, stricter filter immediately before its writer, and a canary per layer proves both are in the path.

ItemStatus
DEP-307 Typed loki.secretfilter and stage.luhn in the pipeline schema⬜
DEP-308 The ingest filter ahead of parsing, with the bulk tier ahead of the scan⬜
DEP-309 Close the two raw-line paths: the sampleDebug tap and live debugging⬜
DEP-310 Card-number redaction with stage.luhn⬜
DEP-311 Rule files: the Materialize rules, the name allowlist, and the egress file⬜
DEP-312 A corpus harness on the pinned Alloy binary, run on every Alloy bump⬜
DEP-313 Shadow mode, counting matches on real logs before any rule enforces⬜
DEP-314 The values surface and fail-open validators⬜
DEP-315 mz-monitoring-canary, a canary producer per layer⬜
DEP-316 Alerts, runbooks and a dashboard row⬜
DEP-317 A strict, fail-closed egress filter on each control-plane branch⬜
DEP-318 A backstop overlay for the control-plane gateway⬜
DEP-320 A customer-facing secret-filtering page, and the design’s revisions⬜

Call-home from self-managed#

Designed in Call-Home and tracked as DEP-270. BYOC environments are ones Materialize operates; self-managed installs are invisible to Materialize between escalations.

ItemStatus
DEP-280 The callHome.level consent ladder and profiles, defaulting to off, with a render assertion that the default sends nothing⬜
DEP-281 A heartbeat producer and three-part install identity⬜
DEP-282 The alerts level: alert state (ALERTS) over the gateway fan-out⬜
DEP-283 A previewOnly destination mode: the full chain, counted locally, never sent⬜
DEP-284 An egress meter, a per-level byte-rate ceiling, and a local export-failure alert⬜
DEP-285 A generated egress schedule per level⬜
DEP-286 The OTLP/HTTP default wire, forward-proxy and corporate-CA configuration⬜
DEP-288 The elevated window: time-boxed, self-expiring, never set remotely⬜
DEP-289 An Alertmanager webhook receiver path⬜
DEP-290 Terraform variables for the level, the credential reference and the proxy⬜
DEP-291 Control-plane per-tenant limits sized for install count, and a stated response model⬜
A customer-facing page: the levels and the schedule for each, previewing, the local meter, and retention, access and deletion⬜

The channel is the BYOC channel; the feature is consent. A self-managed customer bought software they run themselves, so every byte that leaves is a concession. The design is a bounded, monotone, locally visible ladder (off, heartbeat, alerts, metrics, diagnostics) over the existing pipeline. The cheapest useful level is alerts, which has something to forward now that a default rule set installs. A TLS-intercepting corporate proxy defeats mTLS outright, which makes token exchange (DEP-287) a reachability requirement in this segment. Receiving a signal creates an obligation. A stated response model, and a stated non-response, is a prerequisite for the alerts level rather than a follow-up.

Backlog#

Tracked in Observability Backlog (Backlog): work with no committed date. Each milestone below is a candidate project of its own, and moves out of the backlog when it is prioritized.

External dependencies#

Designed in Monitoring Materialize’s External Dependencies and tracked as DEP-233. The client’s measurement of a dependency is the SLI, and the dependency’s own telemetry is the diagnosis. Several pieces shipped with 1.0: the client vantage point, the provider pulls, and the client and provider dashboards. So did the persist-failures split and the ext:consensus_* recording rules. What remains is the flavor-specific half. It stays opt-in behind the normalized ext:* contract. That contract keeps seven database flavors and five object stores from multiplying the dashboard and alert set.

ItemStatus
DEP-295 The rest of the ext:* layer: ext:consensus_connections_used_ratio and the ext:objstore_* families🔨
DEP-297 The exporter vantage point: a multi-target postgres_exporter subchart, transaction-ID age, per-database sizes, a documented grant⬜
DEP-298 PostgreSQL, CNPG and self-hosted CockroachDB adapters, with the CockroachDB Cloud alerts reclassified as their own adapter⬜
DEP-299 An on-premise object-store adapter for MinIO, Garage and Ceph⬜
DEP-300 Version reporting (ext:*_version_info) across every adapter⬜
DEP-303 The externalDependencies values block and a dependency-monitoring profile⬜
DEP-366 The per-cloud wrappers declare the metadata database in externalDependencies.consensus⬜
DEP-368 A metric tier for recorded series, so a filtered destination keeps ext:* and ALERTS⬜
DEP-304 An AbortIncompleteMultipartUpload lifecycle rule on the persist bucket, in materialize-terraform-self-managed⬜

No adapter flavor is the one that waits. Every shipped wrapper provisions managed PostgreSQL, while CockroachDB and CNPG both run in customer clusters. The on-premise shape has no Terraform wrapper, but terraform/test/generic-cloud already runs it. Its adapters are therefore the ones this repository can test end to end.

Dashboards#

ItemStatus
DEP-208 Troubleshooting: symptom-first entry into the rest, with every panel linking onward⬜
DEP-157 Resizing (Day 2)⬜
DEP-158 Changing sources (Day 2)⬜
DEP-159 Changing external destinations (Day 2)⬜
DEP-160 Managing users (Day 2)⬜
DEP-212 Hydration drilldown⛓️
DEP-213 Freshness drilldown⛓️
DEP-214 Sources drilldown⛓️
DEP-215 Sinks drilldown⛓️
DEP-225 Sizing (Day 1)⬜
DEP-378 infra-pods: one workload’s health, logs and metrics⬜
DEP-379 Meta-monitoring for Grafana, Thanos and Alertmanager⬜
DEP-384 infra-net’s Cloud Networking rows, from the provider pull⬜
DEP-115 A Datadog dashboard set⬜
DEP-217 A Google Cloud Monitoring dashboard set⬜
DEP-218 A Honeycomb dashboard set⬜
DEP-325 infra-* variables that reference variables those dashboards never define⬜
DEP-333 infra-nodes’ node picker resolves to nothing on EKS⬜

Day 2 operations outweigh Day 1. Upgrades, resizing, changing sources and destinations, and managing users matter most to a running deployment, so the Day 1 dashboards come last. The change-operation dashboards focus on objects being added, removed and initially populated, with error detection, rather than on steady state. Troubleshooting is symptom-first where env-top is subsystem-first, so it sequences after the dashboards it links into. The drilldowns wait on the metrics contract, and Freshness is the one customers ask about most.

Alerting#

Designed in Alerting in Self-Managed. The evaluation and notification path shipped with 1.0, and what remains is breadth and operability.

ItemStatus
DEP-372 A runbook for every shipped alert, under operating/runbooks/⬜
DEP-373 A deadman’s switch and alerting meta-alerts⬜
DEP-374 Triage of the ported rules that install only when selected⬜
DEP-335 Workloads declare their alert tier through labels, so freshness can join the default set⬜
DEP-375 Inhibit rollout-sensitive alerts while a rollout is in progress⬜
DEP-376 alerting.alertmanager.mode: external⬜
DEP-377 rules.extra⬜
DEP-339 Alerting through Terraform: the downstream documentation🔨
CLO-266 An alert on 5xx responses from balancerd⬜
DEP-340 E2E: the Loki ruler loads the chart’s log rules, and each fires on a real line⬜
DEP-345 The Thanos ruler’s rule importer never picks up a changed selector⬜
DEP-328 Report the Thanos ruler’s write_relabel_configs panic upstream, then drop the gateway workaround⬜

Freshness and source rules exist but stay opt-in. Calibrated against a production fleet, an absolute threshold fires mostly on clusters that are behind by design. Only a declared tier can tell those apart. Sink health and correctness logs beyond the three Materialize log-derived defaults have no rule yet.

Collection#

ItemStatus
DEP-385 A gateway over its memory limiter’s soft limit refuses its scrapes for hours: sizing, limiter placement, resharding🔨
DEP-380 CNI monitors for Calico and Azure Network Policy Manager⬜
DEP-126 Trace correlation as structured metadata in the log pipeline, without trace storage⬜
CLO-76 Optional continuous profiling, off by default⬜
DEP-252 Cost visibility through an optional OpenCost component, with its design in review as #365🔨

Platform#

ItemStatus
DEP-338 An image vulnerability gate that fails only the change that can act on it⬜
DEP-382 security-gate as a required check on main⬜
DEP-249 Trivy suppressions are per check, so they also mask our own templates⬜
DEP-251 SARIF findings point at temporary render paths⬜
DEP-381 tflint in Terraform CI⬜
DEP-369 The per-cloud wrappers pass the Kubernetes cluster name as cluster_name⬜
DEP-370 Renovate bumps the mzmon-alloy tag and leaves its digest⬜
DEP-371 Release PRs leave Cargo.lock stale⬜
DEP-111 and DEP-118 An ArgoCD and FluxCD CI matrix⬜

Designs in flight#

A design moves to Shipped when the rows it owns on this page are gone.

DesignStatusRemaining work
Observability for Bring-Your-Own-CloudReadyBYOC
A Tenant-Scoped Query APIReadyTenant-scoped read path
Alerting in Self-ManagedReadyAlerting
Call-HomeReadyCall-home from self-managed
Monitoring Materialize’s External DependenciesReadyExternal dependencies
Secret FilteringReadySecret filtering
A Label Vocabulary for Materialize MetricsDraft, in review as #445Metric label standardization
Cost visibility with OpenCostDraft, in review as #365Collection