Metrics Architecture#

This page describes how metrics move through materialize-monitoring. In the bundled stack, every metric enters through the alloy-gateway, which scrapes or receives it, processes it, and writes it to one or more destinations. The default destination is the bundled Thanos. Thanos keeps recent data in Receive, keeps everything older as blocks in object storage, and answers PromQL for Grafana and for the Thanos Ruler.

The other pages in this section cover each stage in depth:

PageCovers
CollectingThe five ways metrics reach the gateway, and its ingress ports
ScrapingMonitor discovery, the kubelet scrape, and scrape configurations for a separately run Prometheus
StoringDestinations, importance tiers, the denylist, object storage, retention and downsampling
QueryingPromQL through Grafana and the Thanos Query API
Alerting > ConfiguringThe Thanos Ruler, and how PrometheusRule resources reach it

The production checklist for this backend is Production Best Practices > Metrics (Thanos). The gateway’s component-level pipeline is in the metrics pipeline reference (internal).

End-to-end dataflow#

flowchart TB
  src["Sources<br/>monitor targets · kubelets<br/>cloud APIs · pushers"]
  graf["Grafana"]
  ruler["Thanos Ruler"]
  gw["alloy-gateway"]
  query["Thanos Query"]
  am["Alertmanager"]
  other["Other destinations<br/>AMP · OTLP backends"]
  recv["Thanos Receive<br/>RF 3"]
  sg["Thanos Store Gateway"]
  bucket[("Object storage")]
  comp["Thanos Compactor"]

  src --> gw
  ruler -->|"rule results"| gw
  ruler -->|"PromQL"| query
  ruler -->|"alerts"| am
  graf -->|"PromQL"| query
  gw -.->|"optional"| other
  gw -->|"remote-write"| recv
  query -->|"recent"| recv
  query -->|"historical"| sg
  recv -->|"2h blocks"| bucket
  sg -->|"read blocks"| bucket
  bucket <-->|"rewrite blocks"| comp

Writes flow downward. The gateway collects from every source, runs the shared processing pipeline, and writes each destination its own filtered copy. Thanos Receive replicates each write to three pods, serves the recent window, and uploads a block to object storage every two hours. The Compactor merges, deduplicates, downsamples, and expires those blocks in place.

Reads flow upward from Grafana into Thanos Query, which merges recent data from Receive with historical data from the Store Gateway. The Thanos Ruler queries the same path, sends alerts to Alertmanager, and writes its results back into the gateway, so rule output takes the same route as every other metric.

Collection#

All metric collection runs in the alloy-gateway Deployment. The alloy-agent DaemonSet collects logs only. Container metrics come from the kubelet, which the gateway scrapes directly, so no per-node metrics agent is needed.

SourceReaches the gateway byDefault
PodMonitor and ServiceMonitor targetsScrape, discovered from every monitor in the clusterOn. The chart ships monitors for Materialize, kube-state-metrics, node-exporter and each stack component
Kubelet /metrics/cadvisorScrape of every node’s kubeletOn, every 60s
kube-proxyScrape, discovered by a server-side pod selectorOn where kube-proxy runs
Cloud provider monitoring APIsPull through an in-process exporterOff
Prometheus remote-writePush to :9090Listener on
OTLPPush to :4317 (gRPC) or :4318 (HTTP)Listener on
Thanos Ruler and Loki RulerRemote-write push to :9090On

The gateway replicas form an Alloy cluster. Every scrape and pull is clustered, so each target has one owner among the replicas. Adding a replica spreads scrape load without duplicating samples. A cloud provider’s API is therefore called once per interval, however many replicas run.

See more: Collecting, Scraping, and Cloud Provider Metrics.

Inside the gateway#

flowchart TB
  scrape["Scrapes and pulls<br/>monitors · kubelet · kube-proxy · providers"]
  rw["prometheus.receive_http :9090<br/>fills job / instance on pushed samples"]
  otlpin["otelcol.receiver.otlp<br/>:4317 / :4318"]
  bridge["inputBridge<br/>Prometheus → OTLP"]

  subgraph proc["Shared processing"]
    transform["inputMetricProcessor"]
    deny["inputMetricDeny<br/>denyMetrics"]
    limit["memory_limiter<br/>refuses at 85%"]
    batch["batch"]
    transform --> deny --> limit --> batch
  end

  egress{"egress seam"}

  subgraph promdest["Per remote-write destination"]
    out["outputBridge<br/>OTLP → Prometheus"]
    tierp["tier filter<br/>minMetricImportance"]
    wal["prometheus.remote_write<br/>own WAL · cluster external label"]
    out --> tierp --> wal
  end

  subgraph oteldest["Per OTLP exporter"]
    tiero["tier filter<br/>minMetricImportance"]
    exp["otlp · googlecloud · datadog exporter"]
    tiero --> exp
  end

  scrape --> bridge
  rw --> bridge
  bridge --> transform
  otlpin --> transform
  batch --> egress
  egress --> out
  egress -.->|"when enabled"| tiero

Prometheus input is bridged into OTLP at the start, processed by otelcol.processor.* components, and bridged back for remote-write. Every scrape keeps the target’s metric type, so a counter or histogram arrives typed at an OTLP destination. The memory limiter refuses new data at 75% of its limit, and the refusal propagates back to the receivers as backpressure.

The stream forks after the shared stages. Each destination has its own importance-tier filter upstream of its writer. A remote-write destination’s write-ahead log therefore holds only the series that destination accepts, and a stuck destination backs up its own WAL and no other. Each remote-write destination also stamps a cluster external label.

See more: Storing > The remote-write destinations, Controlling what each destination stores, and the metrics pipeline reference (internal).

Thanos write path#

Thanos Receive#

flowchart TB
  gw["alloy-gateway<br/>thanos destination"]
  svc["thanos-receive Service :10908"]
  router["any Receive pod<br/>routes the request"]
  ring(("ketama hashring<br/>3 owners per series"))
  own0["receive-0<br/>head + WAL"]
  own1["receive-1<br/>head + WAL"]
  own2["receive-2<br/>head + WAL"]
  bucket[("Object storage<br/>one 2h block per pod<br/>receive_replica=&lt;pod&gt;")]

  gw -->|"remote-write"| svc --> router --> ring
  ring --> own0
  ring -->|"gRPC :10901"| own1
  ring -->|"gRPC :10901"| own2
  own0 --> bucket
  own1 --> bucket
  own2 --> bucket

Receive runs in standalone mode, where every pod both routes and ingests. The pod that takes a request hashes each series onto an auto-generated ketama hashring and forwards it to the pods that own it. With replication factor 3 a write succeeds once two of the three owners accept it.

At the default 3 replicas every pod owns every series. The thanos-large profile runs 6, and each series then lives on 3 of the 6.

Each pod keeps its head and WAL on an emptyDir, cuts a block every two hours, and uploads it to object storage. Local retention is 6h, which serves the recent window to queries until the Store Gateway has the uploaded blocks. Every pod uploads its own copy of what it holds, labelled with its own receive_replica. The Compactor later merges those copies into one.

See more: Production Best Practices > Receive for the quorum table and zone spread, and Storage: ephemeral by default for why Receive has no volume.

Object storage and the block lifecycle#

flowchart TB
  b0["2h block<br/>receive-0"]
  b1["2h block<br/>receive-1"]
  b2["2h block<br/>receive-2"]
  merge["deduplicate on receive_replica<br/>vertical compaction"]
  levels["compact 8h → 2d → 14d"]
  ds["downsample<br/>5m from ≥ 40h · 1h from ≥ 10d"]
  ret["retention<br/>raw 30d · 5m 90d · 1h 365d"]
  del["marked for deletion<br/>removed after 48h"]

  b0 --> merge
  b1 --> merge
  b2 --> merge
  merge --> levels --> ds
  levels -->|"replaced sources"| del
  ds --> ret --> del

Object storage is the only durable copy of metric data. Every other Thanos component either writes blocks into it or reads them out.

The Compactor is a singleton, because two compactors against one bucket corrupt it. It works only on blocks older than its 30m consistency delay, so it never reads a block whose upload is still in progress.

StageWhat happens
DeduplicateBlocks are grouped by their external labels with receive_replica removed, so each replica’s copy of a window lands in one group. Vertical compaction merges them into a single block and drops identical samples
CompactBlocks are merged into longer ranges, 2h into 8h, 8h into 2d, and 2d into 14d
DownsampleA raw block spanning 40h or more gets a 5m copy. A 5m block spanning 10d or more gets a 1h copy
RetainEach resolution expires on its own horizon, by default raw 30d, 5m 90d and 1h 365d
DeleteA block that a compaction replaced, or that retention expired, is marked for deletion and removed 48h later. The delay lets the Store Gateway load the replacement before the source disappears

Vertical compaction also absorbs the few seconds of overlap Receive writes when a pod restarts. Without it, the first overlap halts the Compactor, and a halted Compactor runs none of the stages above.

The Store Gateway reads blocks directly from the bucket and keeps an index-header cache on its own volume.

See more: Storing > Retention and downsampling, Production Best Practices > Retention & compaction, and ThanosCompactHalted.

Read path#

flowchart TB
  graf["Grafana<br/>Thanos datasource"]
  qfe["Thanos Query Frontend<br/>splitting · result cache"]
  query["Thanos Query<br/>deduplicates on<br/>receive_replica · ruler_replica · prometheus_replica"]
  receive["Thanos Receive<br/>head + last 6h of blocks"]
  store["Thanos Store Gateway<br/>uploaded blocks"]
  ruler["Thanos Ruler"]
  bucket[("Object storage")]

  graf -->|"PromQL :9090"| query
  graf -.->|"thanos-large: PromQL"| qfe
  qfe -.-> query
  query -->|"StoreAPI gRPC"| receive
  query -->|"StoreAPI gRPC"| store
  query -->|"rules · alerts gRPC"| ruler
  store --> bucket
  ruler -->|"PromQL :9090"| query

Thanos Query presents a Prometheus-compatible API and holds no data of its own. It discovers Receive, the Store Gateway and the Ruler through DNS SRV lookups, and fans each query out over gRPC.

A query that spans recent and historical data reads both, and Query merges the overlap. It then deduplicates on the replica labels. Each Receive replica’s copy of a series, each Ruler replica’s copy of a rule result, and a deduplicated block from the bucket all collapse to one series.

Grafana’s Thanos datasource points at Thanos Query. The thanos-large profile enables the Query Frontend and moves the datasource to it, while the Ruler stays on Thanos Query, so rules never evaluate against a cached result.

See more: Querying and Alerting > Configuring.

Rule evaluation#

flowchart TB
  pr["PrometheusRule resources<br/>flavor ≠ logql"]
  sidecar["import sidecar<br/>lists every 60s · reloads"]
  ruler["Thanos Ruler<br/>2 replicas · stateless"]
  query["Thanos Query"]
  am["Alertmanager<br/>every replica"]
  gw["alloy-gateway :9090"]
  receive["Thanos Receive"]

  pr --> sidecar --> ruler
  ruler -->|"PromQL"| query
  ruler -->|"alerts"| am
  ruler -->|"ALERTS · recording results"| gw
  gw --> receive

The Thanos Ruler is stateless. It evaluates by querying Thanos Query, keeps no TSDB of its own, and remote-writes every result to the gateway. Rule output therefore reaches every destination the gateway writes to, and not only the bundled Thanos.

Both replicas evaluate every rule. Each stamps its own ruler_replica, Thanos Query deduplicates on that label, and Alertmanager deduplicates the notifications. The gateway fills in job="thanos-ruler" and a constant instance on the Ruler’s samples. The two replicas’ series then differ only in ruler_replica, which is the label Query deduplicates on.

The Loki Ruler’s recording rules arrive the same way, as remote-write to the gateway.

See more: Alerting > Configuring for rule delivery and stateless mode, and Alert Architecture for what Alertmanager does next.

Destinations beyond the bundled Thanos#

The gateway writes to any number of destinations at once, each filtered to its own importance tier. The bundled Thanos is the default entry in that map, and it can be removed.

ProfileShape
aws-amp-fanoutFull fidelity in the bundled Thanos, and only essential series in Amazon Managed Prometheus
otlp-metrics-honeycombGateway metrics to a generic OTLP backend
otel-metrics-fanoutSeveral OTLP backends at once, each on its own tier

A deployment that already runs its own Prometheus can skip the gateway entirely and scrape Materialize with the configurations on Scraping.

See more: Storing > Other Metric Storage Backends.

See more#