Recorded Series#

These are the series the chart’s recording rules write, generated from the query registry. A recorded series is a metric name this repository mints, so it reads the same whatever produced the measurement underneath it. The Thanos ruler evaluates the rules and writes the results back through the alloy-gateway, so a recorded series is queryable wherever the gateway’s other metrics are.

Every ext:* series is part of the normalized layer for an external dependency. One name is recorded by each adapter that can measure it, and the flavor label names the adapter.

LabelOnHolds
flavorEvery ext:* seriesThe adapter that recorded it: persist is Materialize’s own measurement, and rds, cloudsql and azure-postgres are a cloud provider’s
namespaceSeries from persistThe Materialize environment’s namespace
resourceSeries from a cloud providerThe database’s name at the provider, as declared in externalDependencies.consensus

An adapter that cannot measure a series records nothing for it, rather than a zero. The provider adapters record only the databases externalDependencies.consensus declares; see Configuring Alerting.

Recorded-series names are covered by the stability policy from the release that first ships them, like alert names.

ext:consensus_up #

flavor="persist" from ext-consensus, group ext_consensus_persist

1 while Materialize’s calls to the metadata database succeed, per environment namespace, and 0 when none has succeeded in five minutes. An environment makes about 80 calls a second while idle, so five minutes without a success is a database it cannot reach rather than one it had no reason to call. A call that hangs counts as not succeeding. Absent when the environment is not being scraped.

Recorded where: materialize.

sum by (namespace) (
  rate(mz_persist_external_succeeded_count{op=~"consensus_.*"}[5m])
) > bool 0

flavor="cloudsql" from ext-consensus, group ext_consensus_cloudsql

1 while Cloud SQL reports each metadata database instance up, and 0 when it reports it down.

Recorded where: cloud-monitoring.

max by (resource) (
  label_replace(
    last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_up{database_id=~"(?i).+:(${consensusCloudsqlResources})"}[15m]),
    "resource", "$1", "database_id", ".+:(.+)"
  )
)

flavor="azure-postgres" from ext-consensus, group ext_consensus_azure_postgres

1 while Azure reports each metadata database flexible server alive, and 0 when it reports it down.

Recorded where: azure-monitor.

max by (resource) (
  label_replace(
    last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_is_db_alive_minimum_count{resourceName=~"(?i)${consensusAzurePostgresResources}"}[15m]),
    "resource", "$1", "resourceName", "(.+)"
  )
)

ext:consensus_commit_latency_seconds:p99 #

flavor="persist" from ext-consensus, group ext_consensus_persist

The 99th-percentile time a commit to the metadata database takes, over five minutes, per environment namespace. A commit is persist’s compare-and-set, op="consensus_cas", the only consensus operation with a latency histogram. Measured by the client, so it includes the network and any wait for a pooled connection. Absent when there was no commit in the window.

Recorded where: materialize.

histogram_quantile(0.99,
  sum by (namespace, le) (
    rate(mz_persist_external_op_latency_bucket{op="consensus_cas"}[5m])
  )
) >= 0

ext:consensus_xid_used_ratio #

flavor="rds" from ext-consensus, group ext_consensus_rds

The fraction of PostgreSQL’s transaction-ID space each RDS metadata database has used, against the 2^31 at which it stops accepting writes. From CloudWatch’s MaximumUsedTransactionIDs.

Recorded where: cloudwatch.

max by (resource) (
  label_replace(
    last_over_time(aws_rds_maximum_used_transaction_ids_maximum{dimension_DBInstanceIdentifier=~"(?i)${consensusRdsResources}"}[15m]),
    "resource", "$1", "dimension_DBInstanceIdentifier", "(.+)"
  )
) / 2147483648

flavor="cloudsql" from ext-consensus, group ext_consensus_cloudsql

The fraction of PostgreSQL’s transaction-ID space each Cloud SQL metadata database has used. Cloud SQL publishes this as a ratio of the same 2^31 limit the other adapters divide by.

Recorded where: cloud-monitoring.

max by (resource) (
  label_replace(
    last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_postgresql_transaction_id_utilization{database_id=~"(?i).+:(${consensusCloudsqlResources})"}[15m]),
    "resource", "$1", "database_id", ".+:(.+)"
  )
)

flavor="azure-postgres" from ext-consensus, group ext_consensus_azure_postgres

The fraction of PostgreSQL’s transaction-ID space each metadata database flexible server has used, against the 2^31 at which it stops accepting writes.

Recorded where: azure-monitor.

max by (resource) (
  label_replace(
    last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_maximum_used_transactionids_maximum_count{resourceName=~"(?i)${consensusAzurePostgresResources}"}[15m]),
    "resource", "$1", "resourceName", "(.+)"
  )
) / 2147483648

ext:consensus_storage_used_ratio #

flavor="cloudsql" from ext-consensus, group ext_consensus_cloudsql

The fraction of each Cloud SQL metadata database’s provisioned disk in use.

Recorded where: cloud-monitoring.

max by (resource) (
  label_replace(
    last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_disk_utilization{database_id=~"(?i).+:(${consensusCloudsqlResources})"}[15m]),
    "resource", "$1", "database_id", ".+:(.+)"
  )
)

flavor="azure-postgres" from ext-consensus, group ext_consensus_azure_postgres

The fraction of each metadata database flexible server’s provisioned storage in use.

Recorded where: azure-monitor.

max by (resource) (
  label_replace(
    last_over_time(azure_microsoft_dbforpostgresql_flexibleservers_storage_percent_maximum_percent{resourceName=~"(?i)${consensusAzurePostgresResources}"}[15m]),
    "resource", "$1", "resourceName", "(.+)"
  )
) / 100