Cloud Provider Metrics#

The gateway can pull what a cloud provider’s monitoring API publishes about the metadata database and the buckets a Materialize deployment depends on, and about whether the provider can supply the nodes its cluster asks for. The result is written beside every other metric, with the same retention, the same PromQL, and the same destinations. A provider series is therefore joinable with mz_persist_* in a single expression. A Grafana CloudWatch, Cloud Monitoring or Azure Monitor datasource cannot offer that.

The Infrastructure Cloud Provider dashboard (infra-cloud) draws the database and bucket pulls, and Infrastructure Autoscaling (infra-autoscaling) the node and quota pulls on its Cloud Capacity tab, each with rows for whichever provider it finds; see Available Dashboards.

Provider collection is off by default and adds to what the clients already report about the same dependencies. Persist, Loki and Thanos measure every request they make against the database and the bucket, at full resolution and at no cost. The provider adds what no client can see: CPU, memory and storage headroom, the burst credits that throttle a volume, transaction-ID consumption, and bucket growth. For the cluster’s nodes it adds the provider’s own view: host and instance health, node groups that cannot reach their desired size, compute quota, and on AKS the managed cluster autoscaler, which runs where nothing in the cluster can scrape it. The external-dependency design records the reasoning.

What is pulled#

ProviderServiceDefault metricsResolution
cloudwatchRDS instanceCPUUtilization, CPUCreditBalance, FreeableMemory, FreeStorageSpace, DatabaseConnections, ReadLatency, WriteLatency, DiskQueueDepth, BurstBalance, EBSIOBalance%, EBSByteBalance%, MaximumUsedTransactionIDsOne minute, summarised per five-minute period
cloudwatchS3 bucketBucketSizeBytes per storage class, NumberOfObjectsDaily
gcpCloud SQL instancecpu/utilization, memory/utilization, disk/utilization, postgresql/num_backends, postgresql/transaction_id_utilization, upOne minute
gcpGCS bucketstorage/v2/total_bytes, storage/v2/total_count, split into live, noncurrent and soft-deleted objectsDaily, repeated every five minutes
azurePostgreSQL Flexible Servercpu_percent, cpu_credits_remaining, memory_percent, storage_percent, active_connections, connections_failed, disk_queue_depth, disk_iops_consumed_percentage, disk_bandwidth_consumed_percentage, maximum_used_transactionIDs, is_db_aliveOne minute, summarised per five-minute bucket
azureBlob Storage accountBlobCapacity, BlobCountHourly, refreshed about once a day
azureBlob Storage accountAvailability, SuccessServerLatency, SuccessE2ELatencyOne minute, summarised per five-minute bucket
cloudwatchEKS cluster’s nodesStatusCheckFailed_System, StatusCheckFailed_Instance, StatusCheckFailed_AttachedEBS per nodeOne minute, summarised per five-minute period
cloudwatchEKS cluster’s managed node groupsGroupDesiredCapacity, GroupInServiceInstances, GroupPendingInstances, GroupMaxSizeOne minute, summarised per five-minute period
cloudwatchThe region, when any EKS cluster is listedOn-Demand vCPUs in use in the standard families, account-wide (AWS/Usage ResourceCount)One minute, summarised per five-minute period
gcpCompute Engine region and its zonesquota/cpus_per_vm_family and quota/local_ssd_total_storage_per_vm_family: usage, limit and refusalsA few times a day; refusals as they happen
azureAKS clustercluster_autoscaler_unschedulable_pods_count, cluster_autoscaler_cluster_safe_to_autoscale, cluster_autoscaler_scale_down_in_cooldown, cluster_autoscaler_unneeded_nodes_countOne minute, summarised per five-minute bucket
azureAKS cluster’s node scale setsVmAvailabilityMetric per node VMOne minute, summarised per five-minute bucket

The metric sets are fixed, and values only name the resources. Each pull is a custom component in the gateway’s gateway-provider pipeline, so changing what is pulled is a change to that pipeline. CloudWatch takes one instance of it per listed resource; GCP and Azure take one per service, whose filter names every listed resource. The Cloud SQL entries are metric-type prefixes, so postgresql/num_backends also pulls num_backends_by_state and num_backends_by_application, and up also pulls uptime.

Azure applies one aggregation list to every metric in a call. The Flexible Server pull therefore asks for all four and keeps the one or two each metric needs. Those are the statistics the RDS pull asks CloudWatch for: average and maximum CPU, minimum credits, maximum connections, and so on. That is 11 series per server, and 12 on a Burstable tier, which also publishes cpu_credits_remaining.

The nodes and the quota behind them#

A node that cannot be had looks the same from inside the cluster whatever the reason: pods stay pending, and the autoscaler reports a failed launch. These pulls say which reason it is, as far as the provider publishes one.

QuestionEKSGKEAKS
Is the host under a node failingStatusCheckFailed_System—VmAvailabilityMetric, with Context saying whether the platform took it down
Is a node group short of its desired sizeDesired against in-service, for managed node groups—Unschedulable pods while the autoscaler is allowed to scale
Is quota the limitvCPUs in use, account-wide; the limit is not in CloudWatchCPUs and local SSD per family, against the limit, and each refusal—

EKS nodes are the one resource found by tag rather than named, because Karpenter and the node groups replace them. EKS and Karpenter tag every node they launch with aws:eks:cluster-name, and each managed node group’s Auto Scaling group with eks:cluster-name, and the pull matches those tags to the listed cluster names exactly. Karpenter’s node pools have no group of their own, so the group metrics cover only managed node groups; Karpenter’s own metrics cover the rest. The vCPU usage counts every On-Demand instance in the account and region, not only the cluster’s, because that is what the quota counts. Its limit is the Service Quotas console’s L-1216C47A.

The Compute Engine quotas are per machine family: C4 and C4A each have their own, and the regional CPUS quota covers only older families. A family appears only once something in the project uses it, so a project running only C4 and C4A has a handful of series per region. Zonal limits read 2^63 − 1 where no zonal quota applies. A region is matched with its zones and nothing else, so europe-west1 does not also pull europe-west10.

AKS runs its cluster autoscaler inside the managed control plane. Its metrics reach no scrape, and Azure Monitor publishes these four instead. Each keeps the aggregation that shows the worst of its five minutes: the most pods waiting and the most unneeded nodes, any cooldown, and any moment it was unsafe to scale. The node pools are scale sets in the cluster’s node resource group, which the pull finds through the cluster.

Metric names#

ProviderName shapeExampleResource label
cloudwatchaws_<service>_<metric>_<statistic>aws_rds_free_storage_space_minimumdimension_DBInstanceIdentifier, dimension_BucketName
gcpstackdriver_<resource type>_<metric type>stackdriver_cloudsql_database_cloudsql_googleapis_com_database_cpu_utilizationdatabase_id (project:instance), bucket_name
azureazure_<resource type>_<metric>_<aggregation>_<unit>, lowercasedazure_microsoft_dbforpostgresql_flexibleservers_storage_percent_maximum_percentresourceName, and resourceID lowercased

The node and quota pulls follow the same shapes.

ProviderFamiliesResource labels
cloudwatchaws_ec2_status_check_failed_{system,instance,attached_ebs}_maximumdimension_InstanceId, plus tag_karpenter_sh_nodepool or tag_eks_nodegroup_name
cloudwatchaws_autoscaling_group_{desired_capacity,in_service_instances,pending_instances,max_size}_averagedimension_AutoScalingGroupName, tag_eks_nodegroup_name
cloudwatchaws_usage_resource_count_maximumdimension_Class="Standard/OnDemand", region
gcpstackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_{cpus,local_ssd_total_storage}_per_vm_family_{usage,limit,exceeded}location (region or zone), vm_family, limit_name
azureazure_microsoft_containerservice_managedclusters_cluster_autoscaler_*resourceName
azureazure_microsoft_compute_virtualmachinescalesets_vmavailabilitymetric_minimum_countdimensionVmname (<scale set>_<index>), dimensionContext, resourceName (the scale set)

A node joins to kube_node_info on its provider ID: the instance ID on EKS, and the scale set and index on AKS. The discovery pull also writes an aws_<service>_info series carrying every tag on the resource, which the pull drops, since tags are free text; the two tags it keeps are set by EKS and Karpenter.

CloudWatch series also carry region, account_id and account_alias, and name, which is rds or s3 on every series. Join on the dimension_* label.

Azure series also carry resourceGroup, subscriptionID and subscriptionName. They also carry the interval and timespan each pull reads, as ISO 8601 durations. The Blob Storage families are named for the blob service, as azure_microsoft_storage_storageaccounts_blobservices_*. No resource tag is copied onto a series; the exporter copies owner by default, and the pull turns that off.

Every provider series carries job="integrations/cloudwatch", job="integrations/gcp" or job="integrations/azure". instance is the resource on RDS and S3 series, since each resource has its own pull, and eks on the EKS pull, which covers every listed cluster. On Cloud Monitoring series it is cloudsql, gcs, compute_quota or compute_quota_exceeded, since one pull covers every listed resource of a service. Azure series are the same, with postgres, blob_capacity, blob_requests, aks_autoscaler or aks_nodes.

Configuring it#

Each provider lists the resources to watch. Nothing else is pulled, and a provider that is enabled with no resources fails at render.

pipeline:
  metrics:
    provider:
      cloudwatch:
        enabled: true
        region: us-east-1
        rds:
          instances: [mz-prod-db]
        s3:
          buckets: [mz-prod-storage-a1b2]
        eks:
          clusters: [mz-prod-eks]
      gcp:
        enabled: true
        projectId: my-project
        cloudSql:
          instances: [mz-prod-pg]
        gcs:
          buckets: [mz-prod-storage]
        compute:
          regions: [us-east1]
      azure:
        enabled: true
        subscriptionId: 00000000-0000-0000-0000-000000000000
        postgres:
          servers: [mz-prod-pg]
        blob:
          storageAccounts: [mzprodstorage]
        aks:
          clusters: [mz-prod-aks]

A Cloud SQL entry is the instance name, not the project:region:instance connection name. A Flexible Server entry is the server name, not its FQDN, and a storage account entry is the account name, not its endpoint. Both kinds of Azure name are unique across Azure, so no resource group is needed. An EKS or AKS entry is the cluster’s name, not its ARN or resource ID. AKS names are unique only within a resource group, so two clusters of one name in the subscription are both pulled, and told apart by resourceGroup. A Compute Engine entry is a region, such as us-east1; its zones are included, and a zone is refused. The chart refuses anything else at render, since the names sit inside the Resource Graph query’s filter. subscriptionId is the subscription’s GUID. On Azure Government or Azure China, set cloudEnvironment to azureusgovernmentcloud or azurechinacloud.

The chart refuses, at render, the values that would otherwise stop the gateway starting: a scrape timeout longer than the interval, and two buckets whose names differ only in . and -, which would become the same component label. Alloy accepts both in alloy validate and exits on them at load, which would stop logs and metrics along with the pull.

The full set of keys is in the values reference, under pipeline.metrics.provider.

Identity and grants#

Credentials never travel through values. The pull runs as the gateway pod’s own cloud identity, bound through alloy-gateway.serviceAccount.annotations.

ProviderIdentityGrant
cloudwatchIRSA (eks.amazonaws.com/role-arn), EKS Pod Identity, or static keys as AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY in the mzmon-alloy-gateway-env Secretcloudwatch:GetMetricStatistics. iam:ListAccountAliases fills account_alias; without it every pull logs a warning. EKS clusters also need cloudwatch:GetMetricData, cloudwatch:ListMetrics, tag:GetResources and autoscaling:DescribeAutoScalingGroups
gcpWorkload Identity (iam.gke.io/gcp-service-account)roles/monitoring.viewer on the project, which also reads its quota
azureWorkload identity (azure.workload.identity/client-id, plus a pod label), or a service principal as AZURE_CLIENT_ID, AZURE_TENANT_ID and AZURE_CLIENT_SECRET in the mzmon-alloy-gateway-env SecretMonitoring Reader on each named server and storage account, and on each named AKS cluster and its node resource group

Azure workload identity needs two things on the gateway, not one. The annotation names the identity. The Entra webhook injects it only into pods labelled azure.workload.identity/use: "true", set through alloy-gateway.controller.podLabels. The chart’s Terraform module sets the label whenever the gateway’s annotations include azure.workload.identity/client-id. On Azure the grant can be scoped to each resource, unlike on the other two clouds. Resource Graph returns only the resources the identity can read, so the grant is also what decides which named resources are pulled.

On GCP, the service account the Terraform module creates for the Google Cloud Metrics exporter holds roles/monitoring.metricWriter, which writes metrics and cannot read them. Reading needs roles/monitoring.viewer added to the same account.

The chart warns at render when a provider is enabled and the gateway’s service account carries no matching annotation. On Azure it also warns when the annotation has no pod label beside it. It cannot tell whether EKS Pod Identity, a static key, or a direct Workload Identity principal is in use instead, so the warning is advisory.

The providers fail differently without a credential. CloudWatch resolves its credential on each pull, so a missing one produces empty pulls and the gateway keeps running. The GCP exporter resolves its credential when it starts. On GKE the metadata server always supplies one, so the result is again empty pulls. Outside Google Cloud there is no metadata server, so GOOGLE_APPLICATION_CREDENTIALS has to point at a mounted key or a Workload Identity Federation configuration. Without one the exporter cannot start, and the gateway fails to load along with every log and metric it carries. Azure, like CloudWatch, resolves its credential on the first pull, so the gateway keeps running. On AKS without the pod label, the pull falls through to the node’s managed identity, which Resource Graph refuses with a 403. up is 0, and the gateway logs service discovery failed.

Lag#

Provider data is minutes old when it arrives, and that decides how it can be used.

SourceAge on arrivalTimestamp
CloudWatch RDSA few minutesThe scrape’s
CloudWatch S3Up to a dayThe scrape’s
Cloud SQLAbout three minutesCloud Monitoring’s
GCSOver ten minutesCloud Monitoring’s
Azure Flexible Server, Blob requestsAbout a minuteThe scrape’s
Azure Blob capacityUp to a dayThe scrape’s
CloudWatch EC2, Auto Scaling, vCPU usageA few minutesThe scrape’s
Compute Engine quota usage and limitUp to a day; measured at three points a dayThe scrape’s
Compute Engine quota refusalsA few minutesCloud Monitoring’s
Azure AKS autoscaler and node VMsAbout a minuteThe scrape’s

CloudWatch and Azure samples are stamped at scrape time. Stamping a daily S3 datapoint with its own time would put it a day in the past, and the bundled Thanos Receive, which accepts no out-of-order samples, refuses samples that old. Compute Engine quota usage and limits are stamped at scrape time for the same reason, since each pull reads back a day to find a point that may be hours old. A usage point therefore reads as current until the next one lands, and a change can take hours to show.

Cloud Monitoring samples keep their own timestamps. An instant query at “now” looks back five minutes by default and finds no GCS sample at all. Queries on provider families need last_over_time(<series>[15m]) or a wider window.

Every gateway restart leaves a gap as long as the lag. Remote write forwards only samples stamped after it started, so that a restart does not resend what was already written. A Cloud Monitoring sample stamped before the restart is therefore never sent. After each restart, Cloud SQL series resume a few minutes later and GCS series over ten minutes later. CloudWatch and Azure samples carry the scrape’s time, so they have no gap.

No provider signal backs a fast page. Provider alerts cover the slow-moving conditions — storage headroom, burst-credit exhaustion, connection ceilings — with for: windows well above the publication delay.

Cost#

The provider bills each pull, and the bill depends on configuration, not on how many dashboards are open.

ProviderCalls per pullMeasured
cloudwatch12 GetMetricStatistics calls per RDS instance and 2 per bucket30 calls for two instances and three buckets, in-cluster
gcpOne descriptor listing per metric prefix, and one time-series listing per matching metric type19 calls for two instances and two buckets, counted by stackdriver_monitoring_api_calls_total
azureOne Resource Graph query per service, then one metrics call per resource, since each set is under Azure’s twenty-per-call limit. Blob Storage is two pulls, capacity and requests9 calls for two servers and two accounts, counted from the exporter’s code, since it exports no call counter
cloudwatch EKSOne tag:GetResources and one DescribeAutoScalingGroups, one ListMetrics per metric, GetMetricData billed per metric requested, and one GetMetricStatistics for vCPU usage26 metrics requested for eight nodes and one node group, counted by yace_cloudwatch_getmetricdata_metrics_requested_total
gcp quotaFour descriptor listings and four time-series listings for usage and limit, and two of each for refusals12 calls, counted by stackdriver_monitoring_api_calls_total
azure AKSTwo Resource Graph queries, then one metrics call per cluster and one per node scale set5 calls for one cluster with two node pools

scrapeInterval is the main lever, and defaults to five minutes. Rates change, so current provider pricing is the reference for what a call costs.

The gateway runs several replicas, and every replica runs the exporter. The scrape of it is clustered, so one replica owns it and each provider is called once per interval.

Which destinations receive it#

Each provider assigns its families a tier through metricImportance, which defaults to extended. The families infra-cloud and infra-autoscaling draw are also named in the query registry, at diagnostic, the lowest tier. A registry tier admits a metric at that tier and above, so metricImportance decides for every destination floor above diagnostic.

Destination minMetricImportanceReceives provider families at the default
all (the bundled Thanos)Yes
diagnostic, extendedYes
recommended, essentialNo

A destination that bills per series, such as Datadog or a BYOC fan-out, therefore does not receive them unless its floor or the provider’s tier is changed.

The Google Cloud destination deserves the same care on an install that also pulls from GCP. Its default floor is recommended, which keeps the pulled families out. Raising either would write each series back into Cloud Monitoring as a prometheus.googleapis.com/ metric: a billed second copy of data Cloud Monitoring already holds. The pull does not read those types back, so it does not loop.

What is deliberately not pulled#

SignalWhy
GCS api/request_countA per-minute DELTA, and the exporter adds only the newest point of each pull to its counter. At a five-minute interval it reports about a fifth of the real count. The clients report the same requests exactly
S3 request metricsOpt-in per bucket on the AWS side, and billed as custom metrics
Azure Blob TransactionsThe window each pull reads ends at the scrape, so its newest five-minute bucket is a minute or so short, and a count read from it under-counts. Availability already falls with throttling and server errors, and the clients report the same requests exactly
Azure resource tagsThe exporter copies the owner tag onto every series by default. Tags are free text, and the pull copies none
Resources found by tagTag discovery pulls every matching resource in the account, and bills for each. EKS nodes are the exception, found by the cluster tag EKS and Karpenter set, because they are replaced too often to name
The EC2 vCPU quota’s limitCloudWatch publishes the usage but not the limit, which only CloudWatch metric math and the Service Quotas API return, and the exporter supports neither
Capacity refusalsNo provider publishes a metric for a launch it could not fulfil, such as InsufficientInstanceCapacity, ZONE_RESOURCE_POOL_EXHAUSTED or AllocationFailed. Karpenter and the cluster autoscaler report them
Azure compute quotaNot an Azure Monitor metric; it is only in the Compute usage API
GKE node pool instance groupsinstance_group/size has no target size beside it, so it says no more than the node count, and the group names are truncated beyond reliable matching
The GKE cluster autoscalerGKE publishes no metrics for it, only its visibility log in Cloud Logging and Kubernetes events
AKS kube_* and node_* platform metricskube-state-metrics and node-exporter already report the same, at full resolution

Checking it works#

On CloudWatch and GCP, up does not say whether a pull succeeded. An exporter whose provider call fails still answers its scrape, with no provider series in it. Measured with credentials missing — for CloudWatch anywhere, and for GCP where a metadata server exists — both exporters return HTTP 200 and report healthy, so up stays 1. up only says the exporter exists and is being scraped.

On Azure it says half of it. A Resource Graph query that fails — no identity, or one with no access to the subscription — fails the scrape, and up is 0. A metrics call that fails for one resource only drops that resource’s series and logs a warning. A resource the identity cannot read is never found at all. up stays 1 in both cases.

QuestionQuery
Is the exporter running and scrapedup{job=~"integrations/(cloudwatch|gcp|azure)"}
Did the last GCP pull failstackdriver_monitoring_last_scrape_error == 1
Is CloudWatch returning data, per instancecount by (dimension_DBInstanceIdentifier) (aws_rds_cpuutilization_average)
Is Cloud Monitoring returning data, per instancecount by (database_id) (last_over_time(stackdriver_cloudsql_database_cloudsql_googleapis_com_database_up[15m]))
Is Azure returning data, per servercount by (resourceName) (azure_microsoft_dbforpostgresql_flexibleservers_is_db_alive_minimum_count)
Are EKS nodes being foundcount(aws_ec2_status_check_failed_system_maximum), against the cluster’s node count
Is Compute Engine quota arrivingcount by (location) (stackdriver_compute_googleapis_com_location_compute_googleapis_com_quota_cpus_per_vm_family_usage)
Are AKS node VMs being foundcount by (resourceName) (azure_microsoft_compute_virtualmachinescalesets_vmavailabilitymetric_minimum_count)

CloudWatch publishes no equivalent of the GCP error series. Its yace_cloudwatch_getmetricstatistics_requests_total counts billed calls, but it is one counter per gateway replica, and every CloudWatch target a replica owns reports it. Measured with five targets across two replicas, the five series read 12, 14, 26, 28 and 2: running totals of 28 and 2 calls, not 82. It cannot be summed across instance, so the cost is best read from the configuration: 12 calls per RDS instance and 2 per bucket, each interval. A CloudWatch pull that fails for want of an identity or a grant is visible only as missing series and as errors in the gateway’s logs, such as Couldn't get account Id. An EKS pull missing tag:GetResources finds no nodes, and one missing autoscaling:DescribeAutoScalingGroups finds no node groups, while the vCPU usage still arrives.

The Collection tab of the Infrastructure Cloud Provider dashboard (infra-cloud) draws these checks for every provider it finds.