Skip to main content

Prometheus metrics

Infrahub exposes the metrics below in the Prometheus text format. The API server serves them on /metrics on its HTTP port, and each task worker serves them on the port set by INFRAHUB_METRICS_PORT (default 8000). Task-manager metrics come from the Prefect exporter included in the observability stack.

To choose which metrics to watch and configure scraping, see Metrics.

HTTP requests​

Recorded by the API server for every request except /health. The path label contains the route template, such as /api/schema/{schema_kind}, not the literal URL.

MetricTypeLabelsDescription
infrahub_requests_totalCountermethod, path, status_code, app_nameTotal HTTP requests
infrahub_request_duration_secondsHistogrammethod, path, status_code, app_nameHTTP request duration, with buckets at 0.1, 0.25 and 0.5 seconds
infrahub_requests_in_progressGaugemethod, app_nameHTTP requests currently in progress

Instance​

MetricTypeLabelsDescription
infrahub_infoGaugeversion, worker_idAlways 1. Identifies the Infrahub version and the process that reports it

GraphQL​

Recorded by the API server for each GraphQL request.

MetricTypeLabelsDescription
infrahub_graphql_duration_secondsHistogramtype, operation, branch, name, query_idQuery duration, in seconds
infrahub_graphql_response_size_bytesHistogramtype, operation, branch, name, query_idResponse size, in bytes
infrahub_graphql_query_depthHistogramtype, operation, branch, name, query_idQuery depth
infrahub_graphql_query_heightHistogramtype, operation, branch, name, query_idQuery height
infrahub_graphql_query_objectsHistogramtype, operation, branch, name, query_idNumber of objects in the query
infrahub_graphql_top_level_queriesHistogramtype, operation, branch, name, query_idNumber of top-level queries
infrahub_graphql_variablesHistogramtype, operation, branch, name, query_idNumber of variables
infrahub_graphql_errorsHistogramtype, operation, branch, name, query_idNumber of errors in the response
infrahub_graphql_generate_schemaHistogrambranchTime to generate the GraphQL schema of a branch

Database​

Recorded by every process that queries the database, so both the API server and the task workers report them.

MetricTypeLabelsDescription
infrahub_db_query_execution_secondsHistogramtype, query, runtime, context1, context2Execution time of a database query
infrahub_db_transaction_retries_totalCounternameTransactions retried after a transient error
infrahub_db_last_connection_pool_usageGaugeaddressLast known number of active connections in the driver pool
infrahub_db_reference_query_floor_secondsGaugeAll-time minimum execution time of the reference permission query
infrahub_db_reference_query_window_min_secondsGaugeMinimum execution time of the reference query over the recent sliding window
infrahub_db_reference_query_stress_ratio_medianGaugeRecent median reference-query time divided by the all-time floor. 1.0 means the database is not under stress

Locks​

MetricTypeLabelsDescription
infrahub_lock_acquire_secondsHistogramlock, typeTime to acquire a lock
infrahub_lock_reserved_duration_secondsHistogramlock, typeTime a lock is held by a client

API admission control​

Recorded by the API server when priority-aware backpressure is enabled (INFRAHUB_API_BACKPRESSURE_ENABLED, default true, see the configuration reference). The priority label is high, medium or low, taken from the X-Priority request header.

MetricTypeLabelsDescription
infrahub_admission_offered_totalCounterpriorityRequests entering the admission layer
infrahub_admission_admitted_totalCounterpriorityRequests admitted
infrahub_admission_rejected_totalCounterpriority, reasonRequests rejected with 429 Too Many Requests
infrahub_admission_in_flightGaugepriorityAdmitted requests currently running
infrahub_admission_waitersGaugepriorityRequests queued for a slot
infrahub_admission_sojourn_secondsHistogrampriorityTime a request waited for a slot
infrahub_admission_max_concurrencyGaugeEffective number of admission slots per API worker
infrahub_admission_missing_priority_totalCounterRequests with a missing or invalid X-Priority header
infrahub_admission_sustained_load_secondsGaugeSeconds the worker has continuously seen database stress at or above the significant-load ratio, 0 below it

Task manager (Prefect exporter)​

The observability stack runs prometheus-prefect-exporter 3.3.0 against the task manager API. The metrics most useful for Infrahub are:

MetricLabelsDescription
prefect_work_queues_late_runs_countwork_queue_name, work_pool_name, priority, is_paused, status, healthy, type, health-check policy labelsFlow runs in the queue in the Late state, which the task manager sets 15 seconds after the scheduled start
prefect_info_work_queuessame as aboveWork queue state; 0 when paused, 1 otherwise
prefect_info_work_poolswork_pool_name, type, is_paused, statusWork pool state; 0 when paused, 1 otherwise
prefect_info_flow_runsdeployment_name, flow_name, state_name, work_queue_nameNumber of flow runs per state, counting only runs whose start time, or scheduled start time if not started, is within the last OFFSET_MINUTES (default 3)
prefect_deployment_failed_flow_runsdeployment_name, flow_name, last_failed_run_id, state_nameLast failed or crashed flow runs per deployment, within FAILED_RUNS_OFFSET_MINUTES (default 7 days)
prefect_flow_runs_total_run_timeflow_nameTotal run time of flow runs, in seconds
prefect_flow_runs_total, prefect_flows_total, prefect_deployments_total, prefect_work_pools_total, prefect_work_queues_totalObject counts

Dependencies​

The observability stack also scrapes Neo4j on port 2004 and RabbitMQ on port 15692. Their metrics are documented by each project: Neo4j metrics and RabbitMQ Prometheus plugin.