Prometheus

Deploy Prometheus to collect and query time-series metrics, define alert rules, and monitor applications and infrastructure.

On this page

Prometheus collects time-series metrics

Prometheus scrapes configured metric endpoints, stores observations as labelled time series, evaluates PromQL queries and can evaluate alerting rules. Exporters expose metrics for systems that do not natively provide a Prometheus endpoint, and service discovery can maintain target lists in changing environments. It is well suited to service and infrastructure monitoring when the metric model is kept disciplined.

It is not a generic event archive. A metric should answer a repeated question about rate, duration, saturation, availability or a bounded state. If the problem needs every request payload, arbitrary event search or long narrative history, use an appropriate log, trace or record system alongside Prometheus.

Instrument for questions, not vanity

The most expensive label is often the one nobody planned to query.

Names and units

Use names and units that make a query readable. Distinguish counters, gauges and histograms according to their documented meaning. Write recording rules only after agreeing what they summarise.

Labels

Keep labels bounded and meaningful. Avoid unbounded identifiers such as request IDs, email addresses or raw URLs; high cardinality increases resource use and makes queries harder to operate.

Service health

Start with indicators tied to user experience and known dependencies. Link each alert to a service owner and a first diagnostic step.

Exporter scope

Expose only the metrics needed and protect endpoints where appropriate. An exporter may reveal topology, versions or workload details that are not for public access.

Plan the scrape and storage path

Document this path before targets are added in bulk.

PartDecisionValidation
TargetsStatic configuration, service discovery or a controlled mix.Target changes appear and disappear as expected without stale entries being ignored.
Prometheus serverRetention, local disk, rule evaluation and operational access.A representative ingestion rate fits storage and query expectations.
AlertmanagerGrouping, routing, silences and receiver ownership.A controlled alert routes to one intended receiver and clears correctly.
Remote storageWhether longer retention or global querying needs remote write or read.Failure behaviour and recovery are documented and tested for the chosen integration.

Choose topology from failure domains

A single Prometheus server can be suitable for a small, bounded environment. Higher availability, multiple environments or longer retention introduce different needs: separate servers, federation, remote storage or other documented patterns. Select a design from the recovery and query requirement, not because a diagram with more components looks safer.

Keep configuration, rule files, service-discovery access and persistent storage under change control. A reload that succeeds is only the first test. Check target health, a known query, rule evaluation and alert delivery after a deployment. Record the version and configuration source so an operator can explain what was running during an incident.

Build alerts that have an owner

Write the failure condition, evaluation period, severity, owner and known false positives. A graph trend is not automatically a paging alert.

Protect metrics and operational access

Metric endpoints, service-discovery credentials and Prometheus administration can reveal internal topology, workload and system state. Limit network access, use appropriate authentication and TLS for the selected deployment, and keep service credentials out of configuration committed to an open repository. A scrape job needs only the access required to collect its target.

Review labels for sensitive values before broad instrumentation. Avoid placing secrets, tokens, personal data or unbounded identifiers in labels. Prometheus may retain and expose them through queries, dashboards and alerts. Data minimisation belongs in instrumentation design, before a metric reaches storage.

Adopt Prometheus in measured steps

Inventory existing monitors, dashboards, alert receivers and target ownership. Select a few services where a metric question and on-call owner already exist. Instrument or configure those targets first, then compare the new alerts and graphs with the existing operational view. This reveals label gaps and meaningless thresholds before the target list becomes large.

Move alert routing one service at a time. Keep the old monitor until the new rule has produced expected data and a controlled notification. Archive obsolete checks with their rationale. Two active alerts for the same symptom create confusion, not redundancy, unless their different purposes are clear.

Prometheus operating duties

Metrics are a shared platform. The platform team and service owner have distinct jobs.

Platform

  • Monitor target health, rule evaluation, disk use and query performance.
  • Back up or protect configuration and document storage recovery.
  • Test upgrade and remote-storage changes.

Service teams

  • Own metric semantics, dashboards and alert response.
  • Bound labels and remove dead instrumentation.
  • Update rules when service behaviour changes.

Incident response

  • Maintain Alertmanager receiver ownership.
  • Review noisy alerts and stale runbooks.
  • Record data gaps discovered during incidents.

Change record

  • Keep scrape, rule, retention and remote-storage changes together.
  • Test alert evaluation and recovery before production rollout.
  • Record the expected data gap and rollback signal for each change.

Prometheus questions to answer

Why does label cardinality matter?

Each distinct label combination creates a time series. Unbounded values can cause memory, storage and query pressure, making the monitoring system less reliable precisely when it is needed.

Does Prometheus need remote storage?

Not always. Decide from retention, query scope and recovery needs. Local storage may fit a bounded environment; longer-term or multi-server designs may need a compatible remote approach.

What should a scrape failure alert say?

Name the target or job, owner and first check. Distinguish a missing target from an application error where possible, so the responder knows whether to inspect discovery, network or the service.

Who owns an alert threshold?

The service owner owns what the signal means and its response. The monitoring platform owner supports the rule and route. Both should be visible in the runbook.

Scope and handover

Target count, scrape interval, label cardinality, retention, rule count, high-availability goals, remote storage and existing alert migration all affect the work. Capacity assumptions should come from representative ingestion and query behaviour, not only host CPU.

Handover includes configuration location, target ownership, rule and receiver ownership, retention setting, storage recovery result, supported exporters, access controls and the upgrade procedure for the deployed version. Also document one working query and alert test for each important service. A future operator needs evidence of expected behaviour, not only a directory of rule files. Add the service-discovery source, job name, expected target count and target health check for every important scrape group. This helps distinguish a dead service from a configuration or discovery failure. Review this inventory when platform teams change discovery labels, namespaces or network boundaries; silent target loss is easier to prevent than to reconstruct during an outage. Include the last target-review date and a contact for the discovery integration owner.

Sources and further reading

Talk to our team.

Tell us what you're working on, whether it's a deployment, an audit, a security test or a cyber range. You'll speak with an engineer who can help you scope it.

  • 30-minute call: free, with no obligation.
  • NDA on request: we can sign before you share details.
  • Clear next steps: a scope and plan after the call.