Finding and controlling high-cardinality metrics

Finding and controlling high-cardinality metrics guidance for observability and data platforms teams.

On this page

Scope and reader

Metric cardinality is the number of distinct time series created by label combinations. Labels such as request ID, email address, session token, raw URL or unbounded error text can create a new series for each value. That increases memory, storage and query load without producing a useful aggregate.

Find expensive metrics by inspecting series counts, label-value growth and the queries that depend on them. Work with service owners before deleting a label: replace a raw path with a route template, move request-level context to traces or logs, or use bounded categories that answer the intended question.

Decision context

Set a review rule for instrumentation changes. New labels need a purpose, expected value range and owner. Monitor cardinality after releases, because a label that was bounded in a test environment can expand when a new tenant, route or integration reaches production.

The boundary should be readable by someone who was not in the original meeting. Name the environment, data path, access path, and expected service behavior. If an assumption is untested, mark it as an open item and give it an owner.

Criteria and tradeoffs

Use these criteria to compare approaches for finding and controlling high-cardinality metrics without hiding the work behind a single recommendation.

Boundary

List the assets and dependencies that make finding and controlling high-cardinality metrics work.

Access

Use named accounts, least privilege, and a reviewable approval path.

Verification

Test the expected result and record the failure signal that would trigger a pause.

Handover

Give the next operator the runbook, owner, recovery path, and open decisions.

Responsibility and evidence

Turn the plan into a review record with one owner and one observable result for each area.

AreaWorking record
ScopeName the systems, people, data, and decisions included in Finding and controlling high-cardinality metrics. Record what remains outside the review.
OwnershipAssign an operational owner for Finding and controlling high cardinality metrics, an approver for changes, and a contact for incidents or blocked work.
EvidenceKeep the configuration, test result, decision record, and exception owner together so another person can review the result.
RecoveryWrite the stop condition, rollback limit, restore dependency, and follow-up review before the change starts.

Implementation questions

Who needs to be involved in Finding and controlling high-cardinality metrics?

Start with the person who owns the service or control, then include the operator who performs the work and the reviewer who accepts the evidence. Finding and controlling high-cardinality metrics guidance for observability and data platforms teams. Keep the final decision with the accountable team.

What should be written down before work begins?

Record the boundary, assumptions, dependencies, allowed access, acceptance check, and rollback condition. For finding and controlling high-cardinality metrics, a short record is more useful than a broad promise.

How do we know the work is complete?

Use an observable check: a test result, configuration comparison, owner sign-off, or restore exercise. State who reviews it and where the record lives.

What happens when the expected path fails?

Stop at the agreed condition, preserve evidence, notify the owner, and use the documented fallback. Do not turn an unreviewed exception into a production default.

A working sequence

Define the boundary and acceptance check for finding and controlling high-cardinality metrics. Confirm the owner, dependencies, access window, and stop condition before touching the target system.

Failure modes to test

A useful review tests the path that is likely to break: an unavailable dependency, an expired credential, an unexpected data shape, a failed update, or an operator without the required access. Choose the failure that fits finding and controlling high-cardinality metrics and define the safe response.

Keep the example illustrative. Do not treat a successful test in one environment as proof that every deployment behaves the same way. Record the limits of the test and the evidence needed to repeat it.

Handover checks

Before the work is accepted, check the operating details that disappear when a project closes.

Owner and access

Named owner, support contact, approved access path, and revocation process.

Recovery record

Backup or fallback tested, dependency order written down, and recovery owner identified.

Open decisions

Exceptions, follow-up dates, and unresolved scope questions are visible to the next reviewer.

Next review

Close the page with the next concrete action for finding and controlling high-cardinality metrics: confirm the owner, gather the missing evidence, run the acceptance check, or schedule a scoped review.

Revisit the record when the system, dependency, access model, or operating responsibility changes. That keeps the page tied to the deployed environment rather than a one-time design discussion.

Sources and further reading

Talk to our team.

Tell us what you're working on, whether it's a deployment, an audit, a security test or a cyber range. You'll speak with an engineer who can help you scope it.

  • 30-minute call: free, with no obligation.
  • NDA on request: we can sign before you share details.
  • Clear next steps: a scope and plan after the call.