Home › Observability › Prometheus Cardinality Explosion: Find and Fix It
Advanced 6 min · September 23, 2026

Prometheus Cardinality Explosion: Find and Fix It

Prometheus OOMKilled by cardinality? Do the series memory math, drop user_id labels at scrape, and guard with sample_limit..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Notes here come from systems that actually shipped.

Follow
✓ Production
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 14 min
  • ✓Prometheus query access plus head-series dashboards
  • ✓Ability to edit and reload prometheus.yml scrape configs
  • ✓Familiarity with relabeling concepts at a basic level
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Each active series costs 2-3KB of head RAM, so a user_id label with 250K users burns half a gigabyte per metric
  • Find the offender with topk series counts by metric name, then count values of the suspect label
  • Drop explosive labels at scrape with metric_relabel_configs and serve dashboards from aggregated recording rules
  • Set per-job sample_limit near 150% of normal size so the next bad deploy fails loudly instead of OOMKilling
✦ Definition~90s read
What is Prometheus Cardinality Explosion From a Label?

Cardinality in Prometheus means the number of distinct series produced by unique combinations of metric names and label values. A counter named http_requests_total with method (4 values) and status (5 values) creates 20 series. Add user_id with 250,000 values and the same counter creates 5 million series — each consuming roughly 2 to 3KB of memory in the TSDB head for its chunks, label pairs, and index entries.

★
Think of Prometheus as a library where every unique label combination needs its own labeled shelf, and each shelf costs money to maintain.

The head holds all active series in memory for fast appends and queries, which is why cardinality kills with OOMKill while disk sits half empty. Retention, compaction, and block storage govern disk; only the head's series count governs this memory cliff. Growth from bounded labels plateaus, but unbounded labels grow with every new user, session, or request ID until the pod limit arrives.

Production control has two levers. metric_relabel_configs drop or rewrite labels after scraping but before storage, deleting explosive dimensions at the door. Recording rules precompute the aggregations dashboards need (sums by tier, rates by endpoint) so raw high-cardinality series can be dropped without losing business answers. sample_limit caps samples per scrape as a circuit breaker, failing loudly on the next bad deploy instead of silently filling the head.

Plain-English First

Think of Prometheus as a library where every unique label combination needs its own labeled shelf, and each shelf costs money to maintain. Labels like method with 4 values need 4 shelves — cheap. A user_id label with 250,000 users needs 250,000 shelves overnight, and the library goes bankrupt by morning. The fix is refusing to build per-person shelves: keep totals by section instead, and look up individual readers in the logbook where that detail belongs.

The page comes at 3 AM: Prometheus OOMKilled, restarted, replaying WAL, OOMKilled again. Disk has terabytes free. Retention is a modest 15 days. The killer is invisible on every disk dashboard because it lives in RAM: active series in the TSDB head, each costing 2 to 3KB, multiplied by a label that seemed harmless in code review.

High-cardinality labels are how this happens. A user_id, request path with embedded IDs, or session token on a single counter multiplies one metric into hundreds of thousands of series. The math is brutal and fast: 250,000 users times a few label combinations crosses a 16 GB pod limit in hours, and the crash loop that follows gaps every dashboard at once.

This guide gives you the memory math, the queries that name the offending label in minutes, and the two production fixes — dropping labels at scrape time and aggregating with recording rules. You'll leave with sample_limit guardrails and alerts that catch the next explosion while it's still a graph wiggle, not a 3 AM page.

Cardinality Math: Why One Label Costs Gigabytes

Every unique combination of metric name and label values becomes one active series, and each one costs roughly 2 to 3KB of head memory for chunks, labels, and index entries. That sounds trivial until multiplication kicks in: a counter with user_id (250K values) times status (3 values) is 750,000 series and around 2 GB of RAM from a single metric.

Do the math before any label ships. Multiply the expected value counts of every label on the metric, multiply by 3KB, and compare against headroom. If one label's growth is unbounded — users, sessions, request IDs — the product grows forever and the pod limit is just a countdown.

Head memory is the constraint that matters. Disk retention, block size, and WAL replay all matter operationally, but OOMKill comes from the head. Dashboards for prometheus_tsdb_head_series and process_resident_memory_bytes belong next to every deploy pipeline, reviewed as carefully as latency graphs.

Break the 2 to 3KB into parts you can reason about. Label pairs cost bytes each plus index entries; the current head chunk holds up to 120 samples before sealing; the inverted index maps every label value to its series. Small labels on few series cost almost nothing, which is why method and status never hurt. The same bytes on 250,000 user values multiply into gigabytes because every part scales with series count. Run the multiplication for each new label in review: values times combinations times 3KB. Numbers on a napkin beat OOMKills at 3 AM every time.

PROMQL
1
2
3
4
5
6
7
8
9
10
11
# Active series in the head — the number that kills pods
prometheus_tsdb_head_series

# Memory tracking head growth (alert at 80% of pod limit)
process_resident_memory_bytes

# Top 10 metrics by series count — the offender is usually #1
sort_desc(topk(10, count by (__name__) ({__name__=~".+"})))

# Scrape size per job — the jumping job owns the bad label
scrape_samples_scraped{job="checkout"}
📊 Production Insight
A checkout counter with user_id and payment_method labels hit 3.7M series at ~2.5KB each — 9 GB of head memory from one metric. The math on a napkin would have blocked the deploy.
🎯 Key Takeaway
Multiply label value counts by 3KB per series before shipping; unbounded labels mean unbounded memory.

Which Labels Explode and Which Stay Safe

Labels split into two families. Bounded labels like method, status_code, and route templates have small fixed value sets — 4 methods, 5 statuses, 50 routes — and their products stay flat for years. They're the reason labels exist and are safe by construction.

Unbounded labels grow with traffic: user IDs, emails, session tokens, request UUIDs, and raw URL paths with embedded IDs. Each new user permanently adds series, so the head grows monotonically until the pod dies. These values belong in logs, traces, or warehouses — systems built for per-entity detail — never in metric labels.

The treacherous middle is labels that look bounded until errors strike: error messages with dynamic text, file paths with temp dirs, pod names in long-lived clusters. Normalize at instrumentation — lowercase, bucket, allowlist — so a new error string can't mint ten thousand series overnight.

Histograms multiply everything they touch. A histogram with 12 buckets and labels for method and status creates 12 times 4 times 5 series per metric name — plus _count and _sum. That is 242 series where a counter would create 20, before any unbounded label arrives. Native histograms compress this dramatically by storing bucket spans instead of indexed per-bucket series, which is why newer clients prefer them. Until you migrate, treat every histogram label as twelve labels in review. The bucket count is a hidden multiplier most teams forget to multiply.

📊 Production Insight
An error_message label with dynamic SQL text minted 40,000 series during one bad migration. Bucketing errors to codes at instrumentation would have kept it at 12 series.
🎯 Key Takeaway
Bounded value sets are safe; anything growing with traffic belongs in logs, not labels.

Finding the Offending Label in Minutes

When head series spike, find the owner fast. Rank metrics by series count with a topk query — the offender is usually number one by a wide margin. Then count values of its labels to find the runaway: a payment_method with 4 values is innocent, a user_id with 250,000 values is guilty.

Correlate with scrape sizes. The scrape_samples_scraped series per job shows which job's exposition ballooned and exactly when, pointing at the deploy that added the label. Cross-check against deploy logs and the diff usually confesses within minutes.

Move quickly but measure first. Record the head series count, memory, and the offending label's value growth before changing anything. Those baselines prove the fix worked and size the guardrails you'll add next.

The TSDB status API gives machine-readable ground truth when PromQL feels indirect. GET /api/v1/status/tsdb returns head statistics including total series, chunk counts, and compaction numbers — scrape it into your capacity dashboard. For label-level forensics, /api/v1/label/<name>/values with match[] selectors counts values per label without writing queries. Compare value counts across a deploy boundary: the label whose cardinality jumps names the commit. When topk queries time out on overloaded servers, the status API answers in milliseconds because it reads head metadata, not series data. Add both endpoints to your incident runbook next to the PromQL versions.

PROMQL
1
2
3
4
5
# Value counts for the suspect label — runaway = tens of thousands+
count by (payment_method) (checkout_completed_total)

# Confirm growth started at a deploy (7-day window)
count(count by (__name__) ({__name__="checkout_completed_total"}))
📊 Production Insight
One team found their offender 4 minutes after the page: topk showed checkout_completed_total at 3.7M series, and a value count showed user_id with 250K values. Total diagnosis time beat the alert's for-duration.
🎯 Key Takeaway
Topk by metric, count values per label, correlate the jump with deploy time.

Dropping and Aggregating: the Two Production Fixes

The production fix drops the explosive label at scrape time, before it ever reaches the head. metric_relabel_configs rules run after scraping against metric labels, and a labeldrop or regex drop on user_id deletes the dimension from every affected series. The metric survives aggregated; the per-user detail never enters memory.

The classic trap is placing the rule in relabel_configs, which runs before scraping against target labels only. The config reloads cleanly, the rule matches nothing, and cardinality climbs untouched. Always use metric_relabel_configs for metric labels — the name is the documentation.

Dashboards that genuinely need segments get recording rules instead. Sum the counter by plan tier or endpoint group into a new aggregated series, point dashboards at it, and drop the raw high-cardinality originals. You keep every business answer while deleting millions of stored series.

Choose the drop granularity deliberately. Dropping a whole metric with source_labels __name__ suits counters that should never have shipped; labeldrop suits metrics that are valuable without one explosive dimension. The keep action inverts the logic for allowlisting: keep only named metrics from a noisy exporter and drop the hundred you never query. Order matters — rules evaluate top to bottom, so place cheap exact matches before expensive regexes. Test every rule with promtool test rules or a staging reload plus series counts; a miswritten regex can drop production alerts silently. Review the rule file quarterly because exporters add metrics on upgrade.

prometheus.ymlYAML
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
# Drop the explosive label before it reaches the head
scrape_configs:
  - job_name: checkout
    sample_limit: 50000
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: 'checkout_completed_total'
        action: drop
        # Or keep the metric but drop just the label:
      - regex: 'user_id'
        action: labeldrop

# Recording rule preserving the business answer by segment
# rules/checkout.yml
- record: job:checkout_completed_total:by_tier
  expr: sum by (plan_tier) (checkout_completed_total)
📊 Production Insight
Dropping user_id at scrape cut a fleet from 4.1M to 380K head series in a day. A plan-tier recording rule kept every dashboard working — nobody missed the raw series.
🎯 Key Takeaway
Drop explosive labels in metric_relabel_configs; serve segments from recording rules, never raw series.

Guardrails: sample_limit and Alerts That Page Early

sample_limit is the circuit breaker that converts silent growth into a loud, immediate scrape failure. Set per job near 150% of its measured normal scrape size: normal variance passes, a label explosion fails the whole scrape and pages someone before the head fills.

Pair it with two alerts. One fires when any job's scrape size crosses 80% of its limit — the early warning that buys weeks. The other watches head-series week-over-week growth above 20% and memory above 80% of the pod limit. Together they catch both the sudden deploy mistake and the slow unbounded creep.

Treat limit tuning as capacity work. Recompute normal sizes quarterly, review every new label in code review with the 3KB math, and require load-test evidence for metrics on hot paths. Guardrails decay as fleets grow; the review habit is what keeps them honest.

Know the two sibling limits. sample_limit caps samples per scrape for one target and fails that scrape loudly; target_limit caps the total across all targets of a job and spreads protection wider. Set sample_limit near 150 percent of normal per-target size and target_limit near 150 percent of the job total. When a scrape exceeds the limit the entire scrape fails — including healthy series — so pair limits with alerts at 80 percent that fire before the cliff. Document each limit beside the measurement that sized it; unexplained limits get deleted in cleanups and the protection vanishes. Limits are capacity decisions wearing config clothes.

PROMQL
1
2
3
4
5
6
7
# Early warning: job approaching its sample limit
scrape_samples_scraped / ignoring() group_left
  scrape_sample_limit > 0.8

# Slow creep: head series growing week over week
predict_linear(prometheus_tsdb_head_series[1d], 7 * 86400)
  > prometheus_tsdb_head_series * 1.2
📊 Production Insight
After adding limits, a team caught the next explosion — a request_path label on an API counter — at 80% of limit during business hours. Fix took 20 minutes; the previous identical incident had been a 3 AM OOMKill loop.
🎯 Key Takeaway
sample_limit at 150% of normal plus growth alerts turn 3 AM OOMKills into daytime warnings.

Recovering Without Just Buying Bigger Pods

Recovery from an active explosion follows a fixed order. First drop the label at scrape to stop the bleeding — head series flatten within a couple of compactions. Then point dashboards at aggregated recording rules so nothing visible breaks. Then set sample_limit and growth alerts so it can't recur silently.

Resist the urge to fix retention or resize pods instead. Retention governs disk blocks, and bigger pods just reschedule the crash by the exact ratio of added memory to growth rate. Both feel productive while the series math keeps compounding underneath.

Finally, move the displaced detail somewhere it belongs. Per-user analytics live happily in warehouses fed by event streams; debugging detail lives in logs and traces. Prometheus keeps the aggregates that drive alerts and dashboards — that's the division of labor that survives scale.

Dropping labels affects new samples only — existing high-cardinality series age out through retention, not instantly. Expect head series to fall over hours as old chunks compact away, not seconds after reload. For immediate relief on a crashing server, the snapshot API plus a storage wipe and replay buys a clean head, but prefer patience unless OOMKills loop. Old series with dropped labels remain queryable under their original names until retention expires, which keeps history intact while stopping growth. Verify the decline with head-series graphs across two compaction cycles before declaring victory. The fix works on geological time; watch it like a geologist.

💡The One Label Rule Worth Enforcing
Never add a label whose values grow with traffic — user IDs, emails, session tokens, UUIDs, raw paths — no matter how useful per-user graphs sound in planning. Each value is a permanent series at 2-3KB, and deleting the label later doesn't refund the OOMKills it caused.
📊 Production Insight
A fleet that cut retention from 30d to 7d saw zero memory change. The 4-line drop rule that followed freed 18 GB overnight. The retention rollback took longer to approve than the actual fix took to work.
🎯 Key Takeaway
Drop, aggregate, guard — in that order; retention cuts and bigger pods only reschedule the crash.
● Production incidentPOST-MORTEMseverity: high

The user_id Label That Built 3.7M Series in 9 Days

Symptom
Prometheus OOMKilled nightly at peak for 9 days straight, each restart taking 25 minutes of WAL replay. Head series hit 4.1M, memory sat at 15.6 GB against a 16 GB limit, and every dashboard gapped during the crash loops.
Assumption
The team assumed retention was the problem and cut it from 30 days to 7, expecting memory to fall. It didn't move, because retention governs disk blocks while the killer was in-memory head series. They then doubled the pod to 32 GB, which bought 11 days before the next OOMKill.
Root cause
A new checkout_completed_total counter shipped with user_id and payment_method labels. With 250,000 daily users, that single metric generated 3.7M series at ~2.5KB each — roughly 9 GB of head memory. The pod's 16 GB limit died nightly during peak traffic, and each restart replayed a larger WAL until crash-looping.
Fix
They dropped user_id at scrape with a metric_relabel_configs rule, added a recording rule summing checkouts by plan tier for the dashboard that needed per-segment data, and set sample_limit: 50000 on the checkout job. Head series fell from 4.1M to 380K in a day, memory dropped from 24 GB to 6 GB, and per-user analytics moved to the warehouse where it belonged.
Key lesson
  • Retention controls disk, not head memory — cutting retention never fixes a cardinality OOMKill.
  • Bigger pods only reschedule the crash; the label math grows with users and always catches up.
  • Every new label needs a cardinality review in code review, with sample_limit as the backstop when review misses one.
Production debug guideFive steps from OOMKill suspicion to a guarded, drop-ruled fleet.5 entries
Symptom · 01
Suspected OOMKill loop with unclear cause
→
Fix
Query prometheus_tsdb_head_series for the 7-day trend and process_resident_memory_bytes alongside it. If series climbed steeply while memory tracked it upward, you've confirmed a cardinality incident rather than a leak or query overload. Note the start time — it points at the deploy that introduced the label.
Symptom · 02
Head series climbing but offending metric unknown
→
Fix
Run sort_desc(topk(10, count by (__name__) ({__name__=~".+"}))) to find which metrics hold the most series, then drill into the winner with count by (suspect_label) to see value counts. A single label with tens of thousands of values (user IDs, paths) names the offender in under a minute.
Symptom · 03
Too many metrics to check by hand
→
Fix
Query scrape_samples_scraped{job="suspect-job"} over the incident window and compare jobs. The job whose scrape size jumped at the incident start owns the bad instrumentation. Check that job's recent deploys for new labels — the diff usually shows the added label line plainly.
Symptom · 04
Offending label confirmed, need the bleeding stopped
→
Fix
Add a metric_relabel_configs drop for the explosive label on the owning job, reload with kill -HUP, and watch head series flatten within two compactions. Keep dashboards on an aggregated recording rule so dropping raw series breaks nothing visible. Verify identical drops on every HA replica.
Symptom · 05
Fixed this time, need it to never page at 3 AM again
→
Fix
Set sample_limit on every job at 150% of its normal scrape size and alert at 80% of that limit. Add a head-series growth alert (20% week-over-week) and a memory alert at 80% of the pod limit. Future explosions now page as warnings weeks before they become OOMKills.
Cardinality Causes Compared — Act on Evidence
Root CauseHow to ConfirmFixPrevention
Unbounded label (user ID, path, UUID)Top series by label count dominated by one label; count grows dailyDrop label at scrape; aggregate via recording ruleLabel review in code review; sample_limit per job
Status-code style label with crash valuesNew label values appear after errors or deploysNormalize values (lowercase, bucketed) at instrumentationAllowlist label values in client libraries
Duplicate jobs scraping same targetsSame series under two job labels; head doubles overnightRemove overlap; keep one job per target setInventory audit; alert on series-per-target jumps
Stale churn from short-lived podsChurn rate high but current series flat; WAL replay slowFaster staleness, stable labels, shorter unmatched retentionStable instance labels; avoid pod-name in labels
⚙ Quick Reference
4 commands from this guide
FileCommand / CodePurpose
prometheus_tsdb_head_seriesCardinality Math
count by (payment_method) (checkout_completed_total)Finding the Offending Label in Minutes
prometheus.ymlscrape_configs:Dropping and Aggregating
scrape_samples_scraped / ignoring() group_leftGuardrails

Key takeaways

1
Each active series costs 2-3KB of head memory
multiply before adding any label.
2
Never put unbounded values (user IDs, UUIDs, raw paths) in labels; aggregate at instrumentation.
3
Drop explosive labels with metric_relabel_configs, not relabel_configs.
4
Set per-job sample_limit at 150% of normal as a circuit breaker.
5
Aggregate into recording rules, then drop raw series to reclaim memory.
6
Alert on head series and memory so growth pages you weeks before OOMKill.

Common mistakes to avoid

5 patterns
×

Adding user_id, email, or session labels to Prometheus metrics

Symptom
Series count grows with your user base forever, and each deploy that adds users pushes memory closer to the OOMKill line.
Fix
Drop the label at scrape time with metric_relabel_configs and keep the detail in logs or traces where high-cardinality data belongs. If you need per-user counts, emit a pre-aggregated metric from the app instead.
×

Keeping raw high-cardinality series for dashboard convenience

Symptom
Dashboards work but the head holds millions of series nobody queries directly, paying memory for convenience.
Fix
Aggregate in the recording rule (sum by without the explosive label) and drop the raw series with a shorter retention or a drop rule. Dashboards read the rule output, never the raw metric.
×

Running scrape jobs with no sample_limit guardrail

Symptom
One bad deploy doubles a job's series and there's no circuit breaker — the whole server OOMs instead of one job failing loudly.
Fix
Set sample_limit per scrape job based on measured scrape sizes, and alert at 80% of the limit. A scrape that exceeds the limit drops the entire scrape, which is loud and immediate.
×

Capacity planning on disk retention while ignoring head memory

Symptom
Disk is half empty when the pod OOMKills, because the killer was in-memory active series, not retained blocks.
Fix
Track prometheus_tsdb_head_series and process_resident_memory_bytes on every capacity review. Budget roughly 2-3KB per series and page before the head reaches 80% of the pod limit.
×

Putting labeldrop rules in relabel_configs instead of metric_relabel_configs

Symptom
The drop rule deploys cleanly, config reloads fine, and cardinality keeps climbing because the rule never matched anything.
Fix
Use labeldrop or labelkeep in metric_relabel_configs, which run after scrape and only affect storage. Relabel_configs run before scrape and can't see metric labels — the drop silently does nothing.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
What is cardinality, and why does one bad label hurt so much?
Q02SENIOR
Head series tripled overnight. Walk through your diagnosis.
Q03SENIOR
Why did a labeldrop rule fail to reduce cardinality?
Q04SENIOR
Explain sample_limit as a circuit breaker and its trade-off.
Q05SENIOR
Why must drop rules deploy identically across HA replicas?
Q01 of 05JUNIOR

What is cardinality, and why does one bad label hurt so much?

ANSWER
Cardinality is the number of unique label-value combinations, and each combination becomes an active series costing 2-3KB of head memory. A user_id label with 200,000 users on one counter creates 200,000 series overnight — roughly half a gigabyte of RAM from a single label on a single metric.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
How much memory does one series cost?
02
Are labels like method and status safe?
03
Can recording rules fix cardinality?
04
Does remote-write to long-term storage help head memory?
05
What sample_limit should a job use?
06
Where should per-user data live instead?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Notes here come from systems that actually shipped.

Follow
✓ Verified
production tested
September 27, 2026
last updated
2,085
articles · all by Naren
🔥

That's Prometheus. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Prometheus Query Returns No Data but Target Is Up
3 / 4 · Prometheus
Next
Prometheus rate() vs increase() — Wrong Alert Thresholds
→