Prometheus Cardinality Explosion: Find and Fix It
Prometheus OOMKilled by cardinality? Do the series memory math, drop user_id labels at scrape, and guard with sample_limit..
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
- ✓Prometheus query access plus head-series dashboards
- ✓Ability to edit and reload prometheus.yml scrape configs
- ✓Familiarity with relabeling concepts at a basic level
- Each active series costs 2-3KB of head RAM, so a user_id label with 250K users burns half a gigabyte per metric
- Find the offender with topk series counts by metric name, then count values of the suspect label
- Drop explosive labels at scrape with metric_relabel_configs and serve dashboards from aggregated recording rules
- Set per-job sample_limit near 150% of normal size so the next bad deploy fails loudly instead of OOMKilling
Think of Prometheus as a library where every unique label combination needs its own labeled shelf, and each shelf costs money to maintain. Labels like method with 4 values need 4 shelves — cheap. A user_id label with 250,000 users needs 250,000 shelves overnight, and the library goes bankrupt by morning. The fix is refusing to build per-person shelves: keep totals by section instead, and look up individual readers in the logbook where that detail belongs.
The page comes at 3 AM: Prometheus OOMKilled, restarted, replaying WAL, OOMKilled again. Disk has terabytes free. Retention is a modest 15 days. The killer is invisible on every disk dashboard because it lives in RAM: active series in the TSDB head, each costing 2 to 3KB, multiplied by a label that seemed harmless in code review.
High-cardinality labels are how this happens. A user_id, request path with embedded IDs, or session token on a single counter multiplies one metric into hundreds of thousands of series. The math is brutal and fast: 250,000 users times a few label combinations crosses a 16 GB pod limit in hours, and the crash loop that follows gaps every dashboard at once.
This guide gives you the memory math, the queries that name the offending label in minutes, and the two production fixes — dropping labels at scrape time and aggregating with recording rules. You'll leave with sample_limit guardrails and alerts that catch the next explosion while it's still a graph wiggle, not a 3 AM page.
Cardinality Math: Why One Label Costs Gigabytes
Every unique combination of metric name and label values becomes one active series, and each one costs roughly 2 to 3KB of head memory for chunks, labels, and index entries. That sounds trivial until multiplication kicks in: a counter with user_id (250K values) times status (3 values) is 750,000 series and around 2 GB of RAM from a single metric.
Do the math before any label ships. Multiply the expected value counts of every label on the metric, multiply by 3KB, and compare against headroom. If one label's growth is unbounded — users, sessions, request IDs — the product grows forever and the pod limit is just a countdown.
Head memory is the constraint that matters. Disk retention, block size, and WAL replay all matter operationally, but OOMKill comes from the head. Dashboards for prometheus_tsdb_head_series and process_resident_memory_bytes belong next to every deploy pipeline, reviewed as carefully as latency graphs.
Break the 2 to 3KB into parts you can reason about. Label pairs cost bytes each plus index entries; the current head chunk holds up to 120 samples before sealing; the inverted index maps every label value to its series. Small labels on few series cost almost nothing, which is why method and status never hurt. The same bytes on 250,000 user values multiply into gigabytes because every part scales with series count. Run the multiplication for each new label in review: values times combinations times 3KB. Numbers on a napkin beat OOMKills at 3 AM every time.
Which Labels Explode and Which Stay Safe
Labels split into two families. Bounded labels like method, status_code, and route templates have small fixed value sets — 4 methods, 5 statuses, 50 routes — and their products stay flat for years. They're the reason labels exist and are safe by construction.
Unbounded labels grow with traffic: user IDs, emails, session tokens, request UUIDs, and raw URL paths with embedded IDs. Each new user permanently adds series, so the head grows monotonically until the pod dies. These values belong in logs, traces, or warehouses — systems built for per-entity detail — never in metric labels.
The treacherous middle is labels that look bounded until errors strike: error messages with dynamic text, file paths with temp dirs, pod names in long-lived clusters. Normalize at instrumentation — lowercase, bucket, allowlist — so a new error string can't mint ten thousand series overnight.
Histograms multiply everything they touch. A histogram with 12 buckets and labels for method and status creates 12 times 4 times 5 series per metric name — plus _count and _sum. That is 242 series where a counter would create 20, before any unbounded label arrives. Native histograms compress this dramatically by storing bucket spans instead of indexed per-bucket series, which is why newer clients prefer them. Until you migrate, treat every histogram label as twelve labels in review. The bucket count is a hidden multiplier most teams forget to multiply.
Finding the Offending Label in Minutes
When head series spike, find the owner fast. Rank metrics by series count with a topk query — the offender is usually number one by a wide margin. Then count values of its labels to find the runaway: a payment_method with 4 values is innocent, a user_id with 250,000 values is guilty.
Correlate with scrape sizes. The scrape_samples_scraped series per job shows which job's exposition ballooned and exactly when, pointing at the deploy that added the label. Cross-check against deploy logs and the diff usually confesses within minutes.
Move quickly but measure first. Record the head series count, memory, and the offending label's value growth before changing anything. Those baselines prove the fix worked and size the guardrails you'll add next.
The TSDB status API gives machine-readable ground truth when PromQL feels indirect. GET /api/v1/status/tsdb returns head statistics including total series, chunk counts, and compaction numbers — scrape it into your capacity dashboard. For label-level forensics, /api/v1/label/<name>/values with match[] selectors counts values per label without writing queries. Compare value counts across a deploy boundary: the label whose cardinality jumps names the commit. When topk queries time out on overloaded servers, the status API answers in milliseconds because it reads head metadata, not series data. Add both endpoints to your incident runbook next to the PromQL versions.
Dropping and Aggregating: the Two Production Fixes
The production fix drops the explosive label at scrape time, before it ever reaches the head. metric_relabel_configs rules run after scraping against metric labels, and a labeldrop or regex drop on user_id deletes the dimension from every affected series. The metric survives aggregated; the per-user detail never enters memory.
The classic trap is placing the rule in relabel_configs, which runs before scraping against target labels only. The config reloads cleanly, the rule matches nothing, and cardinality climbs untouched. Always use metric_relabel_configs for metric labels — the name is the documentation.
Dashboards that genuinely need segments get recording rules instead. Sum the counter by plan tier or endpoint group into a new aggregated series, point dashboards at it, and drop the raw high-cardinality originals. You keep every business answer while deleting millions of stored series.
Choose the drop granularity deliberately. Dropping a whole metric with source_labels __name__ suits counters that should never have shipped; labeldrop suits metrics that are valuable without one explosive dimension. The keep action inverts the logic for allowlisting: keep only named metrics from a noisy exporter and drop the hundred you never query. Order matters — rules evaluate top to bottom, so place cheap exact matches before expensive regexes. Test every rule with promtool test rules or a staging reload plus series counts; a miswritten regex can drop production alerts silently. Review the rule file quarterly because exporters add metrics on upgrade.
Guardrails: sample_limit and Alerts That Page Early
sample_limit is the circuit breaker that converts silent growth into a loud, immediate scrape failure. Set per job near 150% of its measured normal scrape size: normal variance passes, a label explosion fails the whole scrape and pages someone before the head fills.
Pair it with two alerts. One fires when any job's scrape size crosses 80% of its limit — the early warning that buys weeks. The other watches head-series week-over-week growth above 20% and memory above 80% of the pod limit. Together they catch both the sudden deploy mistake and the slow unbounded creep.
Treat limit tuning as capacity work. Recompute normal sizes quarterly, review every new label in code review with the 3KB math, and require load-test evidence for metrics on hot paths. Guardrails decay as fleets grow; the review habit is what keeps them honest.
Know the two sibling limits. sample_limit caps samples per scrape for one target and fails that scrape loudly; target_limit caps the total across all targets of a job and spreads protection wider. Set sample_limit near 150 percent of normal per-target size and target_limit near 150 percent of the job total. When a scrape exceeds the limit the entire scrape fails — including healthy series — so pair limits with alerts at 80 percent that fire before the cliff. Document each limit beside the measurement that sized it; unexplained limits get deleted in cleanups and the protection vanishes. Limits are capacity decisions wearing config clothes.
Recovering Without Just Buying Bigger Pods
Recovery from an active explosion follows a fixed order. First drop the label at scrape to stop the bleeding — head series flatten within a couple of compactions. Then point dashboards at aggregated recording rules so nothing visible breaks. Then set sample_limit and growth alerts so it can't recur silently.
Resist the urge to fix retention or resize pods instead. Retention governs disk blocks, and bigger pods just reschedule the crash by the exact ratio of added memory to growth rate. Both feel productive while the series math keeps compounding underneath.
Finally, move the displaced detail somewhere it belongs. Per-user analytics live happily in warehouses fed by event streams; debugging detail lives in logs and traces. Prometheus keeps the aggregates that drive alerts and dashboards — that's the division of labor that survives scale.
Dropping labels affects new samples only — existing high-cardinality series age out through retention, not instantly. Expect head series to fall over hours as old chunks compact away, not seconds after reload. For immediate relief on a crashing server, the snapshot API plus a storage wipe and replay buys a clean head, but prefer patience unless OOMKills loop. Old series with dropped labels remain queryable under their original names until retention expires, which keeps history intact while stopping growth. Verify the decline with head-series graphs across two compaction cycles before declaring victory. The fix works on geological time; watch it like a geologist.
The user_id Label That Built 3.7M Series in 9 Days
- Retention controls disk, not head memory — cutting retention never fixes a cardinality OOMKill.
- Bigger pods only reschedule the crash; the label math grows with users and always catches up.
- Every new label needs a cardinality review in code review, with sample_limit as the backstop when review misses one.
| File | Command / Code | Purpose |
|---|---|---|
| prometheus_tsdb_head_series | Cardinality Math | |
| count by (payment_method) (checkout_completed_total) | Finding the Offending Label in Minutes | |
| prometheus.yml | scrape_configs: | Dropping and Aggregating |
| scrape_samples_scraped / ignoring() group_left | Guardrails |
Key takeaways
Common mistakes to avoid
5 patternsAdding user_id, email, or session labels to Prometheus metrics
Keeping raw high-cardinality series for dashboard convenience
Running scrape jobs with no sample_limit guardrail
Capacity planning on disk retention while ignoring head memory
Putting labeldrop rules in relabel_configs instead of metric_relabel_configs
Interview Questions on This Topic
What is cardinality, and why does one bad label hurt so much?
Frequently Asked Questions
20+ years shipping production backend systems. Notes here come from systems that actually shipped.
That's Prometheus. Mark it forged?
6 min read · try the examples if you haven't