Home› Observability› Complete Guide
Complete Guide

Complete Observability Tutorial

This complete guide covers all 10 Observability tutorials on TheCodeForge, organised by topic.

Learning Roadmap
Beginner → Build a strong foundation
Intermediate → Deepen your understanding with practical topics
Advanced → Master advanced concepts and real-world applications
10
Topics
4
Beginner
4
Intermediate
2
Advanced
Jump to section
Prometheus (4)Grafana (2)Elasticsearch (3)OpenTelemetry (1)

Monitoring answers questions you thought of in advance. Observability is the property of being able to answer questions you did not — and the difference shows up at three in the morning, when the dashboard is green and customers are complaining. Everything in this track exists to make the second kind of question answerable: high-cardinality labels so you can slice by customer, traces so you can follow one request across seven services, structured logs so you can filter rather than grep.

The recurring failure is that each of those has a cost, and the cost is paid in the storage engine. A label with a million values does not make Prometheus slow gradually; it makes it fall over. An Elasticsearch field mapped as text in one index and keyword in another does not warn you; it rejects the document. Understanding the storage model is what separates an observability stack that helps from one that becomes its own incident.

Cardinality is the only Prometheus concept that really matters

Prometheus stores one time series per unique combination of metric name and label values. http_requests_total with labels for method, status and route across 5 methods, 8 statuses and 40 routes is 1,600 series — entirely fine. Add a user_id label and it becomes 1,600 series per user, which is how an instrumentation change takes down a monitoring cluster in an afternoon.

The rule is mechanical: a label's value set must be small and bounded, and you must be able to say what bounds it. User IDs, request IDs, email addresses, full URLs with query strings, raw error messages and timestamps are all unbounded. They belong in logs or trace attributes, where the storage model is built for them.

SignalStorage modelRight for
Metrics (Prometheus)One series per label combination, sampled over timeBounded dimensions and aggregate questions: rate, error ratio, latency distribution per route
Logs (Elastic, Loki)One document or line per event, indexed or scannedUnbounded detail about a specific event — which user, which id, which message
Traces (OpenTelemetry)A tree of spans per request, usually sampledCausality and latency attribution across services
promql
# Find the label that is costing you
topk(10, count by (__name__)({__name__=~".+"}))

# Which label values are exploding on one metric?
count(count by (route) (http_requests_total))

# A rule worth alerting on before it becomes an outage
prometheus_tsdb_symbol_table_size_bytes > 200e6

rate() versus increase(), and why alerts get this wrong

Both operate on counters and both need a range. rate() returns a per-second average over the window; increase() returns the total growth across it. Mixing them up produces thresholds that are wrong by exactly the window length, which is how an alert set to fire at 5 errors instead fires at 300.

The subtler trap is the range itself. A range shorter than roughly four scrape intervals may contain too few samples to extrapolate from, producing gaps and flapping alerts. A range much longer than your alert's for duration smooths away the spike you wanted to catch. As a rule, make the range at least four times the scrape interval, and keep it consistent across an alert family so the numbers stay comparable.

promql
# Per-second error rate - compare against a rate, never against a count
rate(http_requests_total{status=~"5.."}[5m])

# Errors in the last five minutes - compare against a count
increase(http_requests_total{status=~"5.."}[5m])

# The alert you almost always actually want: a ratio, not an absolute
sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m])) > 0.05

# Always rate() BEFORE sum(). sum() first destroys the per-series
# counter resets that rate() needs in order to be correct.
In practiceA query that returns no data while the target is up is usually a label mismatch, not a broken exporter. Strip the selector back to the bare metric name and add labels one at a time — relabelling rules, a changed job name, or an environment label added by the scrape config are the common culprits.

Elasticsearch: mappings are a schema you cannot change casually

Mapper parsing exception means a document's field conflicts with the type already established for that field in the index. Elasticsearch infers mappings from the first document it sees, so a field that arrives as a number on Monday and a string on Tuesday is rejected on Tuesday — and because mappings are per-index, a rolling index can behave differently day to day.

Unassigned shards and red cluster health are a different layer: the cluster cannot place a shard copy. Disk watermarks, no eligible node satisfying allocation rules, or a lost primary with no replica. The allocation explain API states the reason rather than leaving you to infer it.

bash
# Why is this shard unassigned? The API answers in plain language.
GET _cluster/allocation/explain
{ "index": "logs-2026.09.26", "shard": 0, "primary": true }

# The three causes that account for most red clusters
GET _cat/allocation?v                 # disk watermarks per node
GET _cat/shards/logs-*?v&h=index,shard,prirep,state,unassigned.reason
GET _cluster/settings?include_defaults=true&filter_path=**.disk.watermark*

# Prevent mapping conflicts instead of reacting to them:
# define an explicit template before the first document arrives
PUT _index_template/logs
{
  "index_patterns": ["logs-*"],
  "template": { "mappings": {
    "dynamic": "strict",
    "properties": {
      "message":    { "type": "text" },
      "status":     { "type": "integer" },
      "trace_id":   { "type": "keyword" }
    }}}
}
In practicedynamic: strict turns a silent schema drift into a loud rejection at ingest, which is far cheaper than discovering six months later that half your status values are strings and cannot be aggregated.

Missing spans: context propagation is the whole trick

A trace that breaks between services is nearly always a context propagation failure rather than an instrumentation gap. The traceparent header must survive every hop — through your HTTP client, through a message queue, through a thread pool or async boundary where the ambient context does not follow automatically, and through any proxy configured to strip unknown headers.

Sampling is the second cause, and it produces a confusing symptom: the trace exists but is incomplete, because one service made an independent sampling decision. Head-based sampling has to be consistent across the whole call graph, which is what the parentbased_traceidratio sampler is for — children inherit the root's decision instead of each rolling their own.

When a queue sits in the middle, propagation is your responsibility: inject the context into the message headers on publish and extract it on consume. No auto-instrumentation can do that for you, because only your code knows where the message envelope is.

Frequently Asked Questions

What counts as high cardinality in practice?
Any label whose value set you cannot bound with a sentence. Route templates are bounded — you have a finite number of routes. Raw paths with IDs in them are not. A useful test: if adding a customer would add new series, the label is unbounded and belongs in logs or span attributes, not in a metric.
Why does Grafana show 'no data' when the same query works in Prometheus?
Usually the dashboard's time range or its variable interpolation, not the query. Check that the panel points at the datasource you tested against, that template variables resolved to real values rather than an empty string, and that the range is long enough to contain samples — a range shorter than the scrape interval can legitimately return nothing. The panel inspector shows the fully interpolated query, which settles it immediately.
Should I alert on absolute error counts or error ratios?
Ratios, almost always. An absolute threshold that is right at peak traffic is far too noisy at 3am, and one that is right at 3am never fires during the incident you care about. Ratios are traffic-independent. Keep an absolute floor alongside it so a ratio computed from three requests does not page anyone.
What causes out-of-order sample ingestion errors?
A series receiving a sample with a timestamp earlier than one already stored. Duplicate scrape jobs targeting the same instance, two agents with the same external labels, clock skew between pushers, or a backfill writing into the current head block. Find the duplicate first — it is more often two things scraping one target than a clock problem.
Is OpenTelemetry ready to replace vendor agents?
For traces and metrics, yes in most stacks — the APIs are stable and the collector gives you vendor-neutral routing, which is worth a lot on its own. Logs are the least mature of the three signals, so a common shape is OTel for traces and metrics with an existing log pipeline left in place until the log support catches up in your language.
How much of the three-signal stack does a small team actually need?
Start with metrics on the four golden signals — latency, traffic, errors, saturation — and structured logs with a request id in every line. That combination answers most incidents. Add tracing when you have enough services that 'which hop was slow' has stopped being obvious, which is usually somewhere past three or four.

Prometheus

Grafana

Elasticsearch

OpenTelemetry

Also Explore
DevOps 304 tutorials → System Design 145 tutorials → Data Engineering 13 tutorials → Security 16 tutorials → Cloud 12 tutorials → CS Fundamentals 76 tutorials →
Start from the beginning

Every tutorial starts with a plain-English analogy — then real code, then interview questions.

Browse Observability Tutorials →