This complete guide covers all 10 Observability tutorials on TheCodeForge, organised by topic.
Monitoring answers questions you thought of in advance. Observability is the property of being able to answer questions you did not — and the difference shows up at three in the morning, when the dashboard is green and customers are complaining. Everything in this track exists to make the second kind of question answerable: high-cardinality labels so you can slice by customer, traces so you can follow one request across seven services, structured logs so you can filter rather than grep.
The recurring failure is that each of those has a cost, and the cost is paid in the storage engine. A label with a million values does not make Prometheus slow gradually; it makes it fall over. An Elasticsearch field mapped as text in one index and keyword in another does not warn you; it rejects the document. Understanding the storage model is what separates an observability stack that helps from one that becomes its own incident.
Prometheus stores one time series per unique combination of metric name and label values. http_requests_total with labels for method, status and route across 5 methods, 8 statuses and 40 routes is 1,600 series — entirely fine. Add a user_id label and it becomes 1,600 series per user, which is how an instrumentation change takes down a monitoring cluster in an afternoon.
The rule is mechanical: a label's value set must be small and bounded, and you must be able to say what bounds it. User IDs, request IDs, email addresses, full URLs with query strings, raw error messages and timestamps are all unbounded. They belong in logs or trace attributes, where the storage model is built for them.
| Signal | Storage model | Right for |
|---|---|---|
| Metrics (Prometheus) | One series per label combination, sampled over time | Bounded dimensions and aggregate questions: rate, error ratio, latency distribution per route |
| Logs (Elastic, Loki) | One document or line per event, indexed or scanned | Unbounded detail about a specific event — which user, which id, which message |
| Traces (OpenTelemetry) | A tree of spans per request, usually sampled | Causality and latency attribution across services |
Both operate on counters and both need a range. rate() returns a per-second average over the window; increase() returns the total growth across it. Mixing them up produces thresholds that are wrong by exactly the window length, which is how an alert set to fire at 5 errors instead fires at 300.
The subtler trap is the range itself. A range shorter than roughly four scrape intervals may contain too few samples to extrapolate from, producing gaps and flapping alerts. A range much longer than your alert's for duration smooths away the spike you wanted to catch. As a rule, make the range at least four times the scrape interval, and keep it consistent across an alert family so the numbers stay comparable.
Mapper parsing exception means a document's field conflicts with the type already established for that field in the index. Elasticsearch infers mappings from the first document it sees, so a field that arrives as a number on Monday and a string on Tuesday is rejected on Tuesday — and because mappings are per-index, a rolling index can behave differently day to day.
Unassigned shards and red cluster health are a different layer: the cluster cannot place a shard copy. Disk watermarks, no eligible node satisfying allocation rules, or a lost primary with no replica. The allocation explain API states the reason rather than leaving you to infer it.
dynamic: strict turns a silent schema drift into a loud rejection at ingest, which is far cheaper than discovering six months later that half your status values are strings and cannot be aggregated.A trace that breaks between services is nearly always a context propagation failure rather than an instrumentation gap. The traceparent header must survive every hop — through your HTTP client, through a message queue, through a thread pool or async boundary where the ambient context does not follow automatically, and through any proxy configured to strip unknown headers.
Sampling is the second cause, and it produces a confusing symptom: the trace exists but is incomplete, because one service made an independent sampling decision. Head-based sampling has to be consistent across the whole call graph, which is what the parentbased_traceidratio sampler is for — children inherit the root's decision instead of each rolling their own.
When a queue sits in the middle, propagation is your responsibility: inject the context into the message headers on publish and extract it on consume. No auto-instrumentation can do that for you, because only your code knows where the message envelope is.
Every tutorial starts with a plain-English analogy — then real code, then interview questions.
Browse Observability Tutorials →