OTel Missing Spans: Rejoin Fragmented Traces
OpenTelemetry traces missing spans? Check sampler ratios, W3C headers, client instrumentation, and collector tail policies..
20+ years shipping production backend systems. Everything here is grounded in real deployments.
- ✓A system emitting partial traces to a backend you can query
- ✓Ability to set env vars and restart one service
- ✓Access to collector config and logs
- Compare trace IDs across the gap since changed IDs mean broken propagation while same IDs mean sampling or code gaps
- Print sampler env vars on missing services because a leftover ratio like 0.01 explains absence all by itself
- Unify propagators on W3C tracecontext plus baggage and verify traceparent headers survive every hop
- Instrument clients, queue consumers, and executors, then size collector queues and tail policies for errors
Think of a trace as a relay race where each runner passes a baton (trace ID in headers) to the next. Samplers tell some runners to sit out before starting. Broken propagation drops the baton at one handoff, so later runners start their own race. Uninstrumented hops are runners who never signed in. The collector is the finish-line camera that films only policy-selected races. Find where the baton dropped and the missing runners reappear.
The checkout trace shows 4 spans. The system has 9 services. Payment, fraud, and notify are simply absent — no errors, no slow spans, just nothing, while logs prove all three ran. The trace is a jigsaw with half the pieces missing.
Missing spans come from four layers, and each leaves different evidence. Samplers deliberately drop spans before export. Broken propagation starts fresh trace IDs at some hop. Uninstrumented clients and queues orphan downstream work. Collectors drop batches under load or keep only policy-matched traces. Same symptom, four systems — guessing wastes days.
This guide splits the fault with trace-ID comparisons and header checks, then fixes the guilty layer: sampler env vars, W3C propagator unity, client and queue instrumentation, collector queues and tail policies. You'll end with complete traces and the dashboards that keep them complete. Pick one fragmented trace before reading further: comparing its trace IDs across the gap decides the whole diagnosis. The fix for your gap is in one of the four sections below.
Sampler Config: the Deliberate Span Remover
Samplers decide at span creation, and parent-based ones honor the parent's flag — children of kept traces stay, children of dropped traces go. parentbased_always_on keeps everything (honest, expensive). parentbased_traceidratio keeps a configured fraction at the root via OTEL_TRACES_SAMPLER_ARG and honors parents after. The default across SDKs is parentbased_always_on, but frameworks and forgotten load-test env vars override it silently.
Ratio sampling drops whole traces by trace ID, which is why missing spans under a 1% ratio look like broken code: the spans were never created for export, on purpose. Confirm by printing env on the missing service and checking SDK startup logs. The temporary always_on override is the definitive test — spans returning within minutes convict sampling absolutely.
Choose ratios deliberately per environment. Production pairs modest head ratios (0.05-0.2) with tail policies that keep every error; staging keeps everything for development velocity. Whatever you choose, assert it: log the sampler description at startup and alert when production env drifts from the declared value.
Audit sampler settings the way you audit credentials: dump effective env per deployment, diff staging against production in CI, and alert when production values drift from declared ones. Load-test leftovers are the classic drift — a ratio set for a traffic experiment survives long after the experiment ends. Keep a documented default per environment and require review for overrides, just like feature flags. Samplers are configuration with production blast radius; govern them accordingly.
W3C Propagation: Passing the Baton Intact
Cross-service linkage rides HTTP and message headers: W3C traceparent carries version, trace ID, parent span ID, and flags; tracestate carries vendor data; baggage carries business context. OTEL_PROPAGATORS defaults to tracecontext,baggage, and every hop must extract on entry and inject on exit with a compatible format.
Breaks happen at boundaries. Proxies strip unknown headers, hand-rolled HTTP clients skip injection, queue producers forget message attributes, and B3-configured services meet W3C-configured ones with mutual incomprehension. The signature is unmistakable: trace IDs change at exactly one hop, with healthy spans on both sides that refuse to join.
Verify mechanically. Dump outgoing headers and demand a well-formed traceparent (00-<32-hex-trace>-<16-hex-parent>-<2-hex-flags>). Unify propagators everywhere first; add dual-format support only where legacy B3 systems truly remain. Header contract tests in CI — asserting traceparent survives each hop — make this class extinct.
Propagator formats beyond W3C still roam production. B3 single and multi headers power Zipkin-era systems; Jaeger propagation lingers in older meshes; AWS X-Ray uses its own trace header. Mixed fleets configure composite propagators listing every format in use, with W3C first as the canonical one. Each hop extracts in priority order and injects all configured formats, keeping legacy and modern readers happy. Inventory every hop's format support during onboarding — the one service nobody checked is always the one that forks traces. Unity is a project, not a default.
Broken Instrumentation: the Unsigned Relay Runners
Instrumented servers with bare clients produce the classic half-trace: inbound spans exist, outbound calls start fresh. HTTP clients, gRPC stubs, queue producers and consumers, scheduled executors, and thread pools each need their wrapper — context doesn't cross these hops by itself, in any language.
Queues deserve special attention. Producers must inject context into message headers or attributes; consumers must extract before starting processing spans. Without both sides, consumers mint new trace IDs and their spans form orphan singletons. Async executors similarly need context-aware wrappers, or background work detaches from its request the moment it hops threads.
Roll out per hop with verification. Add one wrapper, redeploy one service, and confirm its spans attach to a live end-to-end trace before moving on. Auto-instrumentation covers frameworks; manual spans cover the custom client nobody remembers. The checklist for new hops — client, queue, executor — belongs in code review.
Messaging systems need explicit context plumbing. Kafka headers, SQS message attributes, and RabbitMQ headers all carry traceparent when producers inject and consumers extract — nothing propagates automatically. Batch consumers must extract per message, not per batch, or one context bleeds across unrelated traces. Links connect related work (a retry, a fan-out) without parentage when strict trees mislead. Document the inject/extract pair for every queue in the system; queues are the number-one trace fork in most architectures. Treat message headers as load-bearing API.
Collector Drops and Tail-Sampling Verdicts
Even perfect code loses spans to a sick pipeline. Receivers refuse when queues fill, processors drop under memory pressure, exporters shed batches on timeout, and tail-sampling processors keep only policy-matched traces by design. From the backend these look identical to instrumentation gaps — absence with no errors in app logs.
Read the collector's own telemetry. Receiver-refused and exporter-dropped counters rising with queue size convict capacity; gaps matching tail-policy rules (only errors kept, decision_wait too short for slow spans) convict policy. Batch, retry with backoff, separate pipelines per signal, and size memory_limiter and queue capacities from load tests — not defaults.
Tail sampling needs deliberate tuning. Policies keep error, high-latency, and probabilistic slices; decision_wait buffers long enough for slow spans to arrive (size from p99 arrival delay); num_traces bounds memory. Too-short waits amputate the slow traces you most wanted; unbounded buffers OOM the collector. Tune from data, verify with canary error traces.
Auto-instrumentation covers frameworks fast and custom code never. Java agents, Python auto-instrumentation, and .NET packages wrap HTTP, RPC, and database calls with zero code changes — enable them first for baseline coverage. Manual spans handle the business logic in between: name operations clearly, attach domain attributes, and propagate context through custom executors. Review coverage with canary traces after every deploy, since dependency upgrades silently unwire bytecode hooks. Automation plus intention beats either alone.
Async, Batch, and Shutdown: Spans Lost at the Finish
Async and batch paths hide the subtlest gaps. In-process spans buffered at shutdown vanish without a graceful flush — short-lived jobs and serverless functions lose their final spans routinely. Batch exporters with full queues drop silently under defaults. Both look like missing spans with healthy code and headers.
Force the flush. Hook SDK shutdown to request drain, extend batch timeouts for job workloads, and use synchronous export in tests to separate buffering gaps from real ones. For serverless, flush explicitly at handler end — the frozen execution environment never runs your atexit hooks.
Clock skew across hosts complicates the picture without removing spans: spans arrive with timestamps that confuse backend assembly and look absent in time-scoped views. NTP health plus backend span-count (not just trace-shape) checks separate skew display artifacts from genuine loss.
Exporter tuning decides whether spans survive bursts. Batch settings (schedule delay, max batch size, queue capacity) trade latency for throughput — bigger batches cost seconds of delay but survive traffic spikes. OTLP exporters retry with backoff on 429 and 503; classic mistakes disable retries for speed and lose spans at the first hiccup. Separate pipelines per signal so a log flood cannot starve trace exports. Load-test the pipeline at 2x peak before the holiday traffic does it for you. Buffers are cheap; lost spans are forever.
Keeping Traces Complete as You Scale
Lock completeness with continuous proof. Dashboard kept-versus-dropped span ratios per service, tail-policy hit rates, collector queue depths, and end-to-end canary traces that page when any hop goes dark. Canaries catch configuration drift — the expired sampler env, the proxy that started stripping headers — before users' traces fragment.
Review sampling quarterly against traffic growth: ratios that fit last year starve debugging this year. Keep the error-keeping tail policies versioned with the collector config, and require propagation tests for every new hop in code review.
Complete traces compound. Once every hop propagates and every error is kept, debugging shifts from archaeology to reading — the trace tells the story the first time. That payoff funds all the unglamorous header checks and env audits that built it.
Sampling policy is a living document, not a deploy-and-forget number. Review head ratios quarterly against traffic growth: last year's 10 percent is this year's budget fire. Tail policies gain rules as new failure modes appear — a new error code deserves a keep-rule the same sprint it ships. Version tail configs with the collector deployment and canary-test rule changes with synthetic error traces. Publish kept-versus-dropped ratios per service so teams see their visibility budget. Sampling governance is observability governance.
The 1% Sampler Left Over From a Load Test
- Sampler env from old load tests outlives every code fix — audit env before rewriting instrumentation.
- Staging-production sampler parity belongs in deploy checks, not memory.
- Tail sampling for errors plus modest head ratios beats both extremes: full cost or blind debugging.
| File | Command / Code | Purpose |
|---|---|---|
| OTEL_TRACES_SAMPLER=parentbased_traceidratio | Sampler Config | |
| collector.yml | processors: | Collector Drops and Tail-Sampling Verdicts |
| provider.force_flush(timeout_millis=30000) | Async, Batch, and Shutdown |
Key takeaways
Common mistakes to avoid
5 patternsAssuming the default sampler everywhere without setting it
Mixing propagator formats across services
Instrumenting servers but not HTTP clients or queues
Setting head sampling to 1% and expecting complete error traces
Blaming code for spans the collector dropped
Interview Questions on This Topic
How do parentbased samplers decide what to keep?
Frequently Asked Questions
20+ years shipping production backend systems. Everything here is grounded in real deployments.
That's OpenTelemetry. Mark it forged?
6 min read · try the examples if you haven't