Home › Observability › OTel Missing Spans: Rejoin Fragmented Traces
Intermediate 6 min · September 23, 2026
OpenTelemetry Traces Missing Spans Across Services

OTel Missing Spans: Rejoin Fragmented Traces

OpenTelemetry traces missing spans? Check sampler ratios, W3C headers, client instrumentation, and collector tail policies..

N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Production
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
Before you start⏱ 12 min
  • ✓A system emitting partial traces to a backend you can query
  • ✓Ability to set env vars and restart one service
  • ✓Access to collector config and logs
 ● Production Incident 🔎 Debug Guide
⚡Quick Answer
  • Compare trace IDs across the gap since changed IDs mean broken propagation while same IDs mean sampling or code gaps
  • Print sampler env vars on missing services because a leftover ratio like 0.01 explains absence all by itself
  • Unify propagators on W3C tracecontext plus baggage and verify traceparent headers survive every hop
  • Instrument clients, queue consumers, and executors, then size collector queues and tail policies for errors
✦ Definition~90s read
What is OpenTelemetry Traces Missing Spans Across Services?

A distributed trace links spans from every service handling one request, connected by shared trace IDs passed in headers. The OpenTelemetry SDK creates spans per configured samplers, propagates context via W3C traceparent/tracestate headers (plus baggage), and exports finished spans through batch processors to a collector pipeline that may tail-sample before the backend stores them.

★
Think of a trace as a relay race where each runner passes a baton (trace ID in headers) to the next.

Each layer can remove spans by design or accident. Head samplers (parentbased_always_on, parentbased_traceidratio with OTEL_TRACES_SAMPLER_ARG) drop spans at creation to bound cost. Propagation breaks when any hop strips or mistranslates headers, forking new trace IDs.

Uninstrumented clients, queue consumers, and executors never create or link their spans. Collectors drop under queue pressure or keep only tail-policy matches (errors, slow traces).

Diagnosis compares trace IDs across each gap: changed IDs convict propagation at that hop, identical IDs convict sampling or instrumentation, and load-correlated gaps convict the pipeline. Fixes follow the verdict — pin sampler env, unify propagators, wrap every hop, size queues and tail policies — verified by canary traces that prove completeness continuously rather than per incident.

Plain-English First

Think of a trace as a relay race where each runner passes a baton (trace ID in headers) to the next. Samplers tell some runners to sit out before starting. Broken propagation drops the baton at one handoff, so later runners start their own race. Uninstrumented hops are runners who never signed in. The collector is the finish-line camera that films only policy-selected races. Find where the baton dropped and the missing runners reappear.

The checkout trace shows 4 spans. The system has 9 services. Payment, fraud, and notify are simply absent — no errors, no slow spans, just nothing, while logs prove all three ran. The trace is a jigsaw with half the pieces missing.

Missing spans come from four layers, and each leaves different evidence. Samplers deliberately drop spans before export. Broken propagation starts fresh trace IDs at some hop. Uninstrumented clients and queues orphan downstream work. Collectors drop batches under load or keep only policy-matched traces. Same symptom, four systems — guessing wastes days.

This guide splits the fault with trace-ID comparisons and header checks, then fixes the guilty layer: sampler env vars, W3C propagator unity, client and queue instrumentation, collector queues and tail policies. You'll end with complete traces and the dashboards that keep them complete. Pick one fragmented trace before reading further: comparing its trace IDs across the gap decides the whole diagnosis. The fix for your gap is in one of the four sections below.

Sampler Config: the Deliberate Span Remover

Samplers decide at span creation, and parent-based ones honor the parent's flag — children of kept traces stay, children of dropped traces go. parentbased_always_on keeps everything (honest, expensive). parentbased_traceidratio keeps a configured fraction at the root via OTEL_TRACES_SAMPLER_ARG and honors parents after. The default across SDKs is parentbased_always_on, but frameworks and forgotten load-test env vars override it silently.

Ratio sampling drops whole traces by trace ID, which is why missing spans under a 1% ratio look like broken code: the spans were never created for export, on purpose. Confirm by printing env on the missing service and checking SDK startup logs. The temporary always_on override is the definitive test — spans returning within minutes convict sampling absolutely.

Choose ratios deliberately per environment. Production pairs modest head ratios (0.05-0.2) with tail policies that keep every error; staging keeps everything for development velocity. Whatever you choose, assert it: log the sampler description at startup and alert when production env drifts from the declared value.

Audit sampler settings the way you audit credentials: dump effective env per deployment, diff staging against production in CI, and alert when production values drift from declared ones. Load-test leftovers are the classic drift — a ratio set for a traffic experiment survives long after the experiment ends. Keep a documented default per environment and require review for overrides, just like feature flags. Samplers are configuration with production blast radius; govern them accordingly.

BASH
1
2
3
4
5
6
7
8
9
10
11
# Pin sampling explicitly — never inherit silently
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1
OTEL_PROPAGATORS=tracecontext,baggage

# Debug override: keep everything on one service temporarily
# OTEL_TRACES_SAMPLER=parentbased_always_on

# Assert in startup logs (Python example)
# from opentelemetry.sdk.trace.sampling import _get_from_env_or_default
# print('sampler:', _get_from_env_or_default().get_description())
📊 Production Insight
A 0.01 ratio from a forgotten load test hid payment spans for 6 weeks through two rewrites. Printing env on the service would have ended the incident in the first hour — the answer was in the environment, not the code.
🎯 Key Takeaway
Print and pin sampler env per service; temporary always_on convicts sampling in minutes.

W3C Propagation: Passing the Baton Intact

Cross-service linkage rides HTTP and message headers: W3C traceparent carries version, trace ID, parent span ID, and flags; tracestate carries vendor data; baggage carries business context. OTEL_PROPAGATORS defaults to tracecontext,baggage, and every hop must extract on entry and inject on exit with a compatible format.

Breaks happen at boundaries. Proxies strip unknown headers, hand-rolled HTTP clients skip injection, queue producers forget message attributes, and B3-configured services meet W3C-configured ones with mutual incomprehension. The signature is unmistakable: trace IDs change at exactly one hop, with healthy spans on both sides that refuse to join.

Verify mechanically. Dump outgoing headers and demand a well-formed traceparent (00-<32-hex-trace>-<16-hex-parent>-<2-hex-flags>). Unify propagators everywhere first; add dual-format support only where legacy B3 systems truly remain. Header contract tests in CI — asserting traceparent survives each hop — make this class extinct.

Propagator formats beyond W3C still roam production. B3 single and multi headers power Zipkin-era systems; Jaeger propagation lingers in older meshes; AWS X-Ray uses its own trace header. Mixed fleets configure composite propagators listing every format in use, with W3C first as the canonical one. Each hop extracts in priority order and injects all configured formats, keeping legacy and modern readers happy. Inventory every hop's format support during onboarding — the one service nobody checked is always the one that forks traces. Unity is a project, not a default.

BASH
1
2
3
4
5
6
7
8
9
# What a healthy hop carries (inspect with a header dump)
# traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
#   version 00 | trace-id (32 hex) | parent-id (16 hex) | flags 01=sampled
# tracestate: vendor data; baggage: user-tier=pro

# Python: verify extraction on entry, injection on exit
# from opentelemetry.propagate import extract, inject
# ctx = extract(request.headers)   # entry: must yield remote parent
# inject(response.headers)         # exit: must emit traceparent
📊 Production Insight
An nginx proxy stripped traceparent at one boundary, splitting every trace in two for a month. Both halves looked healthy alone — only a header dump at the hop revealed the split. One proxy config line rejoined them.
🎯 Key Takeaway
Unify tracecontext plus baggage; demand well-formed traceparent at every hop.

Broken Instrumentation: the Unsigned Relay Runners

Instrumented servers with bare clients produce the classic half-trace: inbound spans exist, outbound calls start fresh. HTTP clients, gRPC stubs, queue producers and consumers, scheduled executors, and thread pools each need their wrapper — context doesn't cross these hops by itself, in any language.

Queues deserve special attention. Producers must inject context into message headers or attributes; consumers must extract before starting processing spans. Without both sides, consumers mint new trace IDs and their spans form orphan singletons. Async executors similarly need context-aware wrappers, or background work detaches from its request the moment it hops threads.

Roll out per hop with verification. Add one wrapper, redeploy one service, and confirm its spans attach to a live end-to-end trace before moving on. Auto-instrumentation covers frameworks; manual spans cover the custom client nobody remembers. The checklist for new hops — client, queue, executor — belongs in code review.

Messaging systems need explicit context plumbing. Kafka headers, SQS message attributes, and RabbitMQ headers all carry traceparent when producers inject and consumers extract — nothing propagates automatically. Batch consumers must extract per message, not per batch, or one context bleeds across unrelated traces. Links connect related work (a retry, a fan-out) without parentage when strict trees mislead. Document the inject/extract pair for every queue in the system; queues are the number-one trace fork in most architectures. Treat message headers as load-bearing API.

📊 Production Insight
Notify spans formed thousands of single-span traces for weeks — the queue consumer never extracted context from message headers. One consumer-side extraction call attached them all to their checkout traces overnight.
🎯 Key Takeaway
Wrap clients, queues, and executors — servers alone orphan everything downstream.

Collector Drops and Tail-Sampling Verdicts

Even perfect code loses spans to a sick pipeline. Receivers refuse when queues fill, processors drop under memory pressure, exporters shed batches on timeout, and tail-sampling processors keep only policy-matched traces by design. From the backend these look identical to instrumentation gaps — absence with no errors in app logs.

Read the collector's own telemetry. Receiver-refused and exporter-dropped counters rising with queue size convict capacity; gaps matching tail-policy rules (only errors kept, decision_wait too short for slow spans) convict policy. Batch, retry with backoff, separate pipelines per signal, and size memory_limiter and queue capacities from load tests — not defaults.

Tail sampling needs deliberate tuning. Policies keep error, high-latency, and probabilistic slices; decision_wait buffers long enough for slow spans to arrive (size from p99 arrival delay); num_traces bounds memory. Too-short waits amputate the slow traces you most wanted; unbounded buffers OOM the collector. Tune from data, verify with canary error traces.

Auto-instrumentation covers frameworks fast and custom code never. Java agents, Python auto-instrumentation, and .NET packages wrap HTTP, RPC, and database calls with zero code changes — enable them first for baseline coverage. Manual spans handle the business logic in between: name operations clearly, attach domain attributes, and propagate context through custom executors. Review coverage with canary traces after every deploy, since dependency upgrades silently unwire bytecode hooks. Automation plus intention beats either alone.

collector.ymlYAML
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
# Collector: tail sampling keeping errors + slow traces, with headroom
processors:
  batch: {}
  memory_limiter:
    check_interval: 1s
    limit_mib: 512
  tail_sampling:
    decision_wait: 10s
    num_traces: 50000
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow
        type: latency
        latency: { threshold_ms: 2000 }
      - name: probabilistic
        type: probabilistic
        probabilistic: { sampling_percentage: 10 }
📊 Production Insight
A team debugged code for weeks while collector logs showed queue-full drops at peak. Doubling queue capacity and adding retry ended the gaps in a day — the instrumentation had been innocent throughout.
🎯 Key Takeaway
Graph refused/dropped counters; tune decision_wait from p99 arrival delay.

Async, Batch, and Shutdown: Spans Lost at the Finish

Async and batch paths hide the subtlest gaps. In-process spans buffered at shutdown vanish without a graceful flush — short-lived jobs and serverless functions lose their final spans routinely. Batch exporters with full queues drop silently under defaults. Both look like missing spans with healthy code and headers.

Force the flush. Hook SDK shutdown to request drain, extend batch timeouts for job workloads, and use synchronous export in tests to separate buffering gaps from real ones. For serverless, flush explicitly at handler end — the frozen execution environment never runs your atexit hooks.

Clock skew across hosts complicates the picture without removing spans: spans arrive with timestamps that confuse backend assembly and look absent in time-scoped views. NTP health plus backend span-count (not just trace-shape) checks separate skew display artifacts from genuine loss.

Exporter tuning decides whether spans survive bursts. Batch settings (schedule delay, max batch size, queue capacity) trade latency for throughput — bigger batches cost seconds of delay but survive traffic spikes. OTLP exporters retry with backoff on 429 and 503; classic mistakes disable retries for speed and lose spans at the first hiccup. Separate pipelines per signal so a log flood cannot starve trace exports. Load-test the pipeline at 2x peak before the holiday traffic does it for you. Buffers are cheap; lost spans are forever.

PYTHON
1
2
3
4
5
# Flush explicitly at shutdown so buffered spans export
provider.force_flush(timeout_millis=30000)

# Serverless: call force_flush() at the end of every handler invocation
# Tests: use SimpleSpanProcessor to rule out buffering gaps
📊 Production Insight
Serverless checkout functions lost final spans on every cold invocation until an explicit handler-end flush. The traces looked service-skipping; the spans had simply never left the frozen runtime.
🎯 Key Takeaway
Flush explicitly at shutdown; sync-export in tests to split buffering gaps from real loss.

Keeping Traces Complete as You Scale

Lock completeness with continuous proof. Dashboard kept-versus-dropped span ratios per service, tail-policy hit rates, collector queue depths, and end-to-end canary traces that page when any hop goes dark. Canaries catch configuration drift — the expired sampler env, the proxy that started stripping headers — before users' traces fragment.

Review sampling quarterly against traffic growth: ratios that fit last year starve debugging this year. Keep the error-keeping tail policies versioned with the collector config, and require propagation tests for every new hop in code review.

Complete traces compound. Once every hop propagates and every error is kept, debugging shifts from archaeology to reading — the trace tells the story the first time. That payoff funds all the unglamorous header checks and env audits that built it.

Sampling policy is a living document, not a deploy-and-forget number. Review head ratios quarterly against traffic growth: last year's 10 percent is this year's budget fire. Tail policies gain rules as new failure modes appear — a new error code deserves a keep-rule the same sprint it ships. Version tail configs with the collector deployment and canary-test rule changes with synthetic error traces. Publish kept-versus-dropped ratios per service so teams see their visibility budget. Sampling governance is observability governance.

⚠ Evidence Before Rewrites
Never debug missing spans by rewriting instrumentation first. Compare trace IDs, print sampler env, dump headers, and graph collector drops — in that order. Two rewrites of working code cost one team 3 weeks; the env answer would have taken an hour.
📊 Production Insight
After canary traces plus sampler-env audits, one org held 100% error-trace completeness for a year across 40 services. The two fragmentations in that span paged via canary within minutes — both were proxy header strips caught the same day.
🎯 Key Takeaway
Canary traces plus kept/dropped dashboards catch drift before users notice fragments.
● Production incidentPOST-MORTEMseverity: high

The 1% Sampler Left Over From a Load Test

Symptom
Payment spans appeared in 1% of checkout traces for 6 weeks with zero errors logged. Two instrumentation rewrites passed staging and changed nothing in production; notify spans formed separate single-span traces throughout.
Assumption
The team assumed the payment service's instrumentation was broken and rewrote it twice over 3 weeks. Both rewrites traced perfectly in staging — because staging set parentbased_always_on while production carried a forgotten 1% ratio from a load test months earlier.
Root cause
Production carried OTEL_TRACES_SAMPLER_ARG=0.01 from a months-old load test, sampling away 99% of payment spans, while staging used parentbased_always_on — so rewrites always looked fixed until deploy. A missing consumer-side extraction on the notify queue independently orphaned notify spans with fresh trace IDs.
Fix
They set OTEL_TRACES_SAMPLER=parentbased_traceidratio with ARG 0.1 in production, added a tail-sampling policy keeping all error and over-2s traces, unified propagators to tracecontext,baggage, and instrumented the queue consumer that had also been orphaning notify spans. Payment visibility went from 1% to 100% of errors overnight.
Key lesson
  • Sampler env from old load tests outlives every code fix — audit env before rewriting instrumentation.
  • Staging-production sampler parity belongs in deploy checks, not memory.
  • Tail sampling for errors plus modest head ratios beats both extremes: full cost or blind debugging.
Production debug guideFive comparisons that convict sampler, headers, code, or pipeline.5 entries
Symptom · 01
Traces show some services but skip others
→
Fix
Pick one fragmented trace and compare trace IDs on spans each side of the gap. Different IDs mean propagation broke at that hop — go to headers. Same IDs with absent spans mean sampling or instrumentation dropped them — check sampler env on the missing services first.
Symptom · 02
Same trace ID but spans absent
→
Fix
Print OTEL_TRACES_SAMPLER and OTEL_TRACES_SAMPLER_ARG on every service in the path and read SDK startup logs. A 0.01 ratio on the missing service explains absence completely. Set parentbased_always_on temporarily on one service and watch its spans return within minutes to convict sampling.
Symptom · 03
Trace IDs change at one boundary
→
Fix
Curl the hop or dump outgoing message headers and look for traceparent (format 00-<trace>-<parent>-<flags>) plus tracestate. Absent headers mean no injection (uninstrumented client); present-but-different IDs downstream mean no extraction. Fix by adding the matching instrumentation and unifying OTEL_PROPAGATORS to tracecontext,baggage everywhere.
Symptom · 04
Headers flow correctly but child spans never appear
→
Fix
Verify the hop's client library, queue consumer, and async executor are wrapped: HTTP clients need client spans with injection, consumers need extraction from message headers, executors need context propagation. Add the missing wrapper, redeploy one service, and confirm its spans attach to the live trace before rolling out.
Symptom · 05
Code looks right but gaps persist under load
→
Fix
Graph receiver-refused and exporter-dropped spans plus queue sizes around gap times, and read tail-sampling policies for decision_wait and match rules. Drops under load mean undersized queues or batching; policy-kept gaps mean rules excluding your case. Fix pipeline first — no code change survives a dropping collector.
Missing-Span Causes Compared
Root CauseHow to ConfirmFixPrevention
Head sampler dropping spansSampler env is ratio-based; kept spans completeparentbased_always_on for debug; tail rules afterPin sampler env; log it at startup
Broken context propagationTrace IDs change at one hop; headers absentUnify W3C propagators; inject/extract at hopHeader contract tests; propagate through queues
Uninstrumented hop in pathGap between two good spans; client lib bareAdd client/queue/executor instrumentationInstrumentation checklist per new hop
Collector drop or tail-sampleReceiver-refused/exporter-dropped counters riseSize queues, batch, retry; fix tail policiesPipeline drop dashboards; load-test collector
⚙ Quick Reference
3 commands from this guide
FileCommand / CodePurpose
OTEL_TRACES_SAMPLER=parentbased_traceidratioSampler Config
collector.ymlprocessors:Collector Drops and Tail-Sampling Verdicts
provider.force_flush(timeout_millis=30000)Async, Batch, and Shutdown

Key takeaways

1
Compare trace IDs across the gap
changed means propagation broke, same means instrumentation is missing.
2
Pin OTEL_TRACES_SAMPLER explicitly per service and log it at startup.
3
Unify on W3C tracecontext plus baggage; verify headers survive every proxy and queue.
4
Instrument clients, queues, and executors
servers alone orphan downstream spans.
5
Size collector queues and tail policies; alert on refused and dropped spans.
6
Pair low head ratios with tail rules so errors stay debuggable.

Common mistakes to avoid

5 patterns
×

Assuming the default sampler everywhere without setting it

Symptom
One service's framework default differs, and its spans vanish while identical code elsewhere traces fine.
Fix
Set OTEL_TRACES_SAMPLER explicitly on every service and assert it in startup logs. Treat sampler env as required config, reviewed like database URLs — never inherited silently.
×

Mixing propagator formats across services

Symptom
Traces break at exactly one service boundary — upstream and downstream each look healthy alone.
Fix
Use one propagator set (W3C tracecontext plus baggage) across all services and verify headers at each hop. Mixed B3 and W3C services need explicit dual-propagator config, not hope.
×

Instrumenting servers but not HTTP clients or queues

Symptom
Inbound spans exist but outbound calls start fresh traces, scattering one request across dozens of fragments.
Fix
Wrap every hop — HTTP clients and servers, queue producers and consumers, async executors — with instrumentation. Unwrapped hops break parent context and orphan everything downstream.
×

Setting head sampling to 1% and expecting complete error traces

Symptom
Errors vanish 99% of the time, and the 1% kept rarely includes the failing span — debugging from traces becomes impossible.
Fix
Size head sampling for steady state and use tail sampling for errors and latency. Debug completeness comes from tail rules, not from sampling everything at head.
×

Blaming code for spans the collector dropped

Symptom
Weeks of instrumentation debugging while the collector logs queue-full drops that nobody graphed.
Fix
Monitor receiver-refused and exporter-dropped span counts with queue-size alerts. Pipeline drops look exactly like instrumentation gaps from the outside.
INTERVIEW PREP · PRACTICE MODE

Interview Questions on This Topic

Q01JUNIOR
How do parentbased samplers decide what to keep?
Q02JUNIOR
How does W3C propagation link spans across services?
Q03SENIOR
Spans stop at one service boundary. How do you split propagation vs inst...
Q04SENIOR
What do decision_wait and num_traces control in tail sampling?
Q05SENIOR
Why do collectors drop spans under load, and how do you stop it?
Q01 of 05JUNIOR

How do parentbased samplers decide what to keep?

ANSWER
parentbased_always_on keeps every span and honors parents; parentbased_traceidratio keeps a configured fraction (OTEL_TRACES_SAMPLER_ARG like 0.1) at the root and honors parents after. Ratio sampling drops whole traces at random — missing spans under ratio config are usually sampled away, not lost.
FAQ · 6 QUESTIONS

Frequently Asked Questions

01
Head vs tail sampling — which keeps errors?
02
Which headers carry trace context?
03
Do child spans follow the parent's sampling decision?
04
Can the backend recover unsampled spans?
05
Why do traces stop at message queues?
06
Same trace ID or different — why does it matter?
N
Naren Founder & Principal Engineer

20+ years shipping production backend systems. Everything here is grounded in real deployments.

Follow
✓ Verified
production tested
September 26, 2026
last updated
2,085
articles · all by Naren
🔥

That's OpenTelemetry. Mark it forged?

6 min read · try the examples if you haven't

←
Previous
Elasticsearch Mapper Parsing Exception on Field Type Conflict
1 / 1 · OpenTelemetry