Skip to content

OTEL Agent Sampling Strategy

Author: Song Longqi

Preface

In distributed service trace tracking, you can observe how a request moves from one service to another within a distributed system. This is essential, but in many cases the data is highly repetitive. Is this data really that important? At such times, using a correct sampling strategy can reduce unnecessary traffic costs.

Sampling refers to accepting data that is ready to be exported or stored. Sampling is sometimes misunderstood as discarding data, which is incorrect.

The previous article OpenTelemetry Sampling Best Practices mainly introduced the sampling strategies of opentelemetry-collector. This article focuses on sampling on the Java Agent side.

Sampling Strategies

Head Sampling

The most common form of head sampling is consistent probability sampling. It can also be called deterministic sampling. In this case, the sampling decision is made based on the TraceID and the desired percentage of sampled traces. This ensures that entire traces are sampled. For example, if 5% of all traces are sampled, then all traces generated will be sampled at a 5% rate.

Advantages of head sampling:

  • Simple configuration
  • Easy to understand
  • High efficiency
  • Can be configured anywhere

The disadvantages are also obvious: it is impossible to make sampling decisions based on data from the entire trace. When an error occurs in a trace, it cannot guarantee that the trace will be sampled and uploaded, which is unfavorable for issue troubleshooting.

Therefore, tail sampling is necessary.

Tail Sampling

Tail sampling refers to the case where the decision to sample a trace is made by considering all or most of the spans in the trace. "Tail sampling" allows you to sample traces based on specific criteria originating from different parts of the trace, which is not possible with "head sampling".

Several scenarios for tail sampling:

  • Sample all traces that contain errors
  • Sample based on overall latency
  • Sample based on specific events
  • Sample when the HTTP status code is not 200
  • Sample based on other criteria

Tail sampling has many challenges:

  • Tail sampling can be difficult to implement. Depending on the type of sampling technique available, it is not always a "set it and forget it" scenario. As the system changes, the sampling strategy also changes. For large, complex distributed systems, the rules for implementing the sampling strategy are also bound to be large and complex.
  • Tail sampling can be difficult to operate. The node implementing tail sampling must accept all data, sometimes trace data from dozens of nodes, before performing calculations. Sometimes the calculation speed cannot keep up with the reception volume. Based on these factors, a correct tail sampling node requires significant resources.
  • Currently, only some large vendors may provide the option of tail sampling.

That is roughly the introduction to head sampling and tail sampling.

Finally, applications that generate a large amount of span data should use head sampling, so that the trace channel does not become congested due to overload.

In the OTEL Java Agent, a strategy based on head sampling is adopted. This should be understood in advance.

For support in other languages or Collector configuration, please refer to: Official Documentation


Sampler in the Agent

In OTEL, there are two environment variables for configuring sampling: OTEL_TRACES_SAMPLER and OTEL_TRACES_SAMPLER_ARG:

How to Choose Sampling Configuration

The configuration for OTEL_TRACES_SAMPLER is:

  • "always_on": AlwaysOnSampler, the default configuration value 1.0 (i.e., no sampling, meaning all traces are sampled)
  • "always_off": AlwaysOffSampler, always discards any trace
  • "traceidratio": TraceIdRatioBased, probability sampling based on trace ID
  • "parentbased_always_on": ParentBased (based on AlwaysOnSampler)
  • "parentbased_always_off": ParentBased (based on AlwaysOffSampler)
  • "parentbased_traceidratio": ParentBased (based on TraceIdRatioBased)
  • "parentbased_jaeger_remote": ParentBased (based on JaegerRemoteSampler)
  • "jaeger_remote": JaegerRemoteSampler
  • "xray": AWS X-Ray Centralized Sampling (third party)

OTEL_TRACES_SAMPLER_ARG configuration:

  • 1.0 – default value is 1.0, meaning all traces are sampled
  • [0-1.0] – probability sampling, e.g., 0.25 means a 25% sampling rate
  • The OTEL_TRACES_SAMPLER_ARG configuration only takes effect when traceidratio or parentbased_traceidratio is used.

Common Configuration

OTEL_TRACES_SAMPLER=parentbased_traceidratio (Parent-based Trace ID Ratio): This sampling strategy determines whether a child span should be sampled based on the parent span's Trace ID. When a new request arrives, it checks whether the parent's Trace ID falls within the specified ratio range. If yes, the child span will also be sampled. This strategy ensures correlation between a request and its related operations. It is very useful for distributed tracing systems.

OTEL_TRACES_SAMPLER=traceidratio (Trace ID Ratio): This sampling strategy determines whether to sample a request based on each request's Trace ID. Each request has a unique Trace ID. With this strategy, each request independently decides whether to be sampled based on the specified ratio. This strategy is suitable for scenarios where you do not need to consider correlation between requests, but only care about the sampling rate of each individual request.

In summary, OTEL_TRACES_SAMPLER=parentbased_traceidratio determines the child's sampling rate based on the parent Trace ID, while OTEL_TRACES_SAMPLER=traceidratio independently determines the sampling rate based on each request's Trace ID. The choice of strategy depends on the level of correlation you need between requests.

Common commands:

-Dotel.traces.sampler=traceidratio
-Dotel.traces.sampler.arg=0.2

Custom Sampler Extension

OTEL agent provides a plugin interface Extensions, allowing you to implement custom sampling by defining a custom Sampler.

This is the interface that a custom Sampler needs to implement:

public class DemoSampler implements Sampler {
  @Override
  public SamplingResult shouldSample(
      Context parentContext,
      String traceId,
      String name,
      SpanKind spanKind,
      Attributes attributes,
      List<LinkData> parentLinks) {
    if (spanKind == SpanKind.INTERNAL && name.contains("greeting")) {
      return SamplingResult.create(SamplingDecision.DROP);
    } else {
      return SamplingResult.create(SamplingDecision.RECORD_AND_SAMPLE);
    }
  }

  @Override
  public String getDescription() {
    return "DemoSampler";
  }
}

In the interface implementation, you can see there are several parameters: parentContext, traceId, name, spanKind, attributes, parentLinks.

Therefore, in a custom Sampler, you can decide whether to sample a trace from multiple dimensions:

  • Attribute filtering
  • Trace ID ratio sampling
  • name filtering, etc.

However, this still cannot avoid discarding spans that contain errors in a ratio-based sampling, because shouldSample() is called when the span is initialized, while errors are usually filled into the corresponding events only when the span ends.

Feedback

Is this page helpful?