Skip to main content

How sampling rate detection works

Reading time: 0 minute(s) (0 words)

Sampling rate anomalies help you spot changes in how often your SLI reports data. This page explains when they appear, when they resolve, and how missing or delayed samples affect the result.

Both anomalies apply to standard SLOs, not composite SLOs.

Availability

Sampling rate anomalies require enablement for your organization. Contact Nobl9 Support to request access.

What Nobl9 compares​

Nobl9 estimates sampling frequency from the time between samples, not their values. Recent intervals have more influence, so the estimate follows sustained changes without reacting immediately to every irregular gap.

  • Sampling rate change compares each stream's recent rate with its learned usual rate, or baseline.
  • Sampling rate mismatch compares good or bad samples with total, using total as the reference. Their timestamps don't need to match exactly.

When detection starts​

Each stream first needs 15 minutes of history and 20 completed sample intervals by default. An interval is the time between consecutive samples. At one sample per minute, warmup takes at least 20 minutes; at one per hour, it takes at least 20 hours. Mismatch needs both streams to finish warmup.

Default settings​

ConditionDefault
Detect an anomalyDifference above 50%, sustained for two learned intervals
Resolve an anomalyDifference at or below 20%, sustained for two learned intervals

Contact Nobl9 Support to adjust these settings for your organization.

With a learned one-minute interval, confirmation takes two minutes after the estimate crosses the threshold, not after the source changes. For multiple streams, the longest learned interval sets the wait. Estimates take time to adjust, so detection and recovery aren't immediate.

Percentages depend on the reference rate. Going from 30 to 60 samples per hour is a 100% increase; going from 60 to 30 is a 50% decrease. Exactly 50% doesn't exceed the default detection threshold. Nobl9 applies this check to its rate estimates, not directly to your query's configured interval.

When an anomaly resolves​

Sampling rate change​

Here, sampling slows from four samples per minute to one at 09:00, then returns to four at 11:00.

The baseline adapts to slower sampling, then to its returnSampling drops from four samples per minute to one at 09:00, then returns to four at 11:00. Each change triggers an anomaly that resolves as the baseline adapts.
Samples per minute
01234DetectedResolvedDetectedResolved08:3009:0009:3010:0010:3011:0011:3012:00
Recent rateLearned baselineAnomaly active

The first anomaly resolves while sampling is still slower: the baseline has learned the new rate. Returning to four samples per minute triggers another change. Resolution means sampling has settled, not necessarily that the collection settings are correct.

Sampling rate mismatch​

Good stays at four samples per minute while total slows to one.

The mismatch stays open until the rates convergeTotal slows at 09:00 and returns to its original rate at 11:00. Shading shows the detected anomaly, including confirmation and recovery time.
Samples per minute
01234DetectedResolved08:3009:0009:3010:0010:3011:0011:3012:00
Good rateTotal rateAnomaly active

The anomaly resolves after total returns to four and the rate estimates converge. Unlike a rate change, a lasting mismatch doesn't become normal. Both streams stopping also doesn't count as recovery.

Both charts use default settings and UTC times. Shading shows when the anomaly is active.

Missing and delayed samples​

A long gap can trigger an anomaly without another sample arriving. Nobl9 allows for the stream's usual spacing and timing variation first: an empty minute is normal for hourly data. Streams that haven't finished warmup can't trigger sampling anomalies; No data detection checks for extended gaps separately.

Nobl9 waits 15 minutes by default before evaluating samples for a given metric time. A longer delay can still trigger an anomaly. For example, if total is delayed while good keeps arriving, the available data may show a mismatch.

Delayed samples may later fill in the SLI chart, but the annotation stays: it records what Nobl9 knew at detection time. New data can change the current estimate, not erase that earlier observation. An hour of minute-spaced samples delivered together still contributes one-minute intervals, not a burst of faster sampling.

The annotation shows the affected stream and rates in samples per hour. Difference registered from marks the start of the sustained difference, using metric timestamps in UTC.

Replay​

Replay uses the same detection rules with all stored samples in the selected range, including those that arrived too late for live detection. It doesn't apply the live evaluation delay.

Each run learns the sampling pattern from scratch, so include enough history for warmup. It also checks for silence through the end of the range. Findings can differ from live detection because the available data and starting history differ. Replay doesn't rewrite live annotations.

Use the Anomalies API for live and Replay findings, or the Annotations API for live chart annotations.