Part 2 · 1 chapters · ~8 min
Sampling
Populations and samples, simple random, stratified and systematic sampling, selection, survivorship and non-response bias, the standard error and √n, sampling in observability (head and tail trace sampling, log sampling), reservoir sampling for streams, and when to stop collecting data.
3
Samples from streams
code
// reservoir sampling: a uniform sample of k items from a stream of unknown length, in O(k) memory
function reservoir<T>(stream: Iterable<T>, k: number): T[] {
const out: T[] = []; let i = 0;
for (const x of stream) { if (i < k) out.push(x); else { const j = Math.floor(Math.random() * (i + 1)); if (j < k) out[j] = x; } i++; }
return out;
}
# OpenTelemetry Collector: tail sampling keeps every error and slow trace, plus 5% of the rest
processors:
tail_sampling:
decision_wait: 10s
policies:
- { name: errors, type: status_code, status_code: { status_codes: [ERROR] } }
- { name: slow, type: latency, latency: { threshold_ms: 500 } }
- { name: baseline, type: probabilistic, probabilistic: { sampling_percentage: 5 } }SAMPLING WITHOUT FOOLING YOURSELF
the sample must look like the population
swipe the figure sideways, or tap expand for full screen
1/4
random
Statistics assume the sample is random. Convenience samples (the first 1,000 rows, today's users) can mislead badly.
known selection chancesnot the first 1,000 rows