Part 8 · 1 chapters · ~8 min

Observability for Asynchronous Work

The metrics that matter (queue depth, age of oldest message, processing rate, failure and retry rates, DLQ depth), tracing across queues with propagated context, structured logs with job and run ids, dashboards per queue, SLOs for asynchronous work, and alerting on staleness rather than volume.

11

Seeing the invisible

signalwhyalert
age of the oldest messagethe user-felt delay; better than depthpayouts older than 5 minutes
queue depth and inflow vs outflowis the backlog growing?outflow < inflow for 15 minutes
failure and retry rates per job typea partner degradingfailure rate above baseline
DLQ depthcustomers affectedany new DLQ message for money jobs
job duration percentilescapacity planning, stuck jobsp99 near the timeout
code
// propagate trace context through the message so traces span request → queue → worker
await queue.add('submit-payout', { payoutId, traceparent: currentTraceparent() });
worker.process(async job => context.with(extractContext(job.data.traceparent), () => submit(job.data.payoutId)));

SLOs for async work are about freshness: "99% of payouts reach a final state within 10 minutes" (SRE part 2). Alert on that, not on queue size alone.