Part 8 · 1 chapters · ~8 min
Observability for Asynchronous Work
The metrics that matter (queue depth, age of oldest message, processing rate, failure and retry rates, DLQ depth), tracing across queues with propagated context, structured logs with job and run ids, dashboards per queue, SLOs for asynchronous work, and alerting on staleness rather than volume.
11
Seeing the invisible
| signal | why | alert |
|---|---|---|
| age of the oldest message | the user-felt delay; better than depth | payouts older than 5 minutes |
| queue depth and inflow vs outflow | is the backlog growing? | outflow < inflow for 15 minutes |
| failure and retry rates per job type | a partner degrading | failure rate above baseline |
| DLQ depth | customers affected | any new DLQ message for money jobs |
| job duration percentiles | capacity planning, stuck jobs | p99 near the timeout |
code
// propagate trace context through the message so traces span request → queue → worker
await queue.add('submit-payout', { payoutId, traceparent: currentTraceparent() });
worker.process(async job => context.with(extractContext(job.data.traceparent), () => submit(job.data.payoutId)));SLOs for async work are about freshness: "99% of payouts reach a final state within 10 minutes" (SRE part 2). Alert on that, not on queue size alone.