Part 6 · 1 chapters · ~8 min

Reliability Habits

Timeouts on every call, bounded retries with budgets, explicit limits and load shedding, graceful shutdown, honest health and readiness checks, dependency classification (critical versus optional), degraded modes, and scheduled drills.

11

Defaults that prevent incidents

code
// graceful shutdown in Node
process.on('SIGTERM', async () => {
  ready = false;                                   // readiness fails: the load balancer stops sending traffic
  await sleep(5_000);                              // let endpoints update
  server.close();                                  // stop accepting, finish in-flight requests
  await worker.close();                            // stop taking jobs, finish the current one
  await pool.end(); process.exit(0);
});
// classify dependencies: critical ones fail the request, optional ones degrade
const [balance, offers] = await Promise.allSettled([ledger.balance(id, { timeout: 300 }), promos.offers(id, { timeout: 150 })]);
if (balance.status === 'rejected') throw balance.reason;           // critical
const shownOffers = offers.status === 'fulfilled' ? offers.value : [];   // optional: degrade silently
RELIABILITY HABITS
defaults every service ships with
timeoutsEvery outbound call has a timeoutthat fits the caller's budget.bounded retriesBackoff with jitter, retrybudgets, only for idempotentoperations.limitsPool sizes, queue lengths, payloadsizes and rate limits areexplicit.graceful shutdownStop accepting, finish in-flight,drain consumers, then exit.health and readinessReadiness reflects ability toserve; liveness only detectsdeadlock.drillsRestore backups, fail over, killpods: on a calendar.
swipe the figure sideways, or tap expand for full screen
1/5
timeouts
The default in many HTTP clients is no timeout or a very long one; one slow dependency then holds every connection.
no default timeoutsset them explicitly