Part 9 · 1 chapters · ~8 min

Incident Log: The Event Storm

A composite incident: two consumers re-emit each other's events, retries amplify traffic until the broker saturates, and the remediation: emit only on real change, causation and correlation ids, hop limits, per-type rate alerts, kill switches and capacity isolation.

12

Timeline and lessons

A composite incident based on a failure mode widely reported with event-driven systems; numbers are illustrative.

code
// emit only on real change, and carry causation so loops are detectable
async function handle(e: CustomerUpdated) {
  const before = await repo.get(e.customerId);
  const after = normalisePhone(before);
  if (deepEqual(before, after)) return;                         // no-op: emit nothing
  await repo.save(after, tx);
  await outbox.add({ type: 'customer.updated', id: e.customerId,
                     causationId: e.eventId, correlationId: e.correlationId,
                     hops: (e.hops ?? 0) + 1 }, tx);
}
// consumers drop events with hops > 10 and alert: a loop is a bug, never normal

Capacity isolation: separate brokers or quotas for critical topics (payments) and bulk topics (profile updates), so a storm in one cannot starve the other.

INCIDENT: THE EVENT STORM
messages per second during a feedback loop between two consumers
normal400/s+2 min3,200/s+5 min26,000/s+9 min (broker saturated)90,000/safter kill switch400/s
swipe the figure sideways, or tap expand for full screen
1/4
the loop
Service A updated a customer when it received customer.updated (to normalise a phone number) and emitted customer.updated. Service B did the same for a different field. Each update triggered the other.
two consumers re-emittingan infinite loop through the broker