Part 4 · 1 chapters · ~8 min

Cron and Scheduling at Scale

Why naive cron breaks with replicas, leader election and locks for single firing, Kubernetes CronJobs, repeatable jobs in queue libraries, run ids and idempotent scheduled work, missed runs and catch-up policies, time zones and daylight saving, and monitoring that scheduled jobs actually ran.

6

Firing once, on time

code
# Kubernetes CronJob: one at a time, in Lagos time, with a deadline
apiVersion: batch/v1
kind: CronJob
spec:
  schedule: "0 1 * * *"
  timeZone: "Africa/Lagos"
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 3600
  jobTemplate: { spec: { backoffLimit: 3, template: { spec: { containers: [ { name: recon, image: ledger, args: ["reconcile", "--date=yesterday"] } ], restartPolicy: Never } } } }

# dead man's switch: the job pings a monitor at the end; no ping by 03:00 → alert (Healthchecks.io, Cronitor, or a metric)
CRON AT SCALE
schedules that fire once, on time, with many replicas running
schedulercron expressionsleader / lockonly one instance firesenqueue jobwith a run idworkersidempotent per run idmissed runscatch up or skip?time zonesAfrica/Lagos, DST elsewhere
swipe the figure sideways, or tap expand for full screen
1/5
the problem
Ten replicas each running cron means a nightly reconciliation runs ten times. And a single cron box is a single point of failure.
every replica fires, or one box failsboth are wrong