Part 4 · 1 chapters · ~8 min
Cron and Scheduling at Scale
Why naive cron breaks with replicas, leader election and locks for single firing, Kubernetes CronJobs, repeatable jobs in queue libraries, run ids and idempotent scheduled work, missed runs and catch-up policies, time zones and daylight saving, and monitoring that scheduled jobs actually ran.
6
Firing once, on time
code
# Kubernetes CronJob: one at a time, in Lagos time, with a deadline
apiVersion: batch/v1
kind: CronJob
spec:
schedule: "0 1 * * *"
timeZone: "Africa/Lagos"
concurrencyPolicy: Forbid
startingDeadlineSeconds: 3600
jobTemplate: { spec: { backoffLimit: 3, template: { spec: { containers: [ { name: recon, image: ledger, args: ["reconcile", "--date=yesterday"] } ], restartPolicy: Never } } } }
# dead man's switch: the job pings a monitor at the end; no ping by 03:00 → alert (Healthchecks.io, Cronitor, or a metric)CRON AT SCALE
schedules that fire once, on time, with many replicas running
swipe the figure sideways, or tap expand for full screen
1/5
the problem
Ten replicas each running cron means a nightly reconciliation runs ten times. And a single cron box is a single point of failure.
every replica fires, or one box failsboth are wrong