Part 5 · 2 chapters · ~16 min

Capacity and Release Safety

Capacity planning from demand forecasts, capacity per unit at the SLO, headroom for failure, the real bottleneck and four kinds of load test; then release safety: small batches, rollback first, automatic rollback, budget-gated pipelines and targeted freezes.

8

Capacity planning and load testing

enough for the next peak, and knowing what breaks first
  1. Forecast organic growth plus known events, and plan for the peak of the peak.
  2. Measure capacity per unit at the point where latency breaks the SLO, not where things crash.
  3. Headroom: N+1 across zones, plus room for retries and autoscaler lag.
  4. The real bottleneck is almost never the stateless tier.
  5. Load, stress, soak and spike tests answer four different questions.
  6. Keep it current, and raise external limits weeks before a peak.
code
// k6: the salary-day spike, whole path, production-like data
import http from 'k6/http';
import { check } from 'k6';
export const options = {
  scenarios: {
    salary_spike: { executor: 'ramping-arrival-rate', startRate: 50, timeUnit: '1s',
      preAllocatedVUs: 500, stages: [ { target: 50, duration: '5m' },
                                      { target: 450, duration: '1m' },     // spike to forecast peak × 1.5
                                      { target: 450, duration: '20m' } ] },
  },
  thresholds: {
    'http_req_duration{name:transfer}': ['p(99)<800'],
    'http_req_failed{name:transfer}':   ['rate<0.001'],
  },
};
export default function () {
  const r = http.post(`${__ENV.BASE}/v1/transfers`, JSON.stringify(randomTransfer()),
    { headers: { 'Content-Type': 'application/json', 'Idempotency-Key': uuid() }, tags: { name: 'transfer' } });
  check(r, { 'accepted': x => x.status === 201 || x.status === 202 });
}
CAPACITY PLANNING AND LOAD TESTING
forecasting demand, finding the real limit, and keeping headroom where it counts
swipe the figure sideways, or tap expand for full screen
1/6
forecast
Forecast demand: organic growth from the trend (transfers per second at the daily peak, growing 6% a month), plus known events (salary days on the 25th to the 1st, a marketing campaign, Black Friday, end-of-year), plus the peak-to-average ratio. Plan for the peak of the peak: the busiest minute of the busiest day.
9

Release safety: canaries, rollbacks and budget gates

every change small, observed, reversible, gated
  1. Changes cause most incidents, and that includes config, flags, infrastructure and migrations.
  2. Small batches beat big ones, and long freezes often backfire.
  3. Roll back first, quickly, with tested and schema-compatible rollbacks.
  4. Automatic rollback on canary analysis or a fast burn soon after a deploy.
  5. Budget gates in the pipeline.
  6. Freezes that are short, targeted and have an exception process.
the schema rule
A rollback is only fast if the old code still works against the new schema. Expand (add the column, nullable), deploy code that writes both, backfill, deploy code that reads the new column, and contract (drop the old one) in a later release. No single release should depend on a schema change it also made.
RELEASE SAFETY
changes cause most incidents, so make every change small, observed, reversible and gated by the budget
swipe the figure sideways, or tap expand for full screen
1/6
changes cause incidents
Changes cause incidents: industry postmortem data consistently attributes most incidents to changes (often quoted around 70%). That is good news: changes are the thing you control. Every kind of change, not only code deploys, needs the same discipline: config, flags, infrastructure, database migrations, DNS, certificates.