Part 4 · 2 chapters · ~18 min
CI/CD at Scale
Pipelines as code from shared templates, job graphs that run only what is affected, layered caching, artefacts built once and signed with provenance enforced at admission, and pipeline speed and trust as metrics; then progressive delivery: rolling, blue-green, canaries with automated analysis, feature flags, and choosing per change.
9
Pipelines as code, caching and artefacts
a fast, trustworthy answer on every commit
- Pipelines as code, built from shared templates the platform publishes.
- A job graph that runs only the affected projects.
- Caches: dependencies, Docker layers and remote build outputs.
- Build once: push by digest, scan, sign, attach an SBOM.
- Provenance plus an admission policy, so clusters refuse images your CI did not build.
- Measure speed and trust: under 10 minutes to a mergeable result, with flaky tests quarantined.
code
# .github/workflows/ci.yml: a service using the platform's shared workflow
name: ci
on: [pull_request, push]
jobs:
service:
uses: acme/platform/.github/workflows/service.yml@v4 # reviewed, versioned, owned by platform
with:
service: payments
node-version: 22
e2e: affected # only if payments or its deps changed
permissions: { id-token: write, contents: read, packages: write, attestations: write }
# inside service.yml (abridged)
# - uses: docker/build-push-action@v6
# with: { push: true, tags: "${{ env.REGISTRY }}/payments:${{ github.sha }}",
# cache-from: type=registry,ref=…:buildcache, cache-to: type=registry,ref=…:buildcache,mode=max,
# provenance: true, sbom: true }
# - uses: sigstore/cosign-installer@v3
# - run: cosign sign --yes $REGISTRY/payments@$DIGEST # keyless, via the CI's OIDC identityPIPELINES AS CODE, CACHING AND ARTEFACTS
a pipeline that stays under ten minutes as the repo, the team and the test suite grow
swipe the figure sideways, or tap expand for full screen
1/6
as code
Pipelines as code: the pipeline is defined in the repository (GitHub Actions, GitLab CI, Buildkite, CircleCI), reviewed like code, and reused through shared workflows or templates the platform team publishes (one "build and deploy a service" template rather than 80 hand-written copies).
10
Progressive delivery
a few first, measured, then everyone
- Rolling updates are the default, for low-risk changes.
- Blue-green means one switch and an instant rollback, at the cost of double capacity.
- Canary moves traffic in steps, pausing at each.
- Automated analysis compares the canary with the baseline and aborts on its own.
- Feature flags separate deploying code from releasing it.
- Choose per change. Database migrations always follow expand, migrate, contract.
code
# Argo Rollouts: canary with automated analysis
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: payments, namespace: prod }
spec:
strategy:
canary:
steps:
- setWeight: 5
- pause: { duration: 5m }
- analysis: { templates: [{ templateName: payments-health }] }
- setWeight: 25
- pause: { duration: 10m }
- analysis: { templates: [{ templateName: payments-health }] }
- setWeight: 100
---
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: payments-health }
spec:
metrics:
- name: error-rate
failureLimit: 1
successCondition: result[0] < 0.01
provider: { prometheus: { address: http://prometheus:9090, query: |
sum(rate(http_requests_total{app="payments",rollout="canary",code=~"5.."}[5m]))
/ sum(rate(http_requests_total{app="payments",rollout="canary"}[5m])) } }
- name: payment-success
successCondition: result[0] > 0.97 # a business metric, not only HTTP codesthe frontend version
Static frontends get canaries too: CDN weighted routing between two deployed versions, or a flag that serves the new bundle to a cohort, with the analysis reading RUM (the Disciplines course part 5): crash-free sessions and INP by version. The Big-company FE course part 6's release rings are this same idea.
PROGRESSIVE DELIVERY
shipping to a few first, measuring, and letting the metrics decide whether to continue
swipe the figure sideways, or tap expand for full screen
1/6
rolling
Rolling update (the default): replace instances a few at a time. Simple, no extra capacity; but every user is soon on the new version, both versions serve at once during the roll, and rollback is another roll. Fine for low-risk changes with good readiness probes.