Part 1 · 2 chapters · ~18 min

Infrastructure as Code

Desired state, recorded state and the real cloud; plans reviewed on every PR with policy as code; pipeline-only applies; drift detection; then modules, versions, environments, state splitting, and how Terraform, CloudFormation, CDK, Pulumi and Crossplane differ.

2

State, plan, review, apply, drift

every infrastructure change is a reviewed diff
  1. Desired state lives in files that are versioned and reviewed.
  2. Recorded state lives in a remote, locked backend, never on a laptop.
  3. The plan compares code, state and reality. A -/+ (replace) line is the dangerous one.
  4. Review the plan on every PR, with policy as code failing unsafe plans.
  5. Apply the saved plan from the pipeline. Humans do not apply to production.
  6. Drift: detect it nightly, then import or revert, and take away console write access.
code
# backend.tf: remote state with locking
terraform {
  backend "s3" {
    bucket         = "acme-tfstate-prod"
    key            = "services/payments/terraform.tfstate"
    region         = "eu-west-1"
    dynamodb_table = "tf-locks"            # (or use_lockfile = true on recent versions)
    encrypt        = true
  }
}

# protect what must never be replaced by accident
resource "aws_db_instance" "ledger" {
  identifier          = "ledger-prod"
  deletion_protection = true
  lifecycle { prevent_destroy = true }
  # …
}
code
# what a reviewer must read before approving
Terraform will perform the following actions:
  # aws_db_instance.ledger must be replaced
-/+ resource "aws_db_instance" "ledger" {
      ~ identifier = "ledger-prod" -> "ledger-prod-v2" # forces replacement
    }
Plan: 1 to add, 0 to change, 1 to destroy.          ← stop here and ask why
import, do not recreate
Resources created by hand can be brought under code with import blocks, so Terraform adopts the existing resource instead of creating a second one. Always plan an import first and make sure the diff is empty before anything is applied.
INFRASTRUCTURE AS CODE: THE LOOP
desired state in files, actual state in the cloud, a recorded state in between, and a plan that is reviewed before anything changes
swipe the figure sideways, or tap expand for full screen
1/6
desired state
Desired state: .tf files describe resources and their arguments (an aws_db_instance with engine postgres, size db.r6g.large, multi_az true). They are code: reviewed, versioned, linted, reused through modules, parameterised per environment.
3

Modules, environments, state splitting and the tools

infrastructure code at the scale of many teams
  1. Modules package everything a service needs, with inputs and outputs.
  2. Version them and pin them, and track adoption as you would for a design system.
  3. Environments are the same modules with different inputs. Promotion means applying the same version to the next environment.
  4. Split state by lifecycle and ownership.
  5. The tools: Terraform/OpenTofu, CloudFormation, CDK, Pulumi, Crossplane.
  6. Choose one tool and give it an owner.
code
# envs/prod/payments.tf: a service in one block, from the platform's module
module "payments" {
  source  = "app.terraform.io/acme/service/aws"
  version = "~> 3.2"

  name            = "payments"
  image           = "111122223333.dkr.ecr.eu-west-1.amazonaws.com/payments@sha256:9f2c…"
  cpu             = 1024
  memory          = 2048
  desired_count   = 6
  autoscale       = { min = 6, max = 30, target_inflight_per_task = 30 }
  public_paths    = ["/api/pay/*", "/api/transfer/*"]
  secrets         = ["prod/payments/paystack", "prod/payments/db"]
  slo             = { availability = 99.95, p99_ms = 800 }       # alarms generated from this
  tags            = { team = "payments", cost_centre = "cc-104" }
}
the module is the paved road
When the service module generates the alarms from the SLO, the IAM role from the secret list, and the cost tags from team inputs, every team gets those practices by default without having to know them. That is platform engineering (part 5) at its simplest.
MODULES, ENVIRONMENTS AND THE TOOLS
structuring infrastructure code for many teams and environments, and how Terraform, CDK, CloudFormation and Pulumi differ
swipe the figure sideways, or tap expand for full screen
1/6
modules
Modules: a module is a reusable unit with inputs and outputs (module "service" { name = "payments", cpu = 512, ... }) that creates everything a service needs: task definition, target group, listener rule, autoscaling, alarms, log group, IAM role with least privilege. Teams use the platform's modules instead of copying resources, so a fix to the module reaches every service on the next upgrade.