Part 6 · 10 chapters · ~50 min
Running it: observability, deployment, and DR
Part 13 defined five signals, symptom-based alerting and designed degradation. This part implements them on cloud-native tooling and finds two things worth knowing: the fifth signal, correctness, is always custom because no cloud has an opinion about whether your ledger balances, and the observability bill is decided almost entirely by cardinality choices engineers make without noticing.
57
The five signals from Part 13, on cloud native tooling
| Signal | AWS | GCP | Note |
|---|---|---|---|
| Latency | CloudWatch metrics, percentiles | Cloud Monitoring distributions | Percentiles, never averages |
| Traffic | Request counts per service | Same | Compare against same hour last week |
| Errors | By type, not just 5xx | Same | Insufficient funds is not an error |
| Saturation | Pool depth, queue lag, CPU | Same | Connection pool is usually the real one |
| Correctness | Custom metrics | Custom metrics | No cloud gives you this. It is yours. |
the fifth signal is always custom
Part 13 argued correctness is a first-class signal because the first four can be green while the money is wrong. No cloud monitoring product has an opinion about whether your ledger balances, so
ledger_invariant_violations_total, suspense ageing and reconciliation break value are custom metrics you emit and alert on. That is the single most important dashboard in the bank and it is entirely your own work.58
CloudWatch, EMF, and the cost of a metric
worked numbers
CloudWatch custom metrics are priced per metric per month,
where a metric is a unique name plus dimension combination.
posting_latency{shard, rail, currency, region}
8 shards × 6 rails × 4 currencies × 4 regions
= 768 metrics from one instrument
add account_id as a dimension and it is
20 million metrics, which is a five-figure monthly bill.the discipline
- Dimensions must be low cardinality. Shard, rail, region and currency are fine. Account id, journal id and customer id are not, ever.
- Use EMF, the embedded metric format: log a structured JSON line and CloudWatch extracts metrics from it. You get high-cardinality fields in the log for querying and low-cardinality metrics for alerting, which is the right split.
- Aggregate before emitting where possible, rather than emitting per request.
- Sample logs, keep errors. Part 13 chapter 145 said the same about traces and the same reasoning applies.
why this is an architecture concern
Part 0 chapter 6 put observability at roughly 9% of the bill, above the warehouse. That number is almost entirely determined by cardinality decisions made by engineers adding dimensions without thinking about the multiplication. A single careless dimension can double the observability bill, and nobody notices until the quarterly review.
59
Cloud Monitoring, and OpenTelemetry as the escape hatch
AWS
GCP
AMP and AMG
CloudWatch plus X-Ray. Deeply integrated with AWS services, and the query language and the tracing UI are both weaker than dedicated tools.
(managed Prometheus and Grafana) are the answer when you want the open ecosystem.
(managed Prometheus and Grafana) are the answer when you want the open ecosystem.
Cloud Monitoring, Cloud Logging and Cloud Trace. Logging query language is genuinely good, and Trace integrates cleanly.
Managed Prometheus is a first-class option here too.
Managed Prometheus is a first-class option here too.
the portability argument for OpenTelemetry
Instrument with OpenTelemetry and export to whatever the platform provides. It costs a little more effort up front and it means the instrumentation itself is not lock-in, which matters because Part 0 chapter 3 identified operational tooling as the deepest lock-in layer. Changing where telemetry goes should be a config change, not a re-instrumentation project, and OTel is the only way to get that property.
60
Tracing across managed services
Part 13 chapter 145 wanted one trace from the API call through the ledger, the outbox, Kafka, four consumers and a provider. Managed services break traces at their boundaries unless you do specific work.
where traces break, and the fix
- Into the queue. The broker does not propagate trace context. Fix: put
traceparentin the message headers, as Part 13 specified. - Through the outbox. The database does not carry context. Fix: store trace id and span id as columns on the outbox row, which Part 13 did explicitly.
- Into managed transforms. A Glue job or a Dataflow pipeline is its own world. Fix: accept the break and correlate on the business id instead.
- Across a provider. NIBSS will not accept your trace header. Fix: correlate on your own reference, which Part 6 put in the connector interface.
the fallback that always works
When trace context cannot cross a boundary, the
journal_id can. Part 13 chose the business identifier as the correlation key precisely because it survives every boundary a trace header does not. A trace is better when available; the business id is what makes the system debuggable when it is not.61
Deployment: blue-green, canary, and money
| Strategy | How | Risk for a money path |
|---|---|---|
| Rolling | Replace instances gradually | Two versions write concurrently. Fine only if the schema is compatible both ways |
| Blue-green | Full parallel stack, switch traffic | Clean cutover, expensive, and database state is shared anyway |
| Canary | Small percentage to the new version | Best for a money path: real traffic, bounded exposure |
| Shadow | Duplicate traffic, discard results | Excellent for read paths, dangerous for writes |
what a money path adds to normal deployment practice
- Schema changes are expand-contract, always. Add the column, deploy code that writes both, backfill, deploy code that reads new, then drop. Four deploys, never one.
- The canary metric is correctness, not errors. Part 13's fifth signal: if the invariant violation count moves during a canary, roll back immediately regardless of what latency and error rate say.
- Never canary the ledger schema. Two versions with different expectations of the same table is how you corrupt data rather than merely break a request.
- Rollback must be tested, not assumed. A deploy you cannot reverse is a one-way door on the most critical system in the company.
62
Infrastructure as code, and what belongs in it
AWS
GCP
CDK
for typed, composable infrastructure in a real language, or Terraform for portability and a larger ecosystem.
CloudFormation underneath CDK, which shows in error messages and in drift handling.
CloudFormation underneath CDK, which shows in error messages and in drift handling.
Terraform is the common choice. Config Connector for Kubernetes-native management.
Deployment Manager exists and is largely superseded.
Deployment Manager exists and is largely superseded.
what belongs in IaC, and what does not
- In: networks, IAM, clusters, databases, topics, buckets, alarms. Everything whose configuration is a decision someone should review.
- In, and often forgotten: alarms and dashboards. An alert created by hand in a console vanishes when someone tidies up, and nobody notices until it does not fire.
- Out: application data, secrets values, and anything with a lifecycle faster than a deploy.
- Out, deliberately: emergency changes during an incident. Fix it by hand, then reconcile the code afterwards, and treat the drift as a work item rather than pretending it did not happen.
the state file point
Terraform state is itself critical infrastructure. It belongs in versioned, locked, backed-up remote storage with restricted access, because whoever can write the state can effectively rewrite the infrastructure. Teams that protect production carefully and leave the state file in a shared bucket have missed where the real access is.
63
Multi-AZ, multi-region, and the residency wall
worked numbers
within a region, across zones: synchronous replication, ~1-2 ms RPO zero achievable automatic failover RTO in seconds across regions: asynchronous only RPO is seconds to minutes manual, declared failover RTO in tens of minutes and for EU customer data, legally unavailable so resilience is built ACROSS ZONES, inside the region.
the cost of the residency wall, quantified
A global active-active design shares capacity across regions, so each region runs at a fraction of full load and absorbs another's failure. A residency-partitioned design cannot do that: each region must independently carry its own peak plus failure headroom, which is roughly 30 to 40% more infrastructure than a design free to shift load. That premium is a legal consequence, and naming it as such rather than as an engineering choice is the accurate framing.
64
Disaster recovery drills on managed services
the drills that must exist, and what each proves
- Instance failover. Force a failover on a shard. Proves the RTO from Part 1 chapter 14, including DNS and pool recovery.
- Zone loss. Remove a zone from the load balancer. Proves capacity headroom is real and not theoretical.
- Restore drill. Restore a random shard from backup and verify the invariant on the restored copy. Proves the backup and measures the true restore time.
- Cache flush. Flush the balance cache in business hours. Proves the Part 3 fallback works and that the cache is not load-bearing for correctness.
- Dependency failure. Break a provider in a controlled way. Proves the Part 6 circuit breaker and the Part 13 degradation order.
- Region loss, tabletop. Walk through it with the team. Proves the plan is understood, and it is the only one of the six you cannot safely do for real.
what managed services change about drilling
They make some drills easier and one much harder. Forcing a failover is an API call rather than pulling a cable, which is a genuine improvement. But you cannot test the provider's own failure modes: you cannot make S3 slow or KMS unavailable to see what happens. For those, the honest answer is a tabletop exercise and a documented assumption, and saying that is better than claiming a resilience you have not tested.
65
Cost observability as an SRE concern
why cost belongs with reliability rather than with finance
- A cost spike is usually an incident. A runaway retry loop, a scan without a partition filter, or a stuck job all show up in the bill before they show up anywhere else.
- Cost controls are availability controls. A query quota that fails a bad query is the same mechanism as a rate limit, and it protects the warehouse for everyone else.
- The team that creates the cost must see it, per Part 12's point about notification spend. Aggregated to a single monthly number, nobody owns it.
- Unit economics from Part 19 need infrastructure cost allocated per transaction, which requires tagging discipline from day one and is nearly impossible to retrofit.
AWS
GCP
Cost Explorer
and Budgets with alerts, cost allocation tags enforced by SCP so untagged resources cannot be created.
CUR into Athena for real analysis.
CUR into Athena for real analysis.
Billing export to BigQuery, which is the cleanest of the two: your cost data lands in the same warehouse as everything else and is queryable with SQL.
Budgets with Pub/Sub notifications for automated response.
Budgets with Pub/Sub notifications for automated response.
the practice worth adopting
An anomaly alert on daily spend per service, compared against the same day last week. It is the Part 13 business-metric pattern applied to cost, it catches runaway loops and forgotten resources within a day rather than a month, and it is perhaps three hours of work.
66
The decision, with numbers
AWS
GCP
EMF
CloudWatch with for metrics, disciplined dimensions, X-Ray sampled by tail, AMP and AMG if the team wants Prometheus. CDK or Terraform. Multi-AZ per region, no cross-region failover where residency forbids.
Observability roughly $3,000 to $5,000 per month and highly sensitive to cardinality choices.
Observability roughly $3,000 to $5,000 per month and highly sensitive to cardinality choices.
Cloud Monitoring and Logging, billing export to BigQuery, Cloud Trace. Terraform. Same multi-AZ posture.
Similar cost, with a better log query experience and cleaner billing analysis.
Similar cost, with a better log query experience and cleaner billing analysis.
the answer
"Instrument with OpenTelemetry regardless of platform, because Part 0 identified operational tooling as the deepest lock-in and instrumentation is the part you can keep portable cheaply. Watch cardinality, because observability is around 9% of the bill and that number is decided by engineers adding dimensions. Put correctness on the dashboard as a first-class signal, because no cloud provides it. And drill the failovers monthly, because Part 1 showed the documented RTO is a fraction of the real one."