Part 1 · 10 chapters · ~20 min

The ledger: where the invariant lives

The component with the least room for compromise. Everything the CBA module built rests on a database that can do a conditional multi-row atomic write and survive a node loss without forgetting an acknowledged commit. This part works out which managed services actually offer that, what each one costs, and why the most impressive option on either platform is the wrong one here.

9

The requirement, restated for a cloud database

Before any service names, restate what Part 1 of the CBA module actually demanded. Most cloud database debates are unresolvable because nobody wrote this down.

what the ledger needs, non-negotiably
  1. Multi-row atomic writes. Journal row plus two or more entries, all or nothing.
  2. A conditional write. "Insert this debit only if the balance permits", evaluated under a lock.
  3. Durability with RPO zero. An acknowledged commit survives node loss.
  4. Strong reads on the balance path, so read-your-own-writes holds.
  5. Ad-hoc aggregation, for the invariant check and reconciliation.
what it does not need, which widens the field
  1. Global distribution. Part 7 made regions a legal boundary, so cross-region consistency is not required and is actively unwanted.
  2. Unbounded elastic scale. Part 3 sharded it; each shard holds ~34 GB and a few hundred writes a second.
  3. Schemaless flexibility. A ledger wants a rigid schema that refuses malformed money.
worked numbers
the second list is why Spanner is not automatically the answer
and why a sharded managed Postgres is usually enough.

state the requirement first, and half the catalogue eliminates itself.
10

RDS, Aurora, and what Aurora actually changed

Aurora is described as "managed Postgres" and it is architecturally a different thing. The difference matters for exactly the properties a ledger cares about.

worked numbers
standard Postgres
  write → buffer pool → WAL → fsync → data pages later
  replica ships and replays the full WAL

Aurora
  write → ships only the redo log to a distributed storage layer
  storage is 6 copies across 3 AZs, quorum 4/6 for write
  replicas read the same storage, so they do not replay

the database and the storage engine were separated.
PropertyStandard RDSAurora
Replica lagReplay-bound, can growTypically under 100 ms, shared storage
Failover60 to 120 s typicalUnder 30 s, often under 15
Write amplificationFull 8 KB pages to replicasRedo only, far less network
Storage growthProvisioned in advanceAutomatic, to 128 TB
CostLowerRoughly 20 to 30% more for equivalent compute
Postgres compatibilityExactVery high, with some extension gaps
Backtrack / PITRPITRPITR plus fast in-place rewind on MySQL flavour
the property that decides it for a ledger
Failover duration is the RTO. Part 13 set a 60-second RTO for a shard primary and said the real number is whatever the drill measures. Aurora's architecture makes that target comfortable where standard RDS makes it tight. For a system where a failover happens during salary-day peak, the 25% cost premium buys a measurably better worst case, and that is an easy argument to make to a CFO.
11

Cloud SQL, AlloyDB, and Spanner

GCP offers three relational answers and they are genuinely different products rather than tiers.

Cloud SQLAlloyDBSpanner
EnginePostgreSQL, actualPostgres-compatible, rearchitectedNot Postgres. Own engine, PG interface
AnalogueRDSAuroraNo AWS equivalent
Scale-out writesNo, one primaryNo, one primaryYes, horizontally
External consistencyWithin instanceWithin instanceGlobal, via TrueTime
CostLowestModerateHighest by a distance
Fit for our ledgerSufficientBetter failoverBuys something we do not need
the trap to avoid
Spanner is the most technically impressive database on either platform and that is not a reason to use it. Part 3 already solved horizontal write scale with application-level sharding, and Part 7 made cross-region consistency legally unavailable. Paying for global external consistency inside a single legal region is paying for a property you cannot use.
12

Spanner: TrueTime, and the one thing it buys you

Worth understanding properly, because it is the clearest example of a cloud-only capability and because explaining it well is a strong interview signal.

worked numbers
the distributed transaction problem:
  two nodes cannot agree on "now" to better than clock skew
  so a transaction ordering needs coordination messages

TrueTime gives an interval, not a timestamp:
  TT.now() → [earliest, latest], with a bounded uncertainty ε
  backed by GPS and atomic clocks in every datacentre

commit protocol: pick a timestamp, then wait out ε before
acknowledging, so no later transaction can claim an earlier time

ε is a few milliseconds, so the commit wait is a few milliseconds.
that is the price, and it buys externally consistent global ordering.
when that is genuinely worth it
  1. A global ledger with no legal partitioning, where any account may transact with any other atomically.
  2. Write throughput beyond one node without application sharding.
  3. A team that cannot operate sharding, where the premium buys away real operational complexity.
why not us
All three fail for our design. Part 7 partitioned by region legally, Part 3 sharded at the application layer deliberately, and the sharding is a few hundred lines of routing we control. Spanner solves a problem we chose not to have. Being able to say that precisely is better than either dismissing it or reaching for it.
13

DynamoDB as a ledger: the honest version

Part 1 named DynamoDB as a viable alternative and moved on. Here is the full version, because at an AWS-first company this is the real conversation.

RequirementDynamoDB answerCost
Multi-row atomicityTransactWriteItems100-item ceiling, and 2x the write cost
Conditional writeConditionExpressionExcellent. Genuinely first-class
Durability3 AZs synchronouslyStrong, no tuning needed
Strong readsConsistentRead=true2x the read cost
Ad-hoc aggregationNoneThe invariant check needs another path
ShardingAutomaticYou never think about it again
Hot partition3,000 RCU / 1,000 WCU per partitionThe Part 3 hot row, again
what changes if you choose it
  1. The invariant check moves to a Streams-driven verifier, since there is no GROUP BY. A Lambda consumes the stream and maintains per-journal sums, alerting on any non-zero.
  2. The balance read is a query on a sort key rather than a SUM, so snapshots become mandatory rather than an optimisation.
  3. Reconciliation exports to S3 and runs in Athena, because it cannot run in the database.
  4. Access patterns freeze early. Single-table design means a new query pattern may need a new GSI or a backfill.
the honest verdict
DynamoDB can hold a ledger and several real banks run one on it. What you give up is the ability to ask the database a question you did not plan for, which for reconciliation and regulatory reporting is a recurring tax. I would choose it at a company where the operational expertise is entirely DynamoDB, and Postgres everywhere else, and I would say exactly that rather than pretending one is objectively correct.
14

Failover behaviour, and measuring real RTO

worked numbers
the documented number is promotion time.
the number that matters is customer-visible outage:

  detection                     5 to 30 s  ← health check interval
  promotion                   5 to 60 s  ← the documented figure
  endpoint propagation     5 to 40 s  ← DNS TTL, often forgotten
  pool recovery             5 to 30 s  ← stale connections
  backlog drain           10 to 120 s ← queued work, thundering herd

total: 30 s to 4 minutes. the documented number was 30 s.
what to actually do about it
  1. Short DNS TTL, or an endpoint that does not depend on DNS. RDS Proxy and Cloud SQL Auth Proxy both help here.
  2. Connection pools that validate before handing out a connection, so a stale one fails fast rather than hanging.
  3. Fast-failing timeouts, per Part 2, so the backlog does not become the outage.
  4. Drill it monthly and publish the measured number, per Part 13. The RTO is what the drill says, not what the docs say.
failover
what the documentation says versus what the drill measures
swipe the figure sideways, or tap expand for full screen
1/7
failure
A shard primary fails at peak. The clock starts here, from the customer’s point of view.
15

Backup, PITR, and restoring 34 GB in anger

Part 15 warned that growth breaks recovery before it breaks the budget. On managed services the backup is automatic and the restore is still your problem.

AWS
GCP
Restore creates a new instance
Automated backups plus continuous WAL archiving to S3. PITR to any second within the retention window. , which takes time proportional to volume size and then warms slowly as pages fault in from S3.
Automated backups plus PITR on Cloud SQL and AlloyDB. Same model: restore is a new instance. Cloud SQL supports cloning, which is faster for a copy than a restore.
worked numbers
restoring one 34 GB shard, measured rather than assumed:

  snapshot restore initiation           ~2 min
  volume available                       ~8 min
  warm-up as pages fault from S3   ~25 min
  WAL replay to the target time         variable
  verify the invariant before serving   ~4 min

~40 minutes for one shard. with 4,096 logical shards over N
nodes, a full-region restore is a multi-hour operation.
the check nobody runs
Restore a random shard every month, automatically, and verify the invariant on the restored copy. Part 13 said a backup that has never been restored is a hypothesis; on managed services people trust the console and the hypothesis goes untested for years. The drill also gives you the real number for the DR plan, which is the one the regulator asks for.
16

Connection management: pgbouncer, RDS Proxy, and Lambda

A Postgres connection costs memory and a backend process. A few hundred is comfortable; a few thousand is not. Serverless compute makes this acute.

worked numbers
Postgres: each connection ≈ 5 to 10 MB and one process
  500 connections ≈ 3 GB and 500 processes  fine
  5,000 connections                               the database falls over

Lambda makes this worse structurally:
  each concurrent execution is its own environment
  1,000 concurrent Lambdas → 1,000 connections
  and they are short-lived, so churn is constant
SolutionWhere it runsNote
pgbouncerYou run itTransaction pooling is the useful mode. Breaks session features like prepared statements unless configured
RDS ProxyAWS managedPooling plus failover-aware connection handling. Designed for the Lambda case
Cloud SQL Auth ProxySidecarPrimarily auth and encryption; pooling is still yours
Application poolIn-processCorrect for long-lived containers, useless for functions
the design consequence
This is one of several reasons Part 3 of this module concludes the posting path belongs in a long-lived container rather than a function. A container holds a warm pool of ten connections for hours; a thousand concurrent functions hold a thousand cold ones. The connection model is a first-class input to the compute decision, not an afterthought.
17

Sharding on managed databases

Part 3 sharded across 4,096 logical shards mapped to physical nodes. On a managed service, that mapping becomes a set of instances and the routing stays ours.

what stays the same
  1. The logical shard layer and the routing table, which are application code and fully portable.
  2. The cross-shard saga with in-transit suspense accounts.
  3. The invariant check per shard, which is a per-instance query.
what the cloud changes
  1. Adding a physical node is minutes rather than a procurement cycle, which makes rebalancing genuinely practical rather than theoretical.
  2. Each shard is a separate billable instance, so the cost model rewards fewer, larger nodes and the operational model rewards more, smaller ones. Resolve that explicitly.
  3. Per-shard backup and failover multiply the operational surface: 16 instances means 16 failovers to drill.
  4. Aurora Limitless and AlloyDB horizontal options exist and are worth evaluating, though both are newer than the workload deserves for a ledger.
worked numbers
a practical starting shape:

  4,096 logical shards     fixed forever, per Part 3
  8 physical instances   512 logical shards each
  each: ~34 GB × 512 ÷ 4096 ≈ ~4 TB, r6g.4xlarge class
  each with a synchronous replica in another AZ

growth is: add instances, reassign logical ranges, copy, cut over.
a config change and a copy, never a rehash. that was the point.
18

The decision, with numbers

AWS
GCP
Aurora PostgreSQL
, 8 writer instances, each with a reader in a second AZ. RDS Proxy in front. PITR with 35-day retention.

Roughly $14k to $18k per month at this shape, dominated by instance hours rather than I/O.
AlloyDB or Cloud SQL for PostgreSQL, same topology. AlloyDB for the better failover and read performance, Cloud SQL if cost dominates.

Roughly $13k to $17k per month, broadly comparable.
CriterionWinnerMargin
Postgres fidelityCloud SQLIt is actual Postgres. Aurora and AlloyDB both diverge slightly
Failover speedAuroraConsistently the fastest of the three
Operational toolingAuroraMore mature, more third-party integration
Read scalingAlloyDBColumnar engine for analytical reads is genuinely novel
CostCloud SQLModest difference, not decisive
If you need horizontal writesSpannerNo AWS equivalent, and we do not need it
the answer
"Aurora PostgreSQL, sharded at the application layer exactly as designed in Part 3, with RDS Proxy for connection management. The deciding property is failover duration, because that is the RTO and a failover will happen during peak. On GCP I would pick AlloyDB for the same reason and the architecture would be identical. What I would not do is reach for Spanner: it solves horizontal write scale and global consistency, and Part 3 sharded deliberately while Part 7 made cross-region consistency legally unavailable. Paying for a property you cannot use is the most common cloud mistake."