Part 5 · 3 chapters · ~22 min

Storage and Data

Object, block and file storage with classes, lifecycles and presigned uploads; managed databases with standbys, replicas, PITR, poolers and cloud-native engines; and caches, queues and streams as the building blocks that absorb load, decouple services and carry events.

15

Object, block and file

three interfaces, three prices
  1. Object storage: immutable blobs by key over HTTP, with eleven nines of durability.
  2. Storage classes and lifecycle rules turn retention into configuration.
  3. Block storage: a zonal disk for one VM, billed as provisioned, backed up by snapshot.
  4. File storage: shared POSIX, dearer and slower per operation.
  5. Access patterns: presigned uploads, and a private bucket that only the CDN can read.
  6. Choosing: object for nearly everything user-facing.
code
// the API signs; the browser uploads straight to the bucket; the API never touches the bytes
import { S3Client, PutObjectCommand } from '@aws-sdk/client-s3';
import { getSignedUrl } from '@aws-sdk/s3-request-presigner';

const s3 = new S3Client({ region: 'eu-west-1' });
export async function presignKycUpload(userId: string, contentType: 'image/jpeg' | 'image/png') {
  const key = `kyc/${userId}/${crypto.randomUUID()}`;
  const url = await getSignedUrl(s3, new PutObjectCommand({
    Bucket: 'acme-kyc-private', Key: key, ContentType: contentType,
    ServerSideEncryption: 'aws:kms',
  }), { expiresIn: 300 });                       // 5 minutes, one object, one content type
  return { url, key };
}

// browser
await fetch(url, { method: 'PUT', headers: { 'Content-Type': file.type }, body: file });
the public bucket
Turn on Block Public Access at the account level (AWS) or enforce public access prevention at the organisation level (GCP). A bucket then cannot be made public by accident. Anything public goes through the CDN with origin access control.
OBJECT, BLOCK AND FILE
three storage models with different interfaces, guarantees and prices, and what each is for
swipe the figure sideways, or tap expand for full screen
1/6
object
Object storage: PUT and GET whole objects by key in a flat namespace (the slashes in invoices/2026/10/inv-881.pdf are part of the key, not directories). Objects are immutable (overwrite replaces), up to 5 TB, durable to eleven nines across zones, and served over HTTP with signed URLs. Since 2020 S3 is strongly consistent for reads after writes; GCS always was.
16

Managed databases

what is managed and what is still yours
  1. Multi-AZ: a synchronous standby means no committed write is lost, but failover drops connections.
  2. Read replicas lag. Read-your-writes queries go to the primary.
  3. Point-in-time recovery from snapshots plus logs. Practise the restore and time it.
  4. Connections are scarce. Use a pooler.
  5. Aurora and AlloyDB separate compute from replicated storage.
  6. NoSQL and specialised stores are chosen by access pattern.
still yours on a managed DBwhy it matters
schema, indexes, queriesthe managed service will happily run a sequential scan on 200 M rows
instance size and storage typeundersized means latency; oversized means the bill
connection strategyautoscaled apps exhaust max_connections without a pooler
backup retention and restore drillsPITR you have never tested is a hope
network exposureprivate subnets only, no public endpoint, security group from app-sg
encryption and keysKMS key choice, rotation, who can decrypt snapshots
major version upgradesscheduled by you, tested by you, before end of support
failover in the client
When the standby is promoted, existing connections die and DNS moves. Clients need short DNS caching (the JVM's default caches forever), connection retry with backoff, and idempotent writes, so a write retried across a failover does not happen twice.
MANAGED DATABASES
what the provider runs for you, what you still decide, and the HA and replica topology underneath
swipe the figure sideways, or tap expand for full screen
1/6
primary, standby
Primary and synchronous standby: in a multi-AZ deployment the primary in zone a streams every write to a standby in zone b and waits for acknowledgement before committing (synchronous), so no committed write is lost if zone a dies. Failover promotes the standby and moves the DNS name, typically in 30 to 120 seconds (Aurora faster), and clients must reconnect.
17

Caches, queues and streams

absorb, decouple, broadcast
  1. Cache-aside with a TTL. Coalesce requests to stop the thundering herd.
  2. Invalidate on change, or accept staleness where it is harmless. Never cache money as truth.
  3. Queues deliver at least once, with visibility timeouts and a dead-letter queue. Consumers must be idempotent.
  4. FIFO ordering per key (per account), with parallelism across keys.
  5. Streams are partitioned, retained and replayable, and each consumer group keeps its own offset.
  6. Choose by need: a cache for reads, a queue for work, a stream for shared ordered events.
code
// an idempotent SQS consumer: the message id is the dedupe key
for (const msg of batch.Records) {
  const { id: eventId, accountId, amountKobo } = JSON.parse(msg.body);
  await db.tx(async t => {
    const seen = await t.oneOrNone('INSERT INTO processed_events(id) VALUES ($1) ON CONFLICT DO NOTHING RETURNING id', [eventId]);
    if (!seen) return;                                   // already applied: ack and move on
    await t.none('UPDATE wallets SET pending_kobo = pending_kobo + $2 WHERE account_id = $1', [accountId, amountKobo]);
  });
}
// throwing here leaves the message to reappear after the visibility timeout;
// after maxReceiveCount it lands in the DLQ with an alarm on its depth
the DLQ alarm
A dead-letter queue nobody watches is a silent data-loss queue. Alarm when its depth goes above zero, give it an owner, and build a redrive tool so fixed messages can be replayed into the main queue.
CACHES, QUEUES AND STREAMS
the three managed building blocks that absorb load, decouple services and carry events
swipe the figure sideways, or tap expand for full screen
1/6
cache-aside
Cache-aside: the app reads the cache; on a miss it reads the database and writes the result to the cache with a TTL. Simple and the most common; the cost is staleness up to the TTL and a thundering herd when a hot key expires and every request misses at once (fix with request coalescing or early refresh).