Part 7 · 2 chapters · ~12 min

File Storage and Uploads

Uploading and serving user files at scale: presigned URLs, multipart and resumable uploads, asynchronous scanning and processing, metadata versus bytes, private downloads, deduplication and lifecycle tiers.

15

Brief, questions and numbers

the brief
  1. Users upload documents (KYC IDs, statements, receipts) and images, and later view or download them securely.
questionanswer we assume
sizes?100 KB to 200 MB
volume?2M uploads/day
privacy?KYC documents are highly sensitive
retention?7 years for regulated documents
processing?virus scan, thumbnails, OCR for IDs
code
uploads  = 2M/day ≈ 23/s average, peaks ~200/s
bytes    = 2M × 2 MB average ≈ 4 TB/day ≈ 1.4 PB/year before lifecycle tiers
metadata = 2M rows/day in Postgres: small
DIRECT UPLOADS WITH PRESIGNED URLS
the API never touches the bytes
clientAPIobject storagequeuescanner / processorPOST /uploads {size, type}presigned PUT URL (10 min)
swipe the figure sideways, or tap expand for full screen
1/4
ask permission
The client asks the API for an upload slot, stating size and type. The API checks quotas and creates a pending file record.
API validates and returns a signed URLfile record starts as pending
16

v1, the break, and v2

v1. Clients upload files to the API server, which writes them to local disk and stores the path in the database.

The break. API servers become bandwidth-bound and stateful (files on one machine), large uploads tie up workers and fail on flaky networks, and unscanned files are served immediately.

v2. Presigned direct uploads to object storage with multipart, a metadata table separate from bytes, event-driven scanning and processing, signed short-lived downloads, encryption with per-tenant keys for KYC files, and lifecycle rules moving old files to cheaper tiers.

the sentence
v2 buys unlimited, resumable, safe uploads with stateless APIs, and pays with asynchronous readiness the UI must show and more moving parts.