File Storage and Uploads
Uploading and serving user files at scale: presigned URLs, multipart and resumable uploads, asynchronous scanning and processing, metadata versus bytes, private downloads, deduplication and lifecycle tiers.
Brief, questions and numbers
- Users upload documents (KYC IDs, statements, receipts) and images, and later view or download them securely.
| question | answer we assume |
|---|---|
| sizes? | 100 KB to 200 MB |
| volume? | 2M uploads/day |
| privacy? | KYC documents are highly sensitive |
| retention? | 7 years for regulated documents |
| processing? | virus scan, thumbnails, OCR for IDs |
uploads = 2M/day ≈ 23/s average, peaks ~200/s bytes = 2M × 2 MB average ≈ 4 TB/day ≈ 1.4 PB/year before lifecycle tiers metadata = 2M rows/day in Postgres: small
v1, the break, and v2
v1. Clients upload files to the API server, which writes them to local disk and stores the path in the database.
The break. API servers become bandwidth-bound and stateful (files on one machine), large uploads tie up workers and fail on flaky networks, and unscanned files are served immediately.
v2. Presigned direct uploads to object storage with multipart, a metadata table separate from bytes, event-driven scanning and processing, signed short-lived downloads, encryption with per-tenant keys for KYC files, and lifecycle rules moving old files to cheaper tiers.