M5: Change-Intensive Apps
A collaborative document editor in four rounds: last write wins and the paragraph it loses, the document as a CRDT with OT as the alternative, presence as ephemeral state anchored by identity, and offline with state-vector sync, compaction and the semantic conflicts no data structure can decide.
The brief and the questions
"A document editor for our proposals team: rich text, comments, a few people on the same proposal at once, like Docs. Field staff edit on trains with bad signal. We have 400 users and about 5,000 documents."
- How many people edit the same document at the same time? Usually one; often two; a kickoff has ten; a company-wide template has fifty viewers and three editors.
- How big is a document? 5 to 50 pages; 20k to 200k characters; images by reference.
- How often does a document change while open? Continuously: every keystroke is a change. Latency to see a collaborator's keystroke: under 500 ms feels live.
- Offline? Yes: trains, basements, 30-minute gaps; edits made offline must not be lost and must merge when back.
- What is the cost of a lost edit? High: proposals are deadline work; "the app deleted my paragraph" is a trust-ending bug.
- What must be preserved? History and versions (who changed what, restore a version); comments anchored to text that survives edits.
- Devices? Laptops and iPads; Chrome and Safari.
- Infrastructure? A server is available; peer-to-peer is not required.
- FR: rich text editing; multiple simultaneous editors; cursors and presence; comments anchored to text; version history and restore; offline editing with merge on reconnect; read-only sharing.
- NFR: 2 to 10 concurrent editors (50 as the ceiling); 200k-character documents; remote keystrokes visible under 500 ms; instant open from the device; no edit ever lost, online or offline; 30-minute offline windows merge without a dialog; memory under 3× the text size; a server, but one that can be a relay.
v1: save the document, and who loses
// v1: a rich-text editor whose state is saved whole on a debounce. perfect for one author
const editor = useEditor({ content: doc.html, onUpdate: debounce(({ editor }) => api.put(`/docs/${id}`, { html: editor.getHTML() }), 2000) })
// two authors: the second PUT overwrites the first's paragraph with a body that never had it. no error. "the app deleted my work"
// with If-Match (ETag): the second PUT gets 409; the client must merge whole bodies: the problem v2 solves- Single-author documents: most of the 5,000 are edited by one person at a time. v1 is correct for them and ships in a week.
- Its failure is silent and total: last write wins on the whole document loses the paragraph the other author wrote. The number that breaks it is concurrent writers, and it breaks at two.
- Optimistic concurrency (ETag + If-Match) turns the silent loss into a visible 409, and is the right answer for coarse, rarely concurrent records (settings, a profile). For prose, "reload and retry" throws away the typing.
- Whole document: LWW or 409 on everything (v1).
- A field or block: per-paragraph versions; conflicts only when two people edit the same block. Cheap to add; fails exactly when collaboration is closest.
- An operation (OT): insert and delete at positions, transformed against concurrent operations by a server that orders them. Google Docs.
- An identity (CRDT): every character has an id and a neighbour; concurrent operations commute by construction; no server ordering needed. Yjs, Automerge, Figma.
Round two: the document as a CRDT
// v2: the document is a CRDT; the editor binds to it; a provider syncs it. no PUT, no debounce, no merge code in the app
import * as Y from 'yjs'
import { WebsocketProvider } from 'y-websocket'
import { IndexeddbPersistence } from 'y-indexeddb' // v4: local-first persistence, one line
const doc = new Y.Doc()
const local = new IndexeddbPersistence('doc-42', doc) // loads instantly from the device; writes every update
const net = new WebsocketProvider(WS_URL, 'doc-42', doc) // sync protocol: state vectors, diffs, reconnection with backoff
const fragment = doc.getXmlFragment('prosemirror') // the editor's tree lives in the CRDT
// Tiptap / ProseMirror binding: the editor renders the fragment and writes changes into it as CRDT ops
const editor = new Editor({ extensions: [StarterKit.configure({ history: false }), // history must be CRDT-aware: use the collaboration undo manager
Collaboration.configure({ document: doc }),
CollaborationCursor.configure({ provider: net, user: { name, color: colorFor(clientId) } })] }) // v3: awareness, relative positions, overlays
// observing for the rest of the UI (title, word count): the CRDT emits granular events
doc.getText('title').observe(e => setTitle(doc.getText('title').toString()))
net.on('status', ({ status }) => setOnline(status === 'connected')) // UI shows offline honestly; editing continues
// undo: per-user, over the user's own ops, across remote interleaving. Y.UndoManager(fragment, { trackedOrigins: new Set([editor]) })- Offline is required and the server may be a relay: that is the CRDT's case. OT would need the server in every merge and a rebase of 30 minutes of offline ops; possible (Docs does something like it) but the CRDT does it by construction.
- Memory at 200k characters: Yjs compresses sequentially typed runs into single items; typical documents are ~2 to 3× the text, inside the 3× requirement. Automerge is heavier; a hand-rolled CRDT is a year of bugs.
- Rich text and structure: Y.XmlFragment carries the ProseMirror tree with attributes; Tiptap's Collaboration extension binds it. The CRDT is the only source of truth; the editor renders it.
- Every keystroke is a CRDT operation applied locally (0 ms) and broadcast through a provider; remote operations arrive and merge; the editor re-renders the changed nodes. Latency to see a collaborator's keystroke is the network round trip, ~100 ms on a decent link.
- The server is small: authenticate, keep a room per document, relay updates, persist an update log and periodic snapshots. ~200 lines, or y-websocket's server as is.
- Undo is per user, over that user's own operations, across remote interleaving: the CRDT's undo manager, with the editor's own history disabled.
- Comments anchor to relative positions (the id of the character they follow), so they ride with the text through everyone's edits.
// the server for a CRDT document is small: authenticate, relay updates to the room, persist snapshots
// y-websocket's server does this; a custom one is ~200 lines. the document logic is in the CRDT library on both sides
wss.on('connection', async (ws, req) => {
const { docId, user } = await authorise(req) // who may read / write this document
const room = rooms.get(docId) ?? await loadRoom(docId) // Y.Doc from the latest snapshot + updates since
room.add(ws, user)
ws.on('message', buf => {
const msg = decode(buf)
if (msg.type === 'sync') room.handleSync(ws, msg) // state vector exchange / update apply + broadcast to others
if (msg.type === 'awareness') room.broadcastAwareness(ws, msg) // ephemeral: relay, never store
})
ws.on('close', () => { room.remove(ws); if (room.empty) scheduleSnapshot(room) })
})
// persistence: append updates to a log table (docId, seq, bytes); every N updates or M minutes, write Y.encodeStateAsUpdate(doc) as a snapshot and truncate
// read-only viewers: same sync, writes rejected by authorise; history / versions: snapshots with timestamps; a "restore" is a new update that sets the content- v2 buys concurrent editing with no lost edits and a relay server; pays: a CRDT library and binding in the critical path (complexity, and a dependency the team must understand), 2 to 3× the text in memory and per-character metadata on the wire (bytes), a different undo model (complexity), and the server's history as an update log plus snapshots instead of rows (complexity, money).
Round three: presence
Ten people on a kickoff document cannot tell where anyone is. They type over each other, two people fix the same typo, someone's cursor jumps because a paragraph above was deleted. The number that broke is simultaneous editors with no awareness of each other, and the fix lives beside the document, with opposite semantics.
- Awareness is a map clientId → { user, cursor, selection, lastSeen }, held by every peer, where each client owns and overwrites its own entry; never persisted, never merged; an entry expires 30 s after its last heartbeat.
- Cursors are relative positions (the character id they follow), resolved against each peer's local state; an index would be stale by the next remote keystroke. The same for selections and for comment anchors.
- Throttle at the source (50 to 100 ms, latest state only) and coalesce at the sink: at 50 peers the presence stream is 500 messages a second into every client; cursors render from a store once per frame (M4's rule), through the editor's decoration API as an overlay layer.
- Derived, not sent: "typing" is a cursor that moved in the last 2 s; "present" is an entry that exists; the avatar row is the map's keys; colours are a hash of clientId so Bob is the same colour on every screen.
- v3 buys the room; pays: a second channel with its own semantics (complexity), relative positions through the CRDT (complexity), a throttle and a per-frame render store (complexity), and presence traffic that grows with peers squared and exceeds document traffic at 50 (money; the server fans out selectively past that).
Round four: offline, sync, and semantic conflicts
A proposal lead edits for 40 minutes on a train, closes the laptop, and opens it in the office. The number that broke is the length of the disconnection: v2 held edits in memory and a crash lost them; a reconnect replayed nothing; and when two people rewrote the same paragraph apart, the merge was a sentence nobody wrote.
- Local-first: the CRDT persists to IndexedDB as it changes (y-indexeddb: an append-only update log plus a snapshot); the editor loads from the device in milliseconds, online or not; the network is an optimisation.
- State vectors: each peer knows the highest counter it has seen per client; two peers exchange vectors and send only the updates the other lacks. A day offline sends a day of updates and receives everyone else's; nothing is resent; a new peer with an empty vector receives a snapshot.
- Reconnection is the provider's job (backoff, resync: M8); the user sees the document update in place, their cursor on the same text, their edits already going out. No "unsaved changes" dialog exists because there is no unsaved state.
- Compaction: snapshots replace the log; tombstones are collected when every known peer has seen past them; a peer silent for N days is expired so it cannot block GC forever, accepting that its very old edits merge less precisely.
- Semantic conflicts are a product feature: the CRDT guarantees convergence, not intent. Detect two large edits to the same block in overlapping windows; surface both versions ("Bob also rewrote this paragraph while you were offline"); or constrain with a soft lock shown through presence. The data layer cannot decide this.
- v4 buys editing on a train and instant open; pays: device persistence (bytes, complexity), a sync protocol (complexity), compaction and stale-peer policy (complexity), and semantic-conflict UX (complexity, consistency).
- Stop: 50 editors, 200k characters, offline windows, no lost edits, history from snapshots. Suggestions mode, granular permissions per block, and an AI co-writer are features; the transport under all of it is M8.
The whole board, and the exercise
| Round | The number | The break | The design | Paid in |
|---|---|---|---|---|
| v1 | One author | Two authors: last write wins loses a paragraph silently | Whole-document PUT on a debounce; ETag + If-Match turns loss into 409 | Nothing; correct for single authors; the unit of change is the document |
| v2 | 2 to 50 concurrent editors; offline required | Positions are meaningless under concurrency | A sequence CRDT (Yjs) bound to the editor; a relay server with an update log and snapshots; CRDT-aware undo; relative-position comments | A library in the critical path; 2 to 3× memory; a different undo; log-based history |
| v3 | Ten people, no awareness | Typing over each other; cursors jumping | Awareness map (own-entry, expire, never persist); relative cursors; throttled, frame-rendered overlays; derived states | A second channel; presence traffic growing with peers squared |
| v4 | 40 minutes offline; two rewrites of one paragraph | Memory-only state; replay-less reconnect; convergence without intent | Local-first in IndexedDB; state-vector sync; snapshots and GC with a stale-peer rule; semantic-conflict detection and surfacing | Device state; a protocol; compaction policy; conflict UX |
- The unit of change is the design: document, field, operation, identity. Pick it from the concurrency the brief admits to.
- OT and CRDTs both converge; topology and offline decide between them, and memory is the price of commutativity.
- Presence is the opposite of the document: overwritten, ephemeral, anchored by identity, rendered per frame.
- Convergence is not intent: the data layer merges; the product decides what to do when two people meant different things.