11 parts · 16 chapters

How Computers Work

Every performance mystery in the higher courses ends here. A JavaScript loop that is ten times faster over an array than over a linked list is fast because of cache lines and the prefetcher. A sort that runs faster on sorted input does so because of branch prediction. A tab that stutters on a 2 GB phone is being paged and throttled. A write that is acknowledged and then lost was never fsynced. This course is the machine those explanations refer to.

Eleven parts, bottom up. Transistors to logic gates to an ALU to the fetch-decode-execute cycle; the modern CPU with pipelines, branch prediction, out-of-order execution and SIMD; the memory hierarchy from registers to RAM and virtual memory; the operating system's processes, threads, scheduling, syscalls and interrupts; file systems and how SSDs really behave; networking from Ethernet frames to sockets and NIC queues; and a browser tab, seen as a process using all of it. Four further parts follow: how numbers are stored, concurrency in hardware, the GPU and the compositor, and measuring performance, each with numbers measured on this machine where it can.

gates to instructions · the CPU · the memory hierarchy · the OS · storage · networking · a browser tab on top · numbers · hardware concurrency · the GPU · measuringevery level · the foundations every staff engineer is expected to reason from
gates to instructionsTransistors, logic gates, adders, the ALU, registers and the fetch-decode-execute cycle.
the CPUPipelining and hazards, branch prediction, out-of-order execution, SIMD, and the cost of a miss.
memory hierarchyRegisters, L1 to L3, RAM, cache lines, locality, virtual memory, pages and the TLB.
the operating systemKernel and user space, processes and threads, scheduling, syscalls and interrupts.
storageBlocks, inodes, journaling, the SSD flash translation layer, and fsync.
networkingEthernet, IP, TCP, UDP, sockets, and the NIC queues under them.
a browser tabWhat the OS gives a renderer process: memory pressure, GPU access, power and throttling.
Under every other courseThe Browser and JavaScript courses run on this machine; the Algorithms course's complexity meets these constants; the Build Your Own kernel slice simulates part of P3. Read it whenever a higher-level explanation says "because of the cache".