Part 13 · 3 chapters · ~18 min
Build and Expert Mode
The capstone (a mini runtime, a native addon and a fully observed production service), reading the Node source and contributing, how the TSC decides, sixty staff-level interview questions in six groups with model answers, and the war stories every Node team eventually tells.
41
The capstone and reading the source
code
# the Node repository map
lib/ JavaScript standard library (lib/internal/* is private)
src/ C++ bindings, Environment, the main entry (node.cc, node_file.cc, tcp_wrap.cc, ...)
deps/ V8, libuv, OpenSSL, llhttp, undici, nghttp2, ...
test/ parallel/, sequential/, ... run with: python3 tools/test.py test/parallel/test-fs-*.js
doc/api/ the documentation source
# follow a call from JS to C++
git grep -n "function readFile(" lib/fs.js
git grep -n "static void Open(" src/node_file.ccContributing: start with a doc fix or a test, read CONTRIBUTING.md and the collaborator guide, and expect review from collaborators. The Technical Steering Committee decides contested questions in public GitHub issues and meetings; reading a few TSC threads shows how runtime trade-offs are argued.
THE CAPSTONE
a mini runtime, a native addon, and a production service with every lesson applied
swipe the figure sideways, or tap expand for full screen
1/5
mini runtime
Build a tiny event loop over kqueue (macOS) or epoll (Linux) in C, embed V8 or QuickJS, and expose setTimeout and a TCP echo. You will reimplement poll timeouts and callback queues from part 2.
your own loop + an embedded enginetimers and sockets, nothing else
42
Sixty questions, in six groups
| group | sample questions (ten in each group in the full bank) | what a strong answer names |
|---|---|---|
| the loop | Order of nextTick, promise and setImmediate inside an IO callback? Why does an idle Node use no CPU? What is ELU? | phases, the poll timeout, queues drained between callbacks |
| V8 | Why can adding a property slow a function down? What is a deopt? Why did the heap OOM in a 2 GB container? | hidden classes, inline caches, feedback, old space vs RSS |
| IO and streams | What uses the threadpool? Why is pipe() dangerous? How does backpressure reach the client? | fs, dns.lookup, crypto, zlib; pipeline; highWaterMark and TCP windows |
| scaling | Workers vs cluster vs replicas? Why did adding pods break the database? How do you size a pool? | isolates, processes, the total-connections trap, PgBouncer |
| production | What happens on SIGTERM in your service? Why 502s after deploys? Liveness vs readiness? | the shutdown order, keepAliveTimeout vs LB, PID 1 |
| security and quality | How does prototype pollution happen? What does the permission model restrict? What does coverage not tell you? | schemas, null-prototype maps, flags, assertions over lines |
43
War stories
| story | what happened | the lesson |
|---|---|---|
| the OOMKill | pods restarted every few hours; heap graphs looked flat. RSS grew from Buffers held by an unbounded upload queue. | watch RSS and arrayBuffers, not just heapUsed; bound every queue |
| the blocked loop | p99 latency jumped to seconds for all routes when one admin endpoint exported a CSV with a synchronous loop over 400k rows. | stream exports; alert on event loop delay |
| the pool exhaustion | an error path returned before client.release(); after 20 such errors every request waited forever. | release in finally; connection timeouts; pool waiting metrics |
| the memory leak | a per-request listener on a shared emitter; MaxListenersExceededWarning had been ignored for months. | treat that warning as a bug; three-snapshot technique |
| the cold-start bill | a Lambda importing the whole AWS SDK and an ORM spent 1.5 s initialising; traffic spikes meant thousands of cold starts. | bundle and tree-shake, lazy-load, provisioned concurrency for steady load |
| the DNS stall | outbound calls slowed under load; four threadpool threads were busy hashing passwords, so dns.lookup queued behind them. | workers for CPU, cache DNS, size the pool |