Part 7 · 4 chapters · ~24 min

Networking and API Design in Node

The net and http stack with llhttp, keep-alive and agents, HTTP/1.1, HTTP/2 and HTTP/3 in Node, undici and fetch, TLS costs and session resumption, WebSockets and SSE, API design at staff level, framework internals and routing algorithms, validation and serialisation cost, client-side resilience, and graceful shutdown done correctly.

22

The http stack, keep-alive, HTTP/2 and HTTP/3

protocolin Nodewhat it gives you
HTTP/1.1node:http, llhttpkeep-alive, one request at a time per connection; clients open several connections
HTTP/2node:http2, nghttp2multiplexed streams on one connection, header compression (HPACK); great for gRPC and many small requests; usually terminated at the load balancer for browsers
HTTP/3QUIC work in progress in core; usually provided by the edge (CDN, Envoy)no TCP head-of-line blocking, faster handshakes, connection migration
code
// server timeouts that matter in production
const server = http.createServer(app);
server.keepAliveTimeout = 65_000;      // longer than the LB idle timeout (AWS ALB default 60 s)
server.headersTimeout = 66_000;        // must exceed keepAliveTimeout
server.requestTimeout = 30_000;        // whole request, defends against slow clients
AN HTTP REQUEST THROUGH NODE
socket to llhttp to your handler and back, on a keep-alive connection
kernellibuvllhttphttp.Serverhandlersocket readablebytes
swipe the figure sideways, or tap expand for full screen
1/5
bytes arrive
The kernel reports the socket readable via epoll/kqueue; libuv reads into a buffer during the poll phase.
kernel readiness → libuv readno thread per connection
23

undici, TLS, WebSockets and SSE

code
// undici: the HTTP/1.1 client behind global fetch, with explicit pooling
import { Agent, setGlobalDispatcher, request } from 'undici';
setGlobalDispatcher(new Agent({ connections: 128, keepAliveTimeout: 30_000, pipelining: 1 }));
const { statusCode, body } = await request('https://api.internal/ledger', { headersTimeout: 2_000, bodyTimeout: 5_000 });
console.log(statusCode, await body.json());
topicwhat to know
undici vs old http clientwritten from scratch for performance, a dispatcher model, connection pools per origin, used by global fetch; always set timeouts
TLS costa full handshake costs a round trip (TLS 1.3) and CPU for key exchange; reuse connections (keep-alive) and enable session resumption so reconnects are cheap; ALPN negotiates h2 vs http/1.1
WebSockets (ws)one long-lived socket per client; memory per connection matters; permessage-deflate costs CPU and memory per socket, enable it only for large compressible messages; scale out with sticky sessions or a pub/sub backplane (Redis)
SSEone-way server push over plain HTTP, automatic reconnect with Last-Event-ID, works through most proxies; simpler than WebSockets when the client only listens
long pollingfallback when nothing else passes through a hostile network
24

API design, frameworks, validation and resilience

API design at staff level
  1. Resources and verbs modelled on the domain, not on tables.
  2. Errors in one machine-readable shape (RFC 9457 problem details) with stable codes.
  3. Pagination by cursor (keyset), never offset, for anything that grows.
  4. Idempotency keys on every unsafe operation that moves money or sends messages.
  5. Versioning by additive change first; breaking changes behind a new version with a deprecation window.
  6. Timeouts, retries and circuit breakers on every outbound call (SRE part 9).
code
// Fastify: schema-validated input and compiled output serialisation
fastify.post('/transfers', {
  schema: {
    body: { type: 'object', required: ['amountKobo', 'to'], properties: { amountKobo: { type: 'integer', minimum: 1 }, to: { type: 'string' } } },
    headers: { type: 'object', required: ['idempotency-key'] },
    response: { 201: { type: 'object', properties: { id: { type: 'string' }, state: { type: 'string' } } } },
  },
}, async (req, reply) => reply.code(201).send(await createTransfer(req.body, req.headers['idempotency-key'])));
ROUTER AND FRAMEWORK OVERHEAD
relative requests per second for a hello-world route (illustrative shape, run your own benchmark)
node:http100Fastify90Hono (Node adapter)80Koa70Express 545NestJS on Express40
swipe the figure sideways, or tap expand for full screen
1/5
the baseline
Raw node:http is the ceiling: no routing, no middleware, no validation. Every framework subtracts from it. The numbers are a shape, not a benchmark: measure your own with autocannon.
the ceiling, with nothing on topillustrative; run autocannon yourself
25

Graceful shutdown done correctly

code
// the shutdown most services get wrong
let shuttingDown = false;
app.get('/readyz', (req, res) => res.status(shuttingDown ? 503 : 200).end());

process.on('SIGTERM', async () => {
  shuttingDown = true;                          // 1. fail readiness so the LB stops sending new requests
  await sleep(10_000);                          // 2. wait for the LB / endpoints to notice (> its check interval)
  server.close();                               // 3. stop accepting; in-flight requests continue
  server.closeIdleConnections();                //    close idle keep-alive sockets now
  const t = setTimeout(() => process.exit(1), 20_000).unref();   // 4. hard deadline
  await queue.close(); await db.end();          // 5. drain workers, close pools after requests finish
  process.exit(0);
});

Order matters: stop being routed to, then stop accepting, then finish in-flight work, then close dependencies, all inside the orchestrator's grace period (Kubernetes default 30 s). Getting it wrong shows up as a burst of 502s on every deploy.