Part 20 · 19 chapters · ~20 min
Deep dive: the web transport stack
Part 14 chose protocols by conversation shape. This part is the layer underneath, and it earns its place because a meaningful share of your p99 happens before application code runs and is invisible to application tracing. Four round trips to a new origin is 360 milliseconds on a Lagos to Frankfurt link, and connection reuse removes three of them with a configuration setting.
233
Why a banking engineer should know this layer
interviewer
“Your p99 is 280 ms and 190 ms of it is before your code runs. Where does it go?”
Part 14 covered protocol choice. This part is the layer underneath, and it earns its place for four concrete reasons rather than general interest.
why it matters here specifically
- A large share of latency is below your application. Handshakes, congestion windows and head-of-line blocking do not appear in your traces unless you look for them.
- Part 8 gave card authorisation a 2-second budget. Two round trips of unnecessary handshake is a meaningful fraction of it, and it is entirely avoidable.
- The failure modes are invisible from the application. A stalled HTTP/2 connection, an exhausted flow-control window, a half-open socket: all of them look like "the service is slow".
- It is the most commonly asked layer that engineers cannot explain. Everyone has read that HTTP/2 multiplexes; very few can say what a frame is or what HPACK does to a proxy.
worked numbers
one uncached HTTPS request to a new origin:
DNS 1 RTT
TCP handshake 1 RTT
TLS 1.3 handshake 1 RTT
the actual request 1 RTT
──────
4 RTT before a byte of response
Lagos to a Frankfurt region ≈ 90 ms RTT → 360 ms
same request on a warm connection → 90 ms
connection reuse is worth 270 ms, and it is a config setting.234
The journey of one request, layer by layer
| Layer | Does | Failure looks like |
|---|---|---|
| DNS | Name to address | Intermittent slowness, hard to attribute |
| TCP | Reliable, ordered byte stream | Connection refused, or a 30-second hang |
| TLS | Confidentiality, integrity, identity | Handshake failure, expired certificate |
| HTTP | Request and response semantics | 4xx, 5xx, or a stalled stream |
| Application | Your code | The thing you actually instrumented |
the attribution problem
Your tracing starts when your handler is invoked. Everything above is invisible unless you deliberately instrument it, which is why "the client says it took 400 ms and we recorded 30 ms" is such a common and frustrating conversation. Part 13 asked for tracing across boundaries; the transport layer is the boundary people forget.
one request
every layer it actually crosses
swipe the figure sideways, or tap expand for full screen
1/8
a new origin
A single HTTPS request to an origin the client has not spoken to before. The round trip time is 90 milliseconds, which is roughly Lagos to Frankfurt.
235
TCP: the handshake, and the cost of a round trip
worked numbers
three-way handshake client → SYN "I want to talk, my seq is X" server → SYN-ACK "fine, my seq is Y, I saw X" client → ACK "I saw Y" ← data can ride along one full round trip before any payload. and connection teardown leaves the closer in TIME_WAIT for 2×MSL, typically 60 s, holding the port tuple.
the consequences that bite in production
- Connection reuse is the single largest transport-level win available. Keep-alive, connection pools, and HTTP/2 all exist substantially for this.
- TIME_WAIT exhaustion. A service making thousands of short-lived outbound connections per second runs out of ephemeral ports. The symptom is intermittent connection failures under load and it is diagnosed with
ss -s, not in application logs. - Who closes matters. The side that initiates close holds TIME_WAIT, so a server closing every connection accumulates them rather than the client.
- SYN backlog. Under a connection flood, the accept queue fills and new connections are dropped silently while the application looks idle.
code
# the settings that matter on a high-connection-rate service net.ipv4.tcp_tw_reuse = 1 # reuse TIME_WAIT for outbound net.core.somaxconn = 4096 # accept queue depth net.ipv4.tcp_max_syn_backlog = 8192 # and from Part 14: detect a dead peer rather than hanging TCP_KEEPIDLE=30 TCP_KEEPINTVL=10 TCP_KEEPCNT=3 TCP_NODELAY=true # disable Nagle on request/response
236
Congestion control, and why slow start matters
worked numbers
TCP does not send as fast as it can. it probes. slow start: begin with a small congestion window, double it every round trip until loss or a threshold initcwnd is typically 10 segments ≈ 14 KB RTT 1: 14 KB RTT 2: 28 KB RTT 3: 56 KB RTT 4: 112 KB a 100 KB response needs ~4 RTT to transfer, regardless of bandwidth. on a 90 ms link that is 360 ms of pure protocol behaviour.
what follows from this
- A new connection is slow even on a fast network. Bandwidth is irrelevant for small responses; round trips are the currency.
- Warm connections have a grown window, which is another reason pooling matters so much. A pooled connection that has been transferring is already at a high window.
- Idle connections shrink. Many stacks reset the window after an idle period, so a pooled-but-idle connection behaves like a new one on its next burst.
- Response size matters more than you expect at small sizes. Getting a response under the initial window, roughly 14 KB, means one round trip instead of two.
- BBR versus CUBIC. BBR models bandwidth and RTT rather than reacting to loss, and is usually better on lossy or long-distance links. Worth knowing the names and the distinction.
the bank-specific consequence
Part 8's authorisation response is small and that is a feature. A compact response fits in the initial congestion window and completes in one round trip. A verbose JSON response that crosses 14 KB silently costs an extra round trip, which on a card path with a 2-second scheme budget is a real and completely invisible tax.
237
TLS: the handshake, and what 1.3 removed
worked numbers
TLS 1.2 full handshake: 2 RTT ClientHello → ServerHello, cert, key exchange → client key exchange, change cipher, finished → server change cipher, finished then data TLS 1.3 full handshake: 1 RTT ClientHello with a key share guess → ServerHello, cert, finished then data TLS 1.3 resumption: 0 RTT, with a caveat early data rides on the first flight
what 1.3 changed, and why each matters
- One fewer round trip, by having the client guess the key-exchange group up front. On a 90 ms link that is 90 ms saved on every new connection.
- Removed the weak options. No RSA key transport, no CBC modes, no renegotiation, no compression. A whole generation of attacks became unreachable by deletion rather than by configuration.
- Forward secrecy is mandatory, so a future compromise of the server key does not decrypt past captured traffic.
- Encrypted more of the handshake, including the certificate.
the 0-RTT caveat, which matters for a bank
0-RTT early data is replayable by design. An attacker who captures it can resend it, and the server cannot distinguish the replay. That is acceptable for an idempotent GET and unacceptable for a payment. The rule: never allow 0-RTT on non-idempotent endpoints. Our idempotency keys from Part 2 would actually protect us, and relying on an application control to compensate for a transport weakness is the wrong layering when disabling it costs one round trip.
238
Certificates, chains, and the failures you will meet
the failures, in rough order of frequency
- Expiry. Still the most common outage cause in this list, and entirely preventable. Alert at 30 days, not at 7, and automate renewal.
- Incomplete chain. The server sends the leaf but not the intermediate. Browsers often recover by fetching it; server-to-server clients usually do not, so it works in testing and fails in production.
- Hostname mismatch. The certificate is valid and is not for this name.
- Clock skew. A client with a wrong clock rejects a valid certificate. Rare, and baffling when it happens.
- Root store differences. A certificate trusted by a browser and not by an old container image, or by a JVM with its own trust store.
- Revocation. CRL and OCSP both have availability problems; OCSP stapling moves the fetch to the server and is what you want.
the mutual TLS addition, from Part 14
With mTLS every failure above can happen in both directions, and client certificates are short-lived by design. A rotation that silently fails produces a service that works until the current certificate expires and then stops entirely. Alert on client certificate age as well as server certificate expiry, because the rotation mechanism failing is invisible until the deadline.
239
HTTP/1.1: the bottlenecks, and the workarounds
worked numbers
HTTP/1.1 is a text protocol over one connection at a time. one request → one response → next request consequences: · head-of-line blocking at the application layer · headers repeated in full on every request · pipelining exists, is broken in practice, and is disabled the workaround everyone used: 6 to 8 connections per origin which multiplies handshakes, congestion windows and memory
| Workaround | What it bought | What it cost |
|---|---|---|
| Domain sharding | More parallel connections | More DNS, more handshakes, more windows |
| Concatenation | Fewer requests | Cache invalidation of the whole bundle |
| Sprites | Fewer requests | Awkward, and obsolete |
| Inlining | No extra request | Not cacheable separately |
why this history matters
Every one of those workarounds becomes an anti-pattern under HTTP/2, and a surprising amount of production configuration still contains them. Domain sharding actively hurts HTTP/2 by preventing connection coalescing. Upgrading the protocol without removing the workarounds gets you a fraction of the benefit, and that is a common and invisible waste.
240
HTTP/2: the binary framing layer
The single change from which everything else follows: HTTP/2 stops being text and becomes a binary framing layer over one connection.
worked numbers
every frame: [ 24-bit length ][ 8-bit type ][ 8-bit flags ][ 31-bit stream id ] then the payload 9-byte header, then up to 2^24-1 bytes of payload (default max frame size is 16 KB, negotiable)
| Frame type | Carries | Note |
|---|---|---|
| HEADERS | Request or response headers | HPACK-compressed |
| DATA | The body | Flow-controlled |
| SETTINGS | Connection parameters | Exchanged at start, and any time |
| WINDOW_UPDATE | Flow-control credit | Per stream and per connection |
| RST_STREAM | Cancel one stream | Cancellation without closing the connection |
| GOAWAY | Graceful connection shutdown | Names the last stream it will process |
| PING | Liveness and RTT measurement | Cheap keepalive |
| PRIORITY | Dependency hints | Largely deprecated, see chapter 244 |
why binary framing is the enabling change
Text framing requires reading until a delimiter, which means one message must be fully read before the next begins. Length-prefixed binary frames can be interleaved, because the receiver knows exactly where each one ends and which stream it belongs to. Multiplexing, flow control, cancellation and header compression are all consequences of that one decision.
241
Streams, multiplexing, and stream states
worked numbers
a stream is an independent bidirectional sequence of frames within one connection, identified by a 31-bit id. client-initiated streams are odd: 1, 3, 5, 7… server-initiated are even (push, chapter 244) ids never reuse, and always increase so a connection has a finite lifetime: 2³¹ ÷ 2 ≈ 1.07 billion streams, then it must be replaced rarely hit, and real on a very long-lived high-volume link
the stream lifecycle, simplified
- idle → nothing sent yet.
- open → HEADERS sent, both sides may send.
- half-closed → one side sent END_STREAM; the other may still send. This is the normal state of a request awaiting a response.
- closed → both done, or RST_STREAM.
| Setting | Controls | Typical |
|---|---|---|
MAX_CONCURRENT_STREAMS | How many at once | 100 to 250. The one that bites |
INITIAL_WINDOW_SIZE | Per-stream flow control | 64 KB default, often too small |
MAX_FRAME_SIZE | Largest frame | 16 KB default |
MAX_HEADER_LIST_SIZE | Header size limit | Prevents header bombs |
the setting that causes silent queueing
MAX_CONCURRENT_STREAMS is announced by the server and enforced by the client. If it is 100 and your client wants 300 concurrent requests, 200 of them queue in the client, invisibly, before any request is sent. The symptom is client-side latency with no matching server-side latency, which is exactly the "we recorded 30 ms and the client saw 400" conversation from chapter 234.242
HPACK: how headers get compressed
worked numbers
HTTP/1.1 repeated every header on every request:
cookies, user-agent, accept, authorization…
often 500 to 800 bytes, on every single request
HPACK uses two tables plus Huffman coding:
static table: 61 predefined common headers
:method GET = index 2 → one byte
dynamic table: headers seen on this connection
an authorization header sent once becomes an index
a repeated request can compress to under 20 bytes of headers.the properties that follow, and matter
- It is stateful per connection. Both ends maintain synchronised tables, so losing sync breaks the connection rather than one request.
- The compression context is shared across streams, which is what makes it so effective and also means header processing is inherently serialised at the HPACK layer.
- It was designed to resist CRIME. Compression of attacker-influenced data alongside secrets leaks information; HPACK avoids this by never compressing across a security boundary the way generic compression did.
- Proxies must re-encode. An intermediary cannot simply forward HPACK-encoded headers, because the table state differs per connection. This is why HTTP/2 proxies are more expensive than HTTP/1.1 proxies, which is a real capacity consideration.
243
Flow control, and the window nobody tunes
worked numbers
HTTP/2 has its own flow control, on top of TCP's. two levels: per stream and per connection default window: 65,535 bytes each a sender may not send more DATA than the window allows until a WINDOW_UPDATE grants more credit. 64 KB on a 90 ms RTT link caps one stream at 65,535 ÷ 0.09 ≈ 728 KB/s, regardless of bandwidth
why it exists, and why the default is wrong for many workloads
- It exists so one stream cannot starve another within a connection, and so a receiver can push back on a fast sender without closing the connection.
- The default 64 KB was chosen conservatively and is far too small for high bandwidth-delay-product links.
- The symptom is throughput that ignores bandwidth: a large response transfers at a fixed rate that does not improve on a faster network.
- Tuning is a SETTINGS frame plus prompt WINDOW_UPDATEs. Most servers allow raising the initial window; many deployments never do.
- It is also the backpressure mechanism, which connects directly to Part 14 chapter 158: a receiver that stops sending WINDOW_UPDATE is applying backpressure correctly.
where this matters for us
Large responses over long links: a statement export, a reconciliation file pull, a bulk payment response. Small request-response traffic never touches the window, which is why most teams never discover it, and why the one team pulling a 200 MB settlement file over HTTP/2 finds their transfer inexplicably capped.
244
Server push, and why it was abandoned
worked numbers
the idea: the server sends a resource the client has not yet asked for, anticipating the need. PUSH_PROMISE on an even stream id, then the data it was removed from Chrome in 2022 and is effectively dead.
why it failed, which is a useful lesson in itself
- The server cannot know the client cache. Pushing something already cached wastes bandwidth, and the cancellation arrives too late to help.
- It competed with the response for the same connection and the same congestion window, frequently making the critical resource slower.
- Implementation complexity was high and the measured benefit was small or negative in most real deployments.
- The problem it solved has a better answer:
103 Early Hintsand preload hints let the client decide, which respects its cache.
the transferable lesson
A mechanism where the sender guesses what the receiver needs, without knowing the receiver's state, tends to lose to one where the receiver decides. It is the same reasoning as Part 3's cache watermark: the client knows what it has already seen, and designs that make correctness the receiver's responsibility outperform designs that rely on the sender guessing correctly.
245
The HTTP/2 failure modes in production
| Symptom | Cause | How to confirm |
|---|---|---|
| Client latency, no server latency | MAX_CONCURRENT_STREAMS reached, requests queued client-side | Client-side stream queue metrics |
| Throughput caps regardless of bandwidth | Flow-control window too small | Window size in SETTINGS; transfer rate ≈ window ÷ RTT |
| Everything stalls at once | TCP-level loss, all streams blocked | Packet loss on the path |
| Uneven load across backends | One connection pinned to one backend | Per-backend request counts |
| Intermittent GOAWAY | Server recycling connections by age or stream count | Connection age at close |
| Works in test, fails through a proxy | Middlebox mangling HTTP/2 or trailers | Test with the proxy in path |
the load-balancing surprise
HTTP/2 uses one long-lived connection, so a client pinned to one backend stays pinned. With HTTP/1.1 and six connections, load spread naturally; with HTTP/2 a small number of clients can concentrate on a subset of backends. The fix is connection-level load balancing that periodically issues GOAWAY to force rebalancing, and it is a genuine operational difference people meet after migrating.
246
HTTP/3 and QUIC: reliability above an unordered transport
worked numbers
the insight: TCP's ordering guarantee is the problem. HTTP/2 removed head-of-line blocking at the application layer, and TCP still delivers bytes in order, so one lost packet stalls every stream until it is retransmitted. QUIC: build reliability per stream, above UDP · UDP provides no ordering, which is the point · QUIC adds ordering and retransmission per stream · a lost packet affects only its own stream this is why Part 14 said UDP is used without being chosen.
what QUIC also brings
- Handshake in 1 RTT, or 0 for resumption, because TLS 1.3 is integrated rather than layered on top.
- Encrypted transport headers, so middleboxes cannot inspect or mangle them. Good for integrity, awkward for network operators.
- Connection migration, via a connection id rather than the four-tuple. Chapter 247.
- Faster evolution. Because it is in userspace rather than the kernel, changes ship with the application rather than with the OS.
the honest costs
- Higher CPU. Userspace processing rather than kernel TCP, though this is narrowing with offload.
- UDP is blocked or deprioritised on some networks, so a fallback to HTTP/2 is mandatory.
- Less mature tooling. Debugging is harder and fewer engineers can do it.
- Encrypted headers mean less network visibility, which some enterprise environments actively refuse.
247
Connection migration, and the mobile case
worked numbers
TCP identifies a connection by the four-tuple: (source IP, source port, dest IP, dest port) phone moves from wifi to mobile data → source IP changes → the connection is dead. reconnect, rehandshake. QUIC identifies by a connection id, carried in the packet → the address changes, the connection continues → no handshake, no interruption
why this matters for a Nigerian bank specifically
Mobile-first users on unreliable connectivity switch networks constantly, and every switch under TCP is a full reconnection: DNS, TCP handshake, TLS handshake, then the request. On a 90 ms link that is nearly 400 ms of dead time, repeatedly, during a session. Connection migration removes it entirely, which is a larger real-world win for a mobile banking app than most backend optimisations, and it is a configuration change at the edge rather than an application change.
248
Head-of-line blocking, traced through all three versions
| Version | Application layer | Transport layer |
|---|---|---|
| HTTP/1.1 | Blocked. One request at a time | Blocked |
| HTTP/2 | Solved. Streams interleave | Still blocked. TCP delivers in order |
| HTTP/3 | Solved | Solved. Per-stream reliability |
the answer to give when asked
"HTTP/2 removes head-of-line blocking at the application layer. TCP still guarantees in-order delivery, so a single lost packet stalls every stream on that connection until it is retransmitted. That is exactly why HTTP/3 moved to QUIC over UDP: it reimplements reliability per stream above an unordered transport, so a loss affects only the stream it belonged to. Saying HTTP/2 fixes head-of-line blocking is the half answer; naming which layer it fixes and which it does not is the full one."
head-of-line blocking
the same problem, pushed down one layer at a time
swipe the figure sideways, or tap expand for full screen
1/8
four requests
Four requests, the first of which is slow because it builds a statement.
249
Choosing a version, per hop
| Hop | Version | Why |
|---|---|---|
| Mobile app → edge | HTTP/3, falling back to /2 | Connection migration and loss resilience matter most here |
| Browser → edge | HTTP/2 or /3 | Either. The CDN decides |
| Edge → origin | HTTP/2 | Stable low-loss network; /3 adds CPU for little gain |
| Service → service | HTTP/2 (gRPC) | Multiplexing and typed contracts, per Part 14 |
| Us → payment provider | Whatever they support | Frequently HTTP/1.1, and not our choice |
| Provider webhooks → us | HTTP/1.1 | Their client, their decision. Support it |
| Card scheme | Not HTTP at all | ISO 8583 over persistent TCP, per Part 14 |
the point worth making
The version is a per-hop decision, not an estate-wide one. The hop where it matters most is the one you control least: a customer on a congested mobile network, where loss resilience and connection migration are worth far more than anything you can do inside the datacentre. Inside the datacentre, loss is rare and HTTP/2 is sufficient, and spending CPU on QUIC there is optimising a problem you do not have.
250
Debugging the transport layer
code
# which version was actually negotiated? ALPN, not assumption.
curl -sI --http2 https://api.example/health -w '%{http_version}\n'
# the full timing breakdown. this answers "where did the 400ms go".
curl -o /dev/null -s -w 'dns:%{time_namelookup} tcp:%{time_connect} \
tls:%{time_appconnect} ttfb:%{time_starttransfer} total:%{time_total}\n' \
https://api.example/health
# connection state: TIME_WAIT accumulation, SYN backlog drops
ss -s
netstat -s | grep -i 'listen\|overflow\|retrans'
# TLS chain, including whether the intermediate is actually sent
openssl s_client -connect api.example:443 -showcerts </dev/null
# HTTP/2 frame-level, when you genuinely need it
nghttp -nv https://api.example/healththe sequence when latency is unexplained
- Confirm the negotiated version. An ALPN misconfiguration silently falls back to HTTP/1.1 and nobody notices for months.
- Get the timing breakdown. It immediately separates DNS, connect, TLS and server time, and usually ends the investigation.
- Check connection reuse. If every request pays a handshake, the pool is not working, which is a configuration bug and the largest single win available.
- Check for retransmissions. Loss means HTTP/2 stalls across all streams, and it points at the network rather than the application.
- Check stream concurrency limits, per chapter 241, if client latency exceeds server latency.
the instrumentation to add rather than debug repeatedly
Emit connection reuse rate, negotiated protocol version and TLS handshake count as metrics. All three are cheap, low-cardinality, and turn a recurring mystery into a graph. Part 13 argued for symptom-based alerting; connection reuse falling is a symptom that precedes a latency incident, and almost nobody measures it.
251
What this means for the bank, concretely
the decisions this part actually changes
- HTTP/3 at the edge for the mobile app, with fallback. Connection migration is worth more to a Nigerian mobile user than most backend work, per chapter 247.
- Connection pooling everywhere, with reuse rate as a monitored metric. Chapter 235: it is worth 270 ms per request on a cold path.
- Keep the authorisation response small, under the initial congestion window, per chapter 236. An extra round trip inside Part 8's budget is expensive and invisible.
- TLS 1.3, and 0-RTT disabled on anything non-idempotent, per chapter 237. Our idempotency keys would cover it, and the right layering is not to need them to.
- Raise the HTTP/2 flow-control window on any path carrying large files: statements, settlement files, bulk exports.
- Certificate expiry alerting at 30 days, both directions, including mTLS client certificates whose rotation can fail silently.
- Periodic GOAWAY for rebalancing, per chapter 245, so long-lived HTTP/2 connections do not pin traffic to a subset of backends.
worked numbers
the summary a staff engineer should be able to give: round trips are the currency, not bandwidth connection reuse is the largest single win HTTP/2 fixed the application layer, not TCP HTTP/3 fixed the transport layer, at a CPU cost the version is a per-hop decision and most of the latency you cannot explain is here.
how to close this part
"Part 14 chose protocols by conversation shape. This part is the layer underneath, and it matters because a meaningful share of p99 happens before application code runs and is invisible to application tracing. The three things I would act on: HTTP/3 at the mobile edge for connection migration, connection reuse measured as a metric because it is worth hundreds of milliseconds and fails silently, and 0-RTT disabled on payment endpoints because early data is replayable by design."