Part 3 · 1 chapters · ~8 min
System Calls
The user and kernel boundary, how a syscall works (registers, trap instruction, return), errno and checking results, measured costs, libc wrappers versus raw syscalls, strace and dtruss, vDSO for cheap time calls, and batching to reduce syscalls (buffering, writev, io_uring).
4
Crossing into the kernel
code
// every syscall can fail: check the result and errno
ssize_t n = write(fd, buf, len);
if (n < 0) { if (errno == EINTR) goto retry; perror("write"); return -1; }
if ((size_t)n < len) { /* partial write: loop until everything is written */ }
// measured on an M3 Pro (clang 17 -O2): getppid 102 ns · write(1 byte) 512 ns · fputc 18.2 ns
// trace syscalls of any program:
strace -c -f node server.js # Linux: count syscalls by type
sudo dtruss -c ./money # macOS equivalent (requires disabling SIP for many binaries)
// fewer syscalls: gather several buffers into one write
struct iovec iov[2] = { { hdr, hdr_len }, { body, body_len } };
writev(fd, iov, 2);How it works on x86-64 Linux: the syscall number goes in rax, arguments in rdi, rsi, rdx, r10, r8, r9; the syscall instruction enters the kernel, which dispatches through a table and returns a result (negative values become errno in libc). Calls like clock_gettime often avoid the kernel entirely through the vDSO, a small kernel-provided library mapped into every process. io_uring (Linux) submits many operations through shared rings with few or no syscalls.
WHAT A SYSCALL COSTS, MEASURED
Apple M3 Pro, clang 17 -O2, 1,000,000 iterations each
swipe the figure sideways, or tap expand for full screen
1/4
the boundary
A syscall switches from user mode to kernel mode and back. Even the trivial getppid() cost 102 ns here, hundreds of times a function call.
mode switch: ~100 ns minimummeasured getppid