Part 1 · 1 chapters · ~8 min

The System Call Path

From libc to the syscall instruction, entry code and pt_regs, the syscall table, SYSCALL_DEFINE macros, copy_to_user and copy_from_user, returning and signal delivery, the vDSO, seccomp filters on the path, Spectre and Meltdown mitigations and their cost, and tracing with strace and perf trace.

2

Following read() through the kernel

code
// fs/read_write.c (simplified)
SYSCALL_DEFINE3(read, unsigned int, fd, char __user *, buf, size_t, count)
{
	return ksys_read(fd, buf, count);
}
ssize_t vfs_read(struct file *file, char __user *buf, size_t count, loff_t *pos)
{
	if (!(file->f_mode & FMODE_READ)) return -EBADF;
	...
	if (file->f_op->read_iter) ret = new_sync_read(file, buf, count, pos);
	...
}
// the __user annotation marks pointers into user space: they must go through copy_to_user/copy_from_user
code
strace -T -e trace=read cat /etc/hostname     # each read with the time spent in it
perf trace -s ls                              # syscall summary with latencies (needs perf)

Mitigation cost: after Meltdown (2018), kernel page-table isolation made each kernel entry switch page tables, which raised syscall cost on affected CPUs; this is one reason measured syscall costs vary by CPU and kernel. C course P3 measured 102 to 512 ns per syscall on macOS on Apple silicon.

A READ() SYSCALL, ENTRY TO RETURN
x86-64 Linux
user (libc read)entry_SYSCALL_64sys_call_tableksys_read → vfs_readfile system / page cachesyscall (rax = 0, rdi = fd, rsi = buf, rdx = count)save registers, switch to kernel stack
swipe the figure sideways, or tap expand for full screen
1/4
entry
The syscall instruction jumps to entry_SYSCALL_64 with the syscall number in rax. The kernel saves user registers and switches to a per-thread kernel stack.
trap into the kernelregisters saved