Explain kernel mode vs user mode and how system calls work

Difficulty: Beginner

Question

What is the difference between kernel mode and user mode? What is a system call and how does it work?

Answer

This is the question where interviewers check whether you understand how an OS protects itself from buggy or malicious programs. The core idea is simple: if any program could execute any instruction, one bad pointer in a browser could overwrite the disk driver or halt the machine. So the CPU itself supports at least two privilege levels, and the OS uses them to build a wall between its own code and everyone else's.

In user mode, the CPU refuses to execute privileged instructions such as those that talk directly to I/O ports, modify page tables, disable interrupts or halt the processor. Code also cannot access kernel memory. In kernel mode (also called supervisor mode, ring 0 on x86), everything is allowed. A mode bit in a CPU status register tracks the current mode. Your applications, from a Python script to Chrome, run in user mode; the scheduler, memory manager and device drivers run in kernel mode. If user code tries a privileged instruction, the CPU raises an exception and the OS typically kills the process, which is what a segmentation fault essentially is.

But user programs obviously need kernel services: reading files, allocating memory, creating processes, sending network packets. The controlled doorway into kernel mode is the system call. Think of it as a bank teller window: you cannot walk into the vault, but you can hand a slip through the window, and a trusted employee does the work and hands the result back. The system call interface is that window.

Here is the flow when a program calls read(). The C library wrapper places the system call number and arguments into registers (for example eax on 32-bit x86 or rax on x86-64 Linux). It then executes a special trap instruction: int 0x80, or the faster syscall/sysenter instructions. The CPU switches to kernel mode, saves the user program counter and stack pointer, and jumps to a fixed entry point in the kernel. The kernel uses the syscall number to index a system call table, validates the arguments (never trusting user pointers), performs the work, places the return value in a register, and executes a return-from-trap instruction that restores user mode and resumes the program right after the call.

Some subtleties are worth mentioning. A system call is not the same as a library function: printf is a library function that may eventually call the write system call, whereas malloc may or may not call brk or mmap depending on the size. A trap is a synchronous, software-generated entry into the kernel, whereas an interrupt is asynchronous and comes from hardware, such as a timer or keyboard. Exceptions like division by zero or page faults are also synchronous and enter kernel mode. All three use the same basic mechanism of saving state and jumping to a handler.

System calls fall into categories: process control (fork, exec, exit, wait), file management (open, read, write, close), device management (ioctl), information maintenance (getpid, time) and communication (pipe, socket, shmget). Because crossing the boundary costs time, well-written programs batch their I/O, for instance by using buffered streams instead of writing one byte at a time.

Code examples

Library call versus raw system call on Linux

#include <unistd.h>
#include <sys/syscall.h>
#include <stdio.h>

int main(void) {
    const char *msg = "hello via syscall\n";
    /* Both lines end up in the kernel's write system call */
    write(1, msg, 18);                 /* libc wrapper */
    syscall(SYS_write, 1, msg, 18);    /* direct syscall number */
    printf("pid = %d\n", (int)syscall(SYS_getpid));
    return 0;
}

The wrapper and the raw syscall() both load the syscall number into a register and execute the trap instruction, moving the CPU into kernel mode.

Key points

Concepts covered

kernel mode, user mode, system call, trap, privileged instructions