What is a context switch and why is it expensive?

Difficulty: Intermediate

Question

What happens during a context switch? Why is it considered overhead, and how does process switching differ from thread switching?

Answer

A context switch is the mechanism that makes multitasking possible, and understanding its cost explains a lot of other design decisions, from why threads exist to why too many processes slow a machine down. Picture a chef who is cooking several dishes for different customers. Before turning to a new dish, the chef must note down exactly where the previous dish stood (temperatures, timers, what has been added), and then recall the state of the new dish. All that note-keeping is pure overhead; no cooking happens during it. That is a context switch.

Technically, the context of a process is everything the CPU needs to resume it exactly where it left off: the program counter, CPU registers, stack pointer, status flags, and the memory management information such as page table base register. When the scheduler decides to switch from process A to B, typically because A's time slice expired, it blocked on I/O, or a higher priority process became ready, the sequence is roughly this. A timer interrupt or system call enters kernel mode. The kernel saves A's CPU state into A's PCB. The scheduler selects B. The dispatcher loads B's state from B's PCB, switches the address space by changing the page table register, and returns to user mode at B's saved program counter. The component that performs the switch is called the dispatcher, and the time it takes is the dispatch latency.

There are direct and indirect costs. Direct costs are saving and restoring registers, running the scheduler code, and switching the page table, which typically takes a few microseconds. The indirect costs are often larger and less obvious. The TLB caches virtual-to-physical translations for the old address space, so after a process switch many entries are invalid or must be flushed (unless the hardware tags entries with an address space ID). The CPU caches, which held B's data before it was switched out, are now full of A's data, so B will suffer a burst of cache misses until it warms up again. Branch predictors also lose their trained state. So a switch that took two microseconds directly can cost much more in lost performance afterward.

That is why switching between threads of the same process is cheaper than switching between processes. Threads share the same address space, so the page table register does not change, the TLB and much of the cache stays useful, and only registers and stack need swapping. It is also why user-level threads, which avoid entering the kernel at all, are cheaper still.

Practical consequences follow. A tiny time quantum in Round Robin gives responsiveness but wastes more time on switches; a huge quantum reduces overhead but degrades into FCFS. Creating thousands of threads for CPU-bound work can make things slower, because the machine spends its time switching. Event-driven designs such as those in Node.js and nginx avoid switches by handling many connections in one thread.

Also mention the difference between a mode switch and a context switch. A system call changes from user to kernel mode within the same process and does not necessarily switch to another process, so it is cheaper. Interviewers sometimes trap candidates by asking whether a system call always causes a context switch; the answer is no.

Key points

Concepts covered

context switch, PCB, TLB flush, cache, dispatcher