Difficulty: Advanced
What is the difference between user-level and kernel-level threads? Explain the many-to-one, one-to-one and many-to-many models.
Once you know what a thread is, the next natural question is who actually manages it: the application or the kernel. This is a slightly more advanced topic, so let me build it up with an analogy. Imagine a company where employees (threads) work for a department (process). In one setup, the department manager privately schedules who works on what, and the company's HR (the kernel) does not even know individual employees exist; HR only knows the department. In another setup, HR knows every employee and schedules each directly. The first is user-level threading; the second is kernel-level threading.
User-level threads are created and managed entirely by a library in user space, such as early Green threads or a coroutine-style runtime. The kernel sees just one schedulable entity, the process. Switching between user threads is very fast because it needs no system call or mode switch; the library just swaps a few registers and stack pointers. They are portable and you can tailor the scheduling policy. The big weakness is blocking: if one user thread makes a blocking system call, like reading a file, the kernel blocks the entire process, so all other threads stall too. They also cannot run in parallel on multiple cores because the kernel schedules only one entity per process.
Kernel-level threads are known to and scheduled by the OS itself. Linux (via clone), Windows and macOS all work this way. If one thread blocks on I/O, the kernel simply schedules another thread from the same process, and multiple threads can run truly in parallel on different cores. The cost is that creation, switching and synchronization involve the kernel, which makes them slower than user threads, and the kernel must maintain per-thread data structures.
The mapping between them gives the three classic models. In many-to-one, many user threads map onto a single kernel thread. It is cheap and simple, but one blocking call blocks everything and there is no multicore parallelism. In one-to-one, each user thread maps to its own kernel thread. This gives full concurrency and parallelism and is what Linux pthreads and Windows use, at the cost of creating one kernel thread per user thread, which limits how many threads you can sensibly create. In many-to-many, M user threads are multiplexed onto N kernel threads (N less than or equal to M). It gets the best of both worlds in principle, allowing many cheap threads and real parallelism, but it is complex to implement, and the scheduling coordination between library and kernel is tricky. A two-level model is a variation that also allows binding a user thread to a specific kernel thread.
Modern examples show that the old debate lives on in new clothes. Go goroutines and Java's virtual threads (Project Loom) are user-level constructs multiplexed onto a smaller pool of kernel threads, essentially a many-to-many design, letting you spawn hundreds of thousands of lightweight tasks. When one blocks on I/O, the runtime parks it and runs another on the same kernel thread.
Summarize the trade-off like this: user-level threads are cheap but kernel-blind, kernel-level threads are heavier but cooperate with the scheduler and blocking behaviour, and hybrid runtimes try to combine both.
user threads, kernel threads, many-to-one, one-to-one, many-to-many