Table of Contents
- 0. From Language-Level Execution Units to the CPU
-
1. How Does Java Schedule Threads?
- 1.1 From Platform Threads to Virtual Threads
- 1.1.1 The Simplest Model: One Java Thread per OS Thread
- 1.1.2 The Problem: Resource and Scheduling Costs of Many OS Threads
- 1.1.3 The Solution: Let the JVM Schedule Lightweight Threads
- 1.1.4 The Result: Platform Threads and Virtual Threads
- 1.2 Execution Layers
- 1.3 From Java Thread Lifecycles to Their Implementation
- 1.3.1 Creation and Start: Created → Runnable
- 1.3.2 Waiting to Be Scheduled: Runnable
- 1.3.3 Running: Runnable → Running
- 1.3.4 Parking and Waiting: Running → Waiting
- 1.3.5 Wakeup and Resumption: Waiting → Runnable → Running
- 1.3.6 Completion and Cleanup: Running → Terminated
- 1.4 End-to-End Sequence Diagrams
- 1.5 Platform Threads vs. Virtual Threads
- 1.6 Returning to the Opening Question
-
2. How Does Go Schedule Goroutines?
- 2.1 From G-M to G-M-P
- 2.1.1 The Simplest Model: A Goroutine Runs on an OS Thread
- 2.1.2 The Problem: Many Goroutines, Fewer OS Threads
- 2.1.3 The Solution: Evolving from G-M to G-M-P
- 2.1.4 The Final Model: G, M, and P
- 2.2 Execution Layers
- 2.3 From a Goroutine's Lifecycle to Its Runtime Implementation
- 2.3.1 Creation: Created → Runnable
- 2.3.2 Waiting for Scheduling: Runnable
- 2.3.3 Execution: Runnable → Running
- 2.3.4 Parking and Waiting: Running → Waiting
- 2.3.5 Wakeup and Resumption: Waiting → Runnable → Running
- 2.3.6 System Calls and Preemption
- 2.3.7 Completion and Cleanup: Running → Dead
- 2.4 End-to-End Sequence Diagram
- 2.5 Returning to the Opening Question
-
3. How Does CPython Schedule Threads?
- 3.1 From Native OS Threads to the GIL and Free-Threaded Builds
- 3.1.1 The Simplest Model: One Python Thread per OS Thread
- 3.1.2 The Problem: Sharing Interpreter State Across OS Threads
- 3.1.3 The Default Solution: The GIL Restricts Interpreter Execution
- 3.1.4 The Next Step: Free-Threaded CPython
- 3.1.5 The Final Model: Thread, PyThreadState, and GIL
- 3.2 Execution Layers
- 3.3 From CPython Thread Lifecycles to Their Implementation
- 3.3.1 Creation and Start: Created → Runnable
- 3.3.2 Waiting for Scheduling: Runnable
- 3.3.3 Execution: Runnable → Running
- 3.3.4 Parking and Waiting: Running → Waiting
- 3.3.5 Wakeup and Resumption: Waiting → Runnable → Running
- 3.3.6 Completion and Cleanup: Running → Terminated
- 3.4 End-to-End Sequence Diagram
- 3.5 Returning to the Opening Question
- 4. What the Three Implementations Have in Common
- 5. Next: Language Memory Models
0. From Language-Level Execution Units to the CPU
The previous article examined how the CPU executes count++ and how atomic instructions, cache coherence, and fences provide the hardware foundations for atomicity, visibility, and ordering.
This article moves one layer up, to the operating system. It asks one central question:
How do Java threads, Go goroutines, and CPython threads pass through their runtimes and the operating system to execute on a CPU?
Answering it requires understanding three things:
- Execution mapping: How does a language-level execution unit map to an OS thread?
- Scheduling responsibilities: What do the runtime scheduler and the operating system scheduler each control?
- Waiting and resumption: When an execution unit waits, wakes, or resumes, what is suspended, and which threads can continue running?
Next, we will examine Java, Go, and CPython in turn. For each, we will trace the evolution of its threading model, then use execution layers, lifecycles, and sequence diagrams to answer these three questions. Finally, we will draw out the language-independent mechanisms they share and return to the opening question.
1. How Does Java Schedule Threads?
1.1 From Platform Threads to Virtual Threads
1.1.1 The Simplest Model: One Java Thread per OS Thread
When a program has only a few concurrent tasks, the most straightforward approach is to map each Java thread to an OS thread. Java creates the thread, and the operating system schedules the corresponding OS thread.
With relatively few threads, this one-to-one mapping is easy to understand and debug.
1.1.2 The Problem: Resource and Scheduling Costs of Many OS Threads
Suppose a service handles many requests, each on its own thread:
Request 1 → Thread 1 → OS Thread 1
Request 2 → Thread 2 → OS Thread 2
Request 3 → Thread 3 → OS Thread 3
...
Request N → Thread N → OS Thread N
Many request-handling tasks spend most of their time waiting:
Read from a database
Wait for a network response
Wait for a message
Wait for other I/O
If Java threads continue to map one-to-one to OS threads, two costs become significant:
- Native-thread resources: Even while a task is waiting, the operating system must retain resources such as its kernel thread and native stack. These costs grow with the thread count.
- Scheduling overhead: When many threads wake frequently and become Runnable, the operating system must schedule among more runnable threads. Context switches and scheduler work consume additional CPU time.
1.1.3 The Solution: Let the JVM Schedule Lightweight Threads
How can we address these costs?
In this setting, logical tasks spend most of their time waiting for I/O and relatively little time using the CPU. They do not need to occupy a dedicated OS thread for their entire lifetime. The key is to decouple logical execution units from OS threads.
Once they are decoupled, the JVM can introduce its own scheduler:
Logical Thread A ─┐
Logical Thread B ─┼─→ JVM Scheduler ─→ Carrier Thread 1 ⇔ OS Thread 1
Logical Thread C ─┤ └→ Carrier Thread 2 ⇔ OS Thread 2
Logical Thread D ─┘
A Carrier Thread is a Platform Thread in the carrier role. The ⇔ symbol indicates its one-to-one native mapping to an OS thread, not a method call.
The JVM can manage many logical threads while the operating system schedules a smaller number of underlying OS threads.
Java calls this lightweight logical thread a Virtual Thread.
1.1.4 The Result: Platform Threads and Virtual Threads
A Platform Thread maps one-to-one to an OS Thread; the JVM Scheduler dispatches Virtual Threads onto Platform Threads acting as carriers, and the OS Scheduler schedules the underlying OS Threads.
1.2 Execution Layers
1.3 From Java Thread Lifecycles to Their Implementation
A Java thread's conceptual lifecycle is:
This is a conceptual lifecycle, not an exact mapping to Java's Thread.State. In particular, RUNNABLE covers both threads awaiting CPU time and threads currently running.
1.3.1 Creation and Start: Created → Runnable
A platform thread creates an OS thread at start(); a virtual thread submits its continuation to the JVM scheduler.
(1) Platform Thread
(2) Virtual Thread
1.3.2 Waiting to Be Scheduled: Runnable
A platform thread waits for OS scheduling; a virtual thread waits for JVM dispatch to a carrier thread.
(1) Platform Thread
(2) Virtual Thread
1.3.3 Running: Runnable → Running
A platform thread executes Thread.run() on its OS thread; a virtual thread mounts and runs its continuation on a carrier.
(1) Platform Thread
(2) Virtual Thread
1.3.4 Parking and Waiting: Running → Waiting
A waiting platform thread normally retains its OS thread; when unmounting is possible, a virtual thread releases its carrier.
(1) Platform Thread
(2) Virtual Thread
1.3.5 Wakeup and Resumption: Waiting → Runnable → Running
A platform thread resumes on the same OS thread; a virtual thread reenters the JVM scheduler and may resume on a different carrier.
(1) Platform Thread
(2) Virtual Thread
1.3.6 Completion and Cleanup: Running → Terminated
A platform thread cleans up its JavaThread and OS thread; a virtual thread terminates without terminating its carrier.
(1) Platform Thread
(2) Virtual Thread
1.4 End-to-End Sequence Diagrams
The following diagrams trace the native startup path for a platform thread and the essential mount, unmount, and resume operations for a virtual thread. The virtual-thread sequence assumes its carrier already has CPU time, so it does not repeat OS scheduling.
1.4.1 Platform Thread
1.4.2 Virtual Thread
1.5 Platform Threads vs. Virtual Threads
Their lifecycle differences can be summarized as follows:
| Lifecycle phase | Platform Thread | Virtual Thread |
|---|---|---|
| Created → Runnable | The JVM creates an OS thread | The JVM submits a continuation to its scheduler |
| Runnable | Waits for the OS scheduler | The JVM scheduler dispatches it to a carrier; the OS scheduler schedules the carrier's OS thread |
| Running | The OS thread executes Java code directly | The virtual thread mounts on a Carrier Thread backed by an OS Thread |
| Waiting | The OS thread generally waits as well | The virtual thread can wait independently and release its carrier if it can unmount |
| Resume | The OS thread becomes runnable again | The virtual thread is resubmitted to the JVM scheduler |
| Terminated | JavaThread / OS thread exit | Its continuation completes and afterDone() moves it to TERMINATED |
The central distinction is the execution relationship: platform threads remain associated with their OS threads, while the JVM dispatches virtual threads to carriers. When a virtual thread can unmount, its carrier becomes available for other tasks.
For implementation details, see OpenJDK's Thread.java, jvm.cpp, javaThread.cpp, and VirtualThread.java. For the virtual-thread design, see JEP 444.
1.6 Returning to the Opening Question
As Section 1.2, Execution Layers shows, a Platform Thread maps one-to-one to an OS Thread, while a Virtual Thread reaches an OS Thread through a Carrier Thread. Therefore, both thread models ultimately execute on OS Threads.
Section 1.3.2, Waiting to Be Scheduled and Section 1.3.3, Running show that the JVM Scheduler dispatches Virtual Threads to Carrier Threads, while the OS Scheduler schedules the underlying OS Threads. Therefore, the scheduling paths differ, but the OS Scheduler ultimately grants CPU execution time to both.
Section 1.3.4, Parking and Waiting and Section 1.3.5, Wakeup and Resumption show that a Virtual Thread can release its Carrier Thread when waiting permits unmounting, then reenter JVM scheduling when ready. Therefore, many waiting tasks need not each occupy an OS Thread for their entire lifetime.
Answering the opening question: Java retains the native-thread model of Platform Threads while allowing Virtual Threads to reuse Carrier Threads, so many concurrent tasks can ultimately execute on the CPU using a relatively small number of OS Threads.
2. How Does Go Schedule Goroutines?
2.1 From G-M to G-M-P
2.1.1 The Simplest Model: A Goroutine Runs on an OS Thread
For a small number of concurrent tasks, the simplest approach is:
Goroutine
↓
OS Thread
↓
Operating System Scheduler
↓
CPU Core
One Goroutine maps to one OS Thread; the Go runtime manages the Goroutine, while the operating system schedules its OS Thread.
At modest thread counts, this one-to-one model is easy to understand and debug.
2.1.2 The Problem: Many Goroutines, Fewer OS Threads
Suppose a service handles many requests, each on its own Goroutine:
Request 1 → Goroutine 1 → OS Thread 1
Request 2 → Goroutine 2 → OS Thread 2
Request 3 → Goroutine 3 → OS Thread 3
...
Request N → Goroutine N → OS Thread N
Many service tasks spend most of their time waiting for databases, network responses, or other I/O.
If Goroutines remained mapped one-to-one to OS Threads, they would face the same two costs as Java Platform Threads:
- Native-thread resources: a waiting task would still retain a kernel thread, native stack, and other resources.
- Scheduling overhead: many OS Threads repeatedly becoming runnable would increase context switching and OS scheduling work.
2.1.3 The Solution: Evolving from G-M to G-M-P
How can the runtime address this?
These logical tasks spend much of their time waiting for I/O rather than using the CPU. They therefore do not need to retain a dedicated OS Thread for their entire lifetime.
The solution is to decouple the logical execution unit from the OS Thread: the Go runtime schedules Goroutines so a limited number of OS Threads can be reused.
2.1.3.1 Step One: G-M with a Global Run Queue
Start with G, M, and a global run queue:
Here, G is a goroutine and M is an OS thread.
Here, runnable Gs share a global run queue, and multiple Ms fetch work from it.
As concurrency grows, two problems emerge:
Multiple M instances
↓
Access the same Global Run Queue
↓
More synchronization contention
And:
Scheduling state for G is centralized
↓
Poorer locality
To reduce contention on the global queue, some scheduling state and runnable work can be distributed across execution contexts:
What if each execution context had its own scheduling resources and local run queue?
2.1.3.2 Step Two: Introduce P and Local Run Queues
To reduce contention for the global queue, the Go runtime introduces P, with each P maintaining its own local run queue:
P holds the runtime resources needed to execute Go user code and maintains its own local run queue. The actual G-M-P scheduler also retains a global queue; goroutines are not permanently assigned to a particular P. We will examine other work sources next.
This gives us the defining constraint of G-M-P:
An M needs a P to execute Go user code.
Scheduling state is no longer concentrated in one queue contested by all M instances; each P maintains its own local run queue.
2.1.3.3 Third Question: What If a P's Local Queue Is Empty?
Local queues reduce contention but can leave work unevenly distributed: one P is busy while another has nothing queued. An M holding an idle P can use work stealing to obtain runnable goroutines from another P's queue.
The scheduler can also find work from other sources:
In the actual implementation, findRunnable() considers fairness, timers, the network poller, and other conditions when searching for runnable Gs. It does not rigidly follow the same sequence every time.
2.1.3.4 Fourth Question: What If M Blocks in a System Call?
Another constraint comes from blocking system calls:
If an M holding a P blocks in the OS, must the P become idle too?
If P remains permanently attached to that M, other goroutines cannot run using its scheduling resources while the M is blocked.
P and M therefore need to be separable:
G0 remains blocked in the system call on M0's OS thread. Meanwhile the runtime can reclaim P0 so M1 can run another runnable G. When M0 returns from the system call, it must try to acquire a P again before resuming Go user code. P and M must therefore be free to separate and recombine.
2.1.3.5 Fifth Question: What If One G Runs for Too Long?
Even without I/O, a goroutine that runs continuously for a long time can starve other runnable work of execution opportunities.
The runtime therefore also needs preemption:
G Running
↓
Preempt
↓
G Runnable
↓
Scheduler
↓
Another G gets CPU time
Modern Go supports preemption, including asynchronous preemption, to reduce the chance that a long-running goroutine blocks progress elsewhere. Preemption does not, by itself, guarantee strict fairness.
2.1.4 The Final Model: G, M, and P
2.1.4.1 G: Goroutine
G represents a goroutine: a logical execution unit managed by the Go runtime.
It holds:
- Stack;
- PC / scheduler context;
- Scheduling state such as Runnable, Running, and Waiting.
go f() creates a G, not a new OS thread.
2.1.4.2 M: Machine
M represents a real OS thread.
M
↓
OS Thread
↓
OS Scheduler
↓
CPU
The OS scheduler schedules the OS thread represented by M onto a CPU.
Each M also has a special g0 goroutine. Its system stack is used by the runtime for scheduling, stack management, and other low-level work.
2.1.4.3 P: Processor
P represents the runtime resources and execution capacity an M must hold to run Go code.
It contains scheduler and allocator state, a local run queue, and runnext for preferential scheduling.
The number of P instances is determined by GOMAXPROCS:
GOMAXPROCS
↓
Number of P instances
↓
Maximum parallelism for Go user code
Their roles can be summarized in one sentence:
G is the task to run; M is the OS thread; P provides the runtime resources an M needs to execute Go code.
2.2 Execution Layers
The Go runtime scheduler and OS scheduler operate at different layers. Switching goroutines on the same M does not necessarily involve an OS thread context switch.
2.3 From a Goroutine's Lifecycle to Its Runtime Implementation
A goroutine's conceptual lifecycle is:
The diagram uses conceptual states such as Created and Waiting. The Go runtime uses more specific internal states, including _Grunnable and _Gwaiting.
2.3.1 Creation: Created → Runnable
go f() creates a G, marks it runnable, and enqueues it for scheduling.
2.3.2 Waiting for Scheduling: Runnable
The G is in a run queue, but an M holding a P has not yet selected it.
2.3.3 Execution: Runnable → Running
An M holding a P selects a runnable G through the runtime scheduler, then runs it through execute(G) and gogo. The OS scheduler independently schedules M's underlying OS thread.
2.3.4 Parking and Waiting: Running → Waiting
gopark() suspends only G, letting its M continue running other runnable goroutines while holding P.
2.3.5 Wakeup and Resumption: Waiting → Runnable → Running
When the wait ends, goready() requeues G so the scheduler can dispatch it to an available M.
2.3.6 System Calls and Preemption
A system call can block M while P becomes available to another M; preemption makes a long-running G eligible for rescheduling.
2.3.7 Completion and Cleanup: Running → Dead
When G's function returns, the runtime cleans up G, marks it dead, and M schedules other work.
2.4 End-to-End Sequence Diagram
This sequence follows one G through creation, execution, waiting, and termination. OS scheduling and goroutine switching are separate layers: the diagram shows M being scheduled onto a CPU once, but later goroutine switches on that same M need not trigger another OS scheduling event.
For implementation details, see Go's proc.go and runtime2.go.
2.5 Returning to the Opening Question
Section 2.1.4, The Final Model and Section 2.2, Execution Layers establish that G is a logical execution unit, M corresponds to an OS thread, and M needs P to execute Go code. Therefore, a G executes through an M holding a P, not as a dedicated OS thread.
Section 2.1.3.3, Work Sources, Section 2.3.2, Waiting for Scheduling, and Section 2.3.3, Execution show the Go scheduler selecting runnable Gs from queues, work stealing, or netpoll for an M, while the OS scheduler schedules the underlying OS thread. Therefore, the two schedulers decide which G runs and which OS thread receives CPU time, respectively.
Section 2.3.4, Parking and Section 2.3.5, Resumption show that M can run other Gs while one G waits. Section 2.3.6, System Calls and Preemption shows that another M may take P if a system call blocks the original M, and a long-running G may be preempted. Therefore, waiting or blocking need not leave the associated scheduling resources idle indefinitely.
Answering the opening question: Go uses G-M-P to schedule many goroutines on M instances holding P, while the OS scheduler runs their underlying OS threads on CPUs, allowing concurrent tasks to make progress on a limited number of CPU cores.
3. How Does CPython Schedule Threads?
3.1 From Native OS Threads to the GIL and Free-Threaded Builds
3.1.1 The Simplest Model: One Python Thread per OS Thread
For a small number of concurrent tasks, the execution relationship is straightforward:
Python Thread
↓
OS Thread
↓
Operating System Scheduler
↓
CPU Core
The OS scheduler decides which OS thread receives CPU time.
3.1.2 The Problem: Sharing Interpreter State Across OS Threads
For example, the OS might schedule two Python threads simultaneously on different CPU cores:
Python Thread A → OS Thread A → CPU Core A
Python Thread B → OS Thread B → CPU Core B
The operating system can schedule both threads in parallel. Inside CPython, however, the threads share interpreter state, including object internals, reference counts, and runtime data structures.
This raises a separate runtime question:
How can multiple OS threads safely access shared state while executing Python interpreter code?
The default GIL-enabled CPython build uses the GIL to protect shared interpreter state. Whether a thread can execute protected Python code therefore depends on two independent conditions: OS scheduling and the GIL.
OS Scheduler
Decides which OS Thread gets CPU time
GIL
Restricts which Thread can execute protected Python interpreter code
3.1.3 The Default Solution: The GIL Restricts Interpreter Execution
In a default GIL-enabled build, a thread must hold the GIL while executing interpreter code protected by it:
OS Thread
↓
OS Scheduler
↓
CPU Time
Python Thread
↓
Acquire GIL
↓
Python Interpreter Execution
Executing Python bytecode therefore requires two things: the OS thread must have CPU time, and the thread must hold the GIL.
If another thread already holds the GIL:
Thread A
↓
Holds the GIL
↓
Executes Python bytecode
Thread B
↓
Attempts to acquire the GIL
↓
Waits
↓
The GIL is released
↓
Competes to acquire it again
Their responsibilities remain distinct:
- The OS scheduler governs OS Thread → CPU.
- The GIL gates Thread → Python Interpreter Execution in a default CPython build.
3.1.4 The Next Step: Free-Threaded CPython
To allow multiple Python threads to execute Python code in parallel, the normal execution path must no longer depend on that global lock:
Python Thread A → OS Thread A → CPU Core A
Python Thread B → OS Thread B → CPU Core B
A free-threaded build does not rely on the GIL as a global execution lock on its normal path. Instead, finer-grained synchronization protects CPython's internal state, allowing multiple Python threads to execute code in parallel. What changes is interpreter synchronization; the OS scheduler still schedules the same kind of native threads.
CPython thus offers two principal execution models:
Some incompatible extensions may still cause a free-threaded runtime to enable the GIL.
3.1.5 The Final Model: Thread, PyThreadState, and GIL
3.1.5.1 Python Thread
Python programs create and manage threads through threading.Thread.
In CPython, _thread ultimately creates a real OS thread.
3.1.5.2 OS Thread
The OS scheduler determines when an OS thread runs on a CPU.
Unlike Go's G-M-P or Java virtual threads, CPython does not multiplex large numbers of threading.Thread objects onto a smaller pool of OS threads.
3.1.5.3 PyThreadState
PyThreadState stores the interpreter-side runtime state associated with a Python thread.
A thread's execution context in CPython consists of two parts:
OS Thread
+
PyThreadState
↓
Execution context for this Thread inside the CPython Runtime
3.1.5.4 GIL
In the default GIL-enabled build, a thread must hold the GIL to execute protected Python interpreter code.
Even when its OS thread has CPU time, it must still satisfy the GIL requirement before executing the corresponding Python bytecode.
3.2 Execution Layers
3.3 From CPython Thread Lifecycles to Their Implementation
CPython relies on OS scheduling for threading.Thread. The default GIL-enabled build also requires GIL ownership to execute protected interpreter code; free-threaded builds use different synchronization mechanisms.
3.3.1 Creation and Start: Created → Runnable
Thread.start() creates an OS thread and its PyThreadState; CPU execution still depends on OS scheduling.
3.3.2 Waiting for Scheduling: Runnable
The OS thread underlying a Python thread waits for OS scheduling; the default build must also satisfy its GIL requirement.
3.3.3 Execution: Runnable → Running
The OS scheduler grants CPU time to the native thread; in the default GIL-enabled build, the thread acquires the GIL before executing Python code.
3.3.4 Parking and Waiting: Running → Waiting
Waiting for I/O, locks, or conditions may block the underlying OS thread. In the default build, operations that release the GIL also allow another thread to execute protected interpreter code.
3.3.5 Wakeup and Resumption: Waiting → Runnable → Running
When ready, the OS thread reenters OS scheduling; the default build must reacquire the GIL before resuming protected Python execution if it was released.
3.3.6 Completion and Cleanup: Running → Terminated
After Thread.run() returns, the runtime cleans up PyThreadState, the native thread exits, and its ThreadHandle reaches the completed state.
3.4 End-to-End Sequence Diagram
Here is the thread lifecycle in default GIL-enabled CPython:
For implementation details, see CPython's _threadmodule.c, pystate.c, and ceval_gil.c; for free-threading, see the official guide.
3.5 Returning to the Opening Question
Section 3.1.5, Final Model and Section 3.2, Execution Layers show that threading.Thread maps to a native OS thread, while PyThreadState records its interpreter state. Therefore, the OS scheduler directly grants CPU execution opportunities; CPython has no separate user-space thread scheduler like Go's G-M-P.
Section 3.1.3, GIL, Section 3.3.2, Waiting for Scheduling, and Section 3.3.3, Execution show that a thread in the default GIL-enabled build must acquire the GIL before executing protected Python interpreter code. Therefore, CPU time alone does not guarantee Python-code execution; the GIL is an execution constraint, not another thread scheduler.
Section 3.3.4, Parking and Waiting, Section 3.3.5, Resumption, and Section 3.1.4, Free-Threaded establish that a blocking operation can suspend an OS thread and release the GIL, whereas free-threaded builds change Python-code synchronization without changing the threading.Thread to OS-thread mapping. Therefore, the interpreter's concurrency constraints change, but the execution carrier does not.
Answering the opening question: CPython relies on the OS scheduler for native threads. The default build uses the GIL to limit parallel execution of Python code within one interpreter, whereas free-threaded builds permit multiple threads to execute Python code in parallel; neither model multiplexes large numbers of threading.Thread instances onto a few OS threads.
4. What the Three Implementations Have in Common
After examining Java, Go, and CPython, we can set aside their specific APIs and runtime names. Three questions remain: what carries an execution unit, who decides when it runs, and how does it resume after waiting?
4.1 How Does an Execution Unit Reach the CPU? — Establish Its Carrier
First, distinguish the execution unit defined by a language or runtime from the OS thread that the operating system schedules. They do not necessarily have a one-to-one relationship.
Suppose a program has three execution units, A, B, and C. The most direct approach gives each its own OS thread:
Execution Unit A → OS Thread 1
Execution Unit B → OS Thread 2
Execution Unit C → OS Thread 3
Tasks may, however, spend most of their time waiting for I/O. If every task permanently occupies an OS thread, native-thread resources must remain allocated even for waiting tasks.
Another approach separates tasks from underlying threads:
Runnable Task A ─┐
Runnable Task B ─┼─→ Runtime Dispatch → OS Thread 1
Runnable Task C ─┘ (reused over time)
One OS thread can execute different tasks at different times. Models that do not require user-space scheduling can keep the direct one-to-one mapping. Both ultimately reach the same execution path:
Therefore, execution units can have different carriers, but the OS scheduler ultimately dispatches OS threads—not language-level tasks—to the CPU.
4.2 Who Selects What Runs Next? — Separate the Scheduling Layers
Establishing a carrier does not mean that a task immediately runs. Two different decisions remain:
- Task selection: If a runtime manages lightweight runnable tasks, it chooses one to assign to an underlying thread.
- Thread scheduling: The OS scheduler selects a runnable OS thread to receive CPU time.
Consider a single CPU core. OS threads T1 and T2 are both runnable, and a runtime has already assigned task A to T1:
Runtime: select Task A → assign to T1
OS Scheduler: T1 and T2 are both Runnable
↓
select T2 first
↓
CPU runs T2
Although the runtime has selected task A, it cannot make progress until T1 receives CPU time. Conversely, an OS thread receiving CPU time does not guarantee that a language-level synchronization condition is satisfied; such a condition is not another thread scheduler.
With direct one-to-one mapping, a separate runtime thread scheduler may be unnecessary, but OS scheduling still applies. Long-running work may also have to yield or be preempted so other tasks can make progress.
Therefore, runtime task dispatch and OS thread scheduling make choices at different levels; Runnable is not the same as Running.
4.3 What Happens After a Wait? — Regain an Execution Opportunity
Suppose task A is executing and then starts an I/O operation, while task B is already runnable. The central question is: can A's underlying OS thread run B while A waits?
Task A running → waits for I/O
│
├─ Suspend only logical Task A
│ └─ OS thread can run Task B
│
└─ Block the underlying OS thread
└─ That thread waits; other OS threads can run
Either path is possible, depending on the execution model and whether the specific wait releases its carrier. Suspending a logical task does not necessarily block its OS thread; blocking one OS thread does not stop the entire process.
When the I/O completes, task A is not guaranteed to run immediately. It becomes eligible for scheduling again:
Waiting
↓ condition ready / wakeup
Runnable
↓ optionally queue in runtime and wait for a carrier
↓ OS thread receives CPU time
Running
If only the logical task was suspended, the OS thread could make progress on other tasks. If the OS thread itself was blocked, it must become runnable and receive CPU time again. Waiting and resumption determine whether execution resources can be reused; wakeup does not mean immediate execution.
Returning to the opening question: a language-level execution unit must first be associated with an OS thread. When present, a runtime scheduler selects the task; the OS scheduler assigns CPU time to its underlying thread. Waiting and resumption determine whether that thread can serve other tasks in the meantime. Together, these mechanisms explain how concurrent execution units reach the CPU.
5. Next: Language Memory Models
The previous chapter explained how the CPU executes instructions. This chapter addresses the missing link: how runtimes and operating systems give concurrent execution units CPU time and resume them after waiting.
With the execution path established, the next chapter turns to the Language Memory Model—the rules governing concurrent reads and writes, which provide the foundation for later discussions of Mutex, Atomic, Volatile, and other synchronization mechanisms.
This article was first published on ThinkerQAQ's personal blog and syndicated here by the author. The original article may be revised over time; please refer to the personal blog for the latest version.









































Top comments (0)