Multiprocessing, IPC & Shared Memory
To bypass the Global Interpreter Lock (GIL) and achieve true CPU parallelism across multiple CPU cores, Python applications use the multiprocessing module. Spawning separate OS processes bypasses GIL contention entirely because each process gets its own CPython interpreter instance, heap memory, and GIL.
This chapter details Process Start Methods (fork, spawn, forkserver), Inter-Process Communication (IPC via Queue and Pipe), zero-copy shared memory (multiprocessing.shared_memory), and process pool management.
1. Process Start Methods: fork vs. spawn vs. forkserver
Python supports 3 OS process creation mechanisms (multiprocessing.set_start_method()):
Process Start Methods Comparison:
1. FORK (Default on legacy Linux):
Clones host process via POSIX fork(2) with Copy-On-Write (COW) memory.
- Pros: Instant startup, inherits global state.
- Cons: UNSAFE with multithreading or C-library locks (causes deadlocks!).
2. SPAWN (Default on macOS & Windows):
Executes a fresh Python interpreter process from scratch.
- Pros: 100% thread-safe; clean isolated heap state.
- Cons: Slower startup; requires pickling all arguments.
3. FORKSERVER:
Spawns a clean single-threaded server process that forks new workers.
- Pros: Fast startup + Thread-safe isolation!2. Inter-Process Communication (IPC): Queue & Pipe
Because OS processes do not share memory space, data passed between processes must be serialized (pickled) over OS IPC primitives:
multiprocessing.Queue: Multi-producer, multi-consumer FIFO queue using an underlying OS pipe and background serialization thread.multiprocessing.Pipe(): Bi-directional 1-on-1 connection pipe between two processes. Executes faster thanQueuefor 1-to-1 process messaging.
import multiprocessing
def worker(q: multiprocessing.Queue):
q.put("Data from worker process")
if __name__ == "__main__":
multiprocessing.set_start_method("spawn")
q = multiprocessing.Queue()
p = multiprocessing.Process(target=worker, args=(q,))
p.start()
print(q.get()) # Receives un-pickled data from worker process
p.join()3. High-Performance Zero-Copy Shared Memory (shared_memory)
Passing large datasets (like 1GB NumPy arrays) through IPC Queue requires serializing (pickle) and deserializing data across process boundaries, creating massive CPU and RAM overhead.
Introduced in Python 3.8, multiprocessing.shared_memory allocates shared OS RAM segments mapped into the virtual address space of multiple processes:
Zero-Copy Shared Memory Architecture:
[ Shared OS RAM Segment (/dev/shm) ]
^ ^
| (Mapped via mmap) | (Mapped via mmap)
[ Process 1 (NumPy Array) ] [ Process 2 (NumPy Array) ]
(Zero pickling! Zero copying! Sub-millisecond multi-gigabyte array access!)from multiprocessing.shared_memory import SharedMemory
import numpy as np
# Process 1: Create shared memory block for a 1,000,000 float array
shm = SharedMemory(create=True, size=8000000)
arr = np.ndarray((1000000,), dtype=np.float64, buffer=shm.buf)
arr[:] = 42.0
# Pass shm.name string to Process 2...
# Process 2: Attach to existing shared memory segment
shm_client = SharedMemory(name=shm.name)
arr_client = np.ndarray((1000000,), dtype=np.float64, buffer=shm_client.buf)
print(arr_client[0]) # Prints 42.0 INSTANTLY in zero-copy memory!4. Production Process Pool Management (multiprocessing.Pool)
Use multiprocessing.Pool (or ProcessPoolExecutor) to manage worker process lifecycles automatically:
from multiprocessing import Pool
def square(x: int) -> int:
return x * x
if __name__ == "__main__":
with Pool(processes=4) as pool:
results = pool.map(square, range(100)) # Distributes work across 4 CPU cores!