FastAPI Async I/O, Concurrency & Worker Strategy
FastAPI is widely known for high performance because it natively supports Python’s async/await syntax powered by Starlette and AnyIO. However, declaring route handlers with async def vs standard def alters how the server dispatches execution. Misunderstanding these execution paths can lead to event loop starvation, causing server-wide latency spikes.
This chapter details the execution engine behind async def vs def, Starlette’s threadpool dispatching mechanism, event loop blocking traps, and Uvicorn worker process architecture.
1. The Execution Engine: async def vs. def
FastAPI handles async def and regular def view functions differently:
FastAPI View Dispatching Architecture:
Incoming HTTP Request
|
v
Is View Handler declared as 'async def' or 'def'?
|
+---> 'async def': Executed DIRECTLY on the Main ASGI Event Loop (uvloop)
| (Must NEVER block; use non-blocking async libraries!)
|
+---> 'def': Offloaded to Starlette's External Threadpool (run_in_threadpool)
(Executes synchronously inside a separate worker thread)1. async def (Event Loop Execution):
Executed directly on the single-threaded main asyncio event loop (uvloop). It handles non-blocking I/O operations using await (e.g. httpx.AsyncClient, asyncpg). It achieves massive concurrency with minimal memory overhead.
2. def (Threadpool Offloading):
If a handler is defined as a standard def function, FastAPI automatically offloads its execution to Starlette’s anyio threadpool (run_in_threadpool). This prevents synchronous operations from blocking the main event loop, but limits concurrency to the threadpool size (default: 40 threads).
2. The Fatal Event Loop Starvation Trap
The #1 performance mistake in FastAPI is running blocking synchronous I/O code inside an async def handler:
# FATAL TRAP: Synchronous blocking call inside 'async def'!
@app.get("/data")
async def get_data():
time.sleep(5) # BLOCKS THE ENTIRE EVENT LOOP FOR 5 SECONDS FOR ALL USERS!
return {"status": "done"}Because async def runs on the main event loop, calling time.sleep(5), requests.get(), or synchronous DB calls (like un-adapted SQLAlchemy sync queries) freezes the entire Uvicorn worker process, causing all concurrent HTTP requests across all users to stall!
3. Uvicorn Worker Process Architecture
Because Python execution is bound by the Global Interpreter Lock (GIL), a single Uvicorn ASGI process utilizes only one CPU core.
In production deployments:
- Process Management: Run Gunicorn as the process manager to control worker lifecycles.
- Worker Class: Specify
UvicornWorker(gunicorn -w 4 -k uvicorn.workers.UvicornWorker main:app). - Rule of Thumb: Allocate
(2 x $NUM_CPU_CORES) + 1worker processes per server node.
Production Process Scaling:
[ Gunicorn Master Process ] <-- Monitors health, handles SIGTERM / SIGHUP
├── [ Uvicorn Worker 1 (CPU Core 0) ] -> Main Event Loop (uvloop)
├── [ Uvicorn Worker 2 (CPU Core 1) ] -> Main Event Loop (uvloop)
├── [ Uvicorn Worker 3 (CPU Core 2) ] -> Main Event Loop (uvloop)
└── [ Uvicorn Worker 4 (CPU Core 3) ] -> Main Event Loop (uvloop)4. Production Decision Matrix: async def vs. def
| Workload Type | Handler Signature | Underlying Execution | Recommended Library |
|---|---|---|---|
| Async I/O (HTTP, Redis, Async DB) | async def | Main Event Loop | httpx, redis-py (async), asyncpg |
| Sync I/O (Blocking SDK, legacy DB) | def | Threadpool (anyio) | requests, psycopg2, standard libraries |
| CPU-Bound (Image processing, ML) | def or Background Task | ProcessPool / Celery | multiprocessing, Celery |