File I/O, Buffering, Encodings & Streaming
File I/O in Python is implemented via the CPython io module. Reading or writing a file passes data through a 3-tier stream stack (io.TextIOWrapper $\rightarrow$ io.BufferedWriter $\rightarrow$ io.FileIO). Understanding stream buffering levels, line-by-line streaming, explicit encoding handling, and chunked memory processing is essential for handling large files safely.
This chapter details CPython’s 3-tier I/O stream stack, buffer size mechanics, UTF-8 byte boundary decoding, and chunked streaming patterns.
1. CPython 3-Tier I/O Stream Architecture
When you call open("file.txt", "r", encoding="utf-8"), CPython constructs a 3-tier layered stream object:
CPython File I/O Stream Stack:
[ User Application Code (read / write) ]
|
v
[ Tier 1: io.TextIOWrapper ] <-- Encodes str <-> bytes (UTF-8, ASCII) & handles \n
|
v
[ Tier 2: io.BufferedReader / BufferedWriter ] <-- 8KB C memory buffer (accumulates reads/writes)
|
v (Buffer full / flush() call)
[ Tier 3: io.FileIO ] <-- Raw unbuffered C system calls: read(2), write(2)
|
v
[ OS Kernel File Descriptor ]Stream Modes:
- Text Mode (
"r","w"): Instantiatesio.TextIOWrapper. Translates raw bytes to Unicodestrusing a specified encoding. - Buffered Binary Mode (
"rb","wb"): Instantiatesio.BufferedReaderorio.BufferedWriter. Operates directly onbyteswithout text encoding overhead. - Unbuffered Binary Mode (
"rb0",buffering=0): Instantiates rawio.FileIO. Issues C system calls immediately.
2. Memory-Safe File Streaming & Chunking
Loading a 10GB file using f.read() materializes the entire file into RAM as a single string, causing an instant Out-Of-Memory (OOM) process crash.
Production Streaming Patterns:
-
Line-by-Line Generator Streaming (Text Files): Iterating over a file object (
for line in f:) usesBufferedReader’s internal buffer to yield lines lazily in $O(1)$ memory space. -
Fixed-Size Chunked Streaming (Binary Files): For binary files (videos, zip archives), read data in fixed-size byte chunks (e.g. 64KB):
from pathlib import Path
def process_large_binary_file(file_path: Path, chunk_size: int = 65536):
with open(file_path, "rb") as f:
while chunk := f.read(chunk_size):
process_chunk(chunk) # Constant 64KB RAM usage!3. Explicit Encoding Declarations & OS Pitfalls
Never invoke open("file.txt") without declaring encoding="utf-8".
If encoding is omitted:
- On Linux/macOS, Python defaults to UTF-8.
- On Windows, Python defaults to
cp1252ormbcs(ANSI)!
Opening a UTF-8 encoded file on Windows without declaring encoding="utf-8" causes UnicodeDecodeError or silently corrupts international characters (charmap codec error).
4. Production Trade-offs & os.fsync()
f.flush()vsos.fsync(f.fileno()):f.flush()pushes data from CPython’sBufferedWriterinto the OS kernel buffer. However, data still resides in the OS kernel RAM. Callingos.fsync(f.fileno())forces the OS kernel to flush its write cache physically to the disk storage medium.