Strings, Unicode, Bytes & Text Encodings
In Python 3, text (str) and binary data (bytes) are strictly decoupled. A str is an abstract sequence of Unicode code points, while a bytes object is an immutable sequence of raw 8-bit integers ($0..255$). Understanding PEP 393 Flexible String Representation, UTF-8 encoding mechanics, and buffer protocols is essential for systems engineering.
This chapter details CPython’s string memory layout, compact Unicode objects, bytes vs bytearray memory buffers, and text encoding boundaries.
1. PEP 393: Flexible String Representation
Prior to Python 3.3, CPython allocated all strings using either 2-byte (UCS-2) or 4-byte (UCS-4) arrays, wasting vast amounts of RAM for simple ASCII text. PEP 393 introduced a dynamic, compact internal representation:
CPython inspects the maximum code point character in a string at creation time and selects one of three compact C array representations:
- 1-Byte Representation (Latin-1 / ASCII): Used if max code point $\le U+00FF$ (1 byte per char).
- 2-Byte Representation (UCS-2): Used if max code point $\le U+FFFF$ (2 bytes per char).
- 4-Byte Representation (UCS-4): Used if max code point $> U+FFFF$ (e.g. Emojis, 4 bytes per char).
CPython PEP 393 String Memory Layout:
ASCII String "hello" (Max char <= U+00FF):
[ PyASCIIObject Header (48B) ] -> [ 'h' | 'e' | 'l' | 'l' | 'o' | '\0' ] (1 byte/char)
Emoji String "hello Python 🐍" (Max char U+1F40D > U+FFFF):
[ PyCompactUnicodeObject Header (72B) ] -> [ 4 bytes per character array ]Because of PEP 393, pure ASCII strings consume 1 byte per character payload, while strings containing a single emoji expand the internal C-array to 4 bytes per character for all characters in that string.
2. Text (str) vs. Binary (bytes & bytearray)
str: Immutable sequence of Unicode code points. Cannot be sent directly over network sockets or written directly to disk without encoding.bytes: Immutable sequence of raw bytes ($0..255$). Represents encoded text, image data, or network payloads.bytearray: Mutable version ofbytes. Allows in-place byte modifications without re-allocating new objects on the heap.
# Encoding: str -> bytes (Unicode Code Points -> UTF-8 Bytes)
text = "Python 🐍"
encoded_bytes = text.encode("utf-8")
print(encoded_bytes) # b'Python \xf0\x9f\x90\x8d' (Emoji takes 4 UTF-8 bytes)
# Decoding: bytes -> str (UTF-8 Bytes -> Unicode Code Points)
decoded_text = encoded_bytes.decode("utf-8")3. Unicode Normalization Forms (NFC vs. NFD)
Visual equality does not guarantee binary equality in Unicode. A character like é can be represented as a single precomposed code point (U+00E9) or a base letter e plus a combining accent (U+0065 + U+0301).
import unicodedata
str1 = "café" # Precomposed (NFC)
str2 = "cafe\u0301" # Decomposed (NFD)
print(str1 == str2) # False! (Binary code points differ)
# Solution: Normalize strings before storing or comparing
norm1 = unicodedata.normalize("NFC", str1)
norm2 = unicodedata.normalize("NFC", str2)
print(norm1 == norm2) # True!4. Production Trade-offs & In-Place Buffers
- String Concatenation in Loops: Strings are immutable. Executing
s += charinside a loop creates an $O(N^2)$ allocation catastrophe. Use''.join(list_of_strings)or abytearraybuffer. bytearrayfor Sockets: When reading binary streams from sockets, pre-allocate abytearrayand read directly into its buffer to eliminate per-chunk heap allocations.