str is Unicode text; bytes is raw binary; encode/decode at boundaries
In Python 3, str represents Unicode code points, while bytes represents raw 8-bit values. You decode bytes to str and encode str to bytes. CPython stores str internally using the most compact representation that fits the widest character: latin-1, UCS-2, or UCS-4. That means len(str) returns the number of code points, not bytes and not grapheme clusters. Text should be decoded at input boundaries and encoded at output boundaries; keep bytes for binary protocols, files, and network payloads.
str + str is text concatenation; bytes + bytes is binary concatenation. Mixing them raises TypeError.
Encoding is str -> bytes; decoding is bytes -> str. Always specify encoding, usually utf-8, at boundaries.
Common mistake: assuming len(str) counts user-perceived characters. Emoji and combining marks can span multiple code points.
Trade-off: UTF-8 bytes are compact for ASCII but variable-width; UTF-32 is fixed-width but memory-heavy.
Version note: Python 3.15 is expected to move toward UTF-8 mode by default via PEP 686; explicit encoding is still best practice.
0-2 years experience
2-5 years experience
5-8 years experience
8+ years experience