nextRound
TechnologiesCoding ProblemsBookmarksLearning PathsLogin
nextRound
TechnologiesCoding ProblemsBookmarksLearning PathsLogin
nextRound

AI-powered interview preparation platform. Practice with curated questions, mock interviews, and personalized learning paths to crack your dream tech interview.

Quick Links

  • Technologies
  • Mock Interviews
  • Saved Questions
  • Pricing

Company

  • About Us
  • Contact Us

Legal

  • Privacy Policy
  • Terms of Use

© 2026 nextRound. All rights reserved.

Questions
14 of 14
1What are the common built-in data types in Python?
2what is an array
3Since Python 3.7, dictionaries preserve insertion order. How is this guaranteed internally, and what changed from earlier versions?
4What is the difference between str.format(), %-formatting, and f-strings? What are the performance and readability tradeoffs?
5When would you use collections.deque instead of a list, and why?
6Why are strings immutable in Python, and what performance implications does this have for repeated concatenation in a loop?
7What problem does collections.defaultdict solve, and how does it differ from using dict.setdefault?
8How would you efficiently remove duplicates from a list while preserving order?
9How would you design a Least Recently Used (LRU) cache using Python's built-in data structures?
10How does a Python dictionary achieve average O(1) lookup time internally?
11What is the difference between a list and a tuple, and when would you choose one over the other?
12What are the time complexities of common list operations (indexing, append, insert, pop, search)?
13What is the difference between a set and a frozenset?
14How does Python handle Unicode internally, and what is the difference between str and bytes?
PythonPython
Basics
Control Flow and Functions
Data Structures
Comprehensions & Functional Programming
Iterators, Generators & Decorators
Object-Oriented Programming
Exception Handling & Debugging
Concurrency & Parallelism
Performance & Optimization
Testing
Security
Modules, Packaging & Environment
Type Hinting & Modern Python
System Design & Architecture with Python
Best Practices & Design Patterns
Edge Cases & Tricky Interview Questions
14 / 14

How does Python handle Unicode internally, and what is the difference between str and bytes?

Difficulty: 8/10
Strings, Bytes, Unicode

str is Unicode text; bytes is raw binary; encode/decode at boundaries

In Python 3, str represents Unicode code points, while bytes represents raw 8-bit values. You decode bytes to str and encode str to bytes. CPython stores str internally using the most compact representation that fits the widest character: latin-1, UCS-2, or UCS-4. That means len(str) returns the number of code points, not bytes and not grapheme clusters. Text should be decoded at input boundaries and encoded at output boundaries; keep bytes for binary protocols, files, and network payloads.

  1. 1

    str + str is text concatenation; bytes + bytes is binary concatenation. Mixing them raises TypeError.

  2. 2

    Encoding is str -> bytes; decoding is bytes -> str. Always specify encoding, usually utf-8, at boundaries.

  3. 3

    Common mistake: assuming len(str) counts user-perceived characters. Emoji and combining marks can span multiple code points.

  4. 4

    Trade-off: UTF-8 bytes are compact for ASCII but variable-width; UTF-32 is fixed-width but memory-heavy.

  5. 5

    Version note: Python 3.15 is expected to move toward UTF-8 mode by default via PEP 686; explicit encoding is still best practice.

javascript

Scenario Questions

0-2 years experience

  1. 1You read a text file. Do you get str or bytes?
  2. 2You need to send text over network. What conversion?

2-5 years experience

  1. 1You decode UTF-8 bytes and get UnicodeDecodeError. How do you debug?
  2. 2Why is len('é') sometimes 1 and sometimes 2?

5-8 years experience

  1. 1You process multilingual text and need grapheme clusters. What library?
  2. 2You compare str and bytes for equality. Why false?

8+ years experience

  1. 1Design an encoding boundary strategy for a system that stores bytes but serves Unicode text.
  2. 2You need to handle invalid UTF-8 without data loss. What approach?

Follow-up Questions

  • Why can len('é') be 1 or 2 depending on normalization?
  • How do you handle invalid UTF-8 without losing data?
Sharethis question

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.