Python Internals for Interviews: Decorators, Generators, GIL & OOP
Part 2 of the Python Interview Prep track. Last updated: October 2026.
Algorithms questions test how you think; Python internals questions test how well you know the tool you're building with. Interviewers use them as a fast filter: they can't be faked with a memorized pattern, and wrong answers reveal experience gaps instantly. Below are the six internals topics that come up most often, each framed as the interviewer actually asks it, with runnable code and the one-line answer to memorize.
1. The mutable default argument gotcha
What the interviewer asks: "What's wrong with def f(x, cache=[])?" This is Python's most famous trap, and it still catches people. Run it:
def add_to_cart(item, cart=[]):
cart.append(item)
return cart
print(add_to_cart("apple")) # first call
print(add_to_cart("banana")) # second call — surprise!
# Output:
# ['apple']
# ['apple', 'banana']
The second call should start from an empty cart, but the banana lands in the same list as the apple. The fix is the None sentinel:
def add_to_cart(item, cart=None):
if cart is None:
cart = []
cart.append(item)
return cart
print(add_to_cart("apple")) # ['apple']
print(add_to_cart("banana")) # ['banana'] -- fresh list every time
Why it happens: default argument values are evaluated once, when the def statement runs — not each time the function is called. So every call shares the same list object. You can prove it: add_to_cart.__defaults__ shows the single list stored on the function.
Memorize: "Defaults are evaluated once at def time, so a mutable default is shared across all calls — use None as the default and create the object inside the function."
2. The GIL, simply
What the interviewer asks: "What is the GIL, and when does it matter?" The Global Interpreter Lock means only one thread executes Python bytecode at a time. For CPU-bound work (math, data crunching), threads give you no speedup at all:
import threading, multiprocessing, time
def busy(n): # pure CPU work
total = 0
for i in range(n):
total += i * i
return total
N = 5_000_000
t0 = time.perf_counter()
for _ in range(4):
busy(N) # sequential
print(f"sequential: {time.perf_counter() - t0:.2f}s")
t0 = time.perf_counter()
threads = [threading.Thread(target=busy, args=(N,)) for _ in range(4)]
[t.start() for t in threads]
[t.join() for t in threads]
print(f"4 threads: {time.perf_counter() - t0:.2f}s")
t0 = time.perf_counter()
with multiprocessing.Pool(4) as pool: # separate processes, separate GILs
pool.map(busy, [N] * 4)
print(f"4 processes: {time.perf_counter() - t0:.2f}s")
# Output (typical):
# sequential: 1.45s
# 4 threads: 1.46s -- zero speedup: the GIL
# 4 processes: 1.36s -- real parallelism, each process has its own GIL
Notice the headline result: 4 threads take the same time as 4 sequential runs. The GIL serialized them. Processes bypass the lock because each gets its own Python interpreter — at the cost of heavier startup and inter-process communication.
But threads aren't useless: for I/O-bound work (network requests, disk reads), a thread releases the GIL while waiting, so threads shine there.
Memorize: "Threads for I/O-bound, processes for CPU-bound — the GIL lets only one thread run Python bytecode at a time."
3. Decorators with arguments + functools.wraps
What the interviewer asks: "Write a decorator that takes arguments" — and then, "what breaks if you skip functools.wraps?" Here's a practical @retry(times=3) decorator:
import functools, time
def retry(times):
def decorator(func): # real decorator, built by the factory
@functools.wraps(func) # preserves name, docstring, signature
def wrapper(*args, **kwargs):
for attempt in range(1, times + 1):
try:
return func(*args, **kwargs)
except Exception as e:
print(f"attempt {attempt}/{times} failed: {e}")
time.sleep(0.1)
raise
return wrapper
return decorator
attempts = {"n": 0}
@retry(times=3)
def flaky():
"""Fetch data from an unreliable API."""
attempts["n"] += 1
if attempts["n"] < 3:
raise ConnectionError("server hiccup")
return "data received"
print("result:", flaky())
print("name:", flaky.__name__) # 'flaky' -- thanks to wraps
print("doc:", flaky.__doc__) # 'Fetch data from an unreliable API.'
# Output:
# attempt 1/3 failed: server hiccup
# attempt 2/3 failed: server hiccup
# result: data received
# name: flaky
# doc: Fetch data from an unreliable API.
The pattern is three nested layers: factory(times) → decorator(func) → wrapper(*args, **kwargs). @retry(times=3) calls the factory, which returns the actual decorator. Now watch what happens without wraps:
def plain_decorator(func):
def wrapper(*args, **kwargs):
return func(*args, **kwargs)
return wrapper
@plain_decorator
def real_name():
"""My real docstring."""
pass
print("without wraps name:", real_name.__name__) # 'wrapper' -- wrong!
print("without wraps doc:", real_name.__doc__) # None -- docstring lost!
Why it matters: without @functools.wraps, the decorated function reports the wrapper's __name__ and loses its docstring — which breaks logging, debuggers, and tools like Sphinx that inspect metadata.
Memorize: "A decorator with arguments is a factory returning the decorator, and @functools.wraps(func) on the wrapper preserves the original function's name and docstring."
4. Generators vs lists
What the interviewer asks: "How would you process a 10 GB file on a machine with 4 GB of RAM?" The answer is lazy evaluation. Compare the memory:
import sys
N = 100_000
lst = list(range(N))
gen = (x for x in range(N)) # generator expression, lazy
print(f"list of {N}: {sys.getsizeof(lst):,} bytes") # 800,056 bytes
print(f"generator: {sys.getsizeof(gen):,} bytes") # 192 bytes
A list stores all 100,000 items (~800 KB); the generator stores almost nothing (192 bytes) because it computes values one at a time. Chain generators into a lazy pipeline and memory stays flat no matter how big the input is:
def read_lines(n): # pretend these come from a huge file
for i in range(n):
yield f"row-{i},data"
def clean(lines):
for line in lines:
yield line.replace("-", "_").upper()
def take(pipeline, k):
for i, item in enumerate(pipeline):
if i >= k:
break
print(item)
take(clean(read_lines(10_000_000)), 3) # 10M rows, constant memory
# Output:
# ROW_0,DATA
# ROW_1,DATA
# ROW_2,DATA
Only 3 rows are ever materialized, even though the source pretends to have 10 million. Now the classic trap:
g = (x * 2 for x in range(3))
print("first pass:", list(g)) # [0, 2, 4]
print("second pass:", list(g)) # [] -- exhausted!
Why: a generator is a one-shot iterator — once consumed, it's empty forever. If you need two passes, rebuild it or convert to a list.
Memorize: "Generators yield values lazily with O(1) memory, perfect for pipelines over huge data — but they're single-pass: exhausted after one iteration."
5. The iterator protocol: __iter__ and __next__
What the interviewer asks: "How does a for-loop actually work?" It doesn't index — it uses the iterator protocol. Any object with __iter__ (returns an iterator) and __next__ (returns the next item or raises StopIteration) works in a for-loop:
class Countdown:
def __init__(self, start):
self.current = start
def __iter__(self):
return self # the object IS its own iterator
def __next__(self):
if self.current <= 0:
raise StopIteration # the signal that ends the loop
self.current -= 1
return self.current + 1
for n in Countdown(3):
print("T-minus", n)
# Output:
# T-minus 3
# T-minus 2
# T-minus 1
Here's exactly what the for-loop does under the hood:
it = iter(Countdown(2)) # for-loop calls iter() once
print("manual next:", next(it)) # 2
print("manual next:", next(it)) # 1
try:
next(it) # no more items...
except StopIteration:
print("loop would end here") # ...so the for-loop stops
Why it matters: iterable vs iterator is a favorite follow-up. A list is an iterable (it has __iter__, which hands you a fresh iterator each time — that's why you can loop a list twice). A generator is an iterator (its __iter__ returns itself — that's why it's single-pass, linking back to the exhaustion trap above).
Memorize: "A for-loop calls iter() once, then next() until StopIteration. An iterable gives you a fresh iterator each time; an iterator is single-pass."
6. OOP for interviews: dunders, super(), and copy semantics
What the interviewer asks: "What are dunder methods?" — plus the perennial shallow vs deep copy question. One compact class shows all three dunders interviewers love:
class Point:
def __init__(self, x, y):
self.x = x
self.y = y
def __repr__(self): # how the object prints and debugs
return f"Point({self.x}, {self.y})"
def __eq__(self, other): # value equality, not identity
return isinstance(other, Point) and (self.x, self.y) == (other.x, other.y)
def __len__(self): # a Point "contains" two coordinates
return 2
class Point3D(Point):
def __init__(self, x, y, z):
super().__init__(x, y) # reuse the parent's __init__
self.z = z
def __repr__(self):
return f"Point3D({self.x}, {self.y}, {self.z})"
p = Point(1, 2)
print(p) # Point(1, 2)
print(p == Point(1, 2)) # True (without __eq__, this would be False)
print(len(p)) # 2
print(Point3D(1, 2, 3)) # Point3D(1, 2, 3)
__repr__ controls what you see in the debugger, __eq__ lets == compare values instead of object identity (note the isinstance guard — a sloppy __eq__ that assumes the type is another interview red flag), and __len__ makes len() work. super().__init__(x, y) reuses the parent's setup instead of duplicating it.
Now the copy question — the aliasing surprise:
import copy
original = [[1, 2], [3, 4]]
shallow = copy.copy(original) # copies the OUTER list only
deep = copy.deepcopy(original) # copies everything, recursively
shallow[0].append(99) # mutates the shared inner list!
print("original:", original) # [[1, 2, 99], [3, 4]] -- changed too!
print("shallow: ", shallow) # [[1, 2, 99], [3, 4]]
deep[1].append(100)
print("original:", original) # [[1, 2, 99], [3, 4]] -- untouched
print("deep: ", deep) # [[1, 2], [3, 4, 100]]
Why: copy.copy duplicates only the outer container, so the inner lists are still shared — mutating through one alias mutates the other. copy.deepcopy clones the entire object graph. Bonus: plain assignment (b = a) doesn't copy anything at all.
Memorize: "copy.copy duplicates the outer container but shares nested objects; copy.deepcopy clones everything recursively."
Key takeaways
- Mutable default arguments are evaluated once at def time — use a None sentinel and build the object inside the function.
- The GIL lets one thread run Python bytecode at a time: threads for I/O-bound, processes for CPU-bound.
- A decorator with arguments is a factory returning the decorator; always use @functools.wraps to preserve name and docstring.
- Generators give O(1) memory pipelines over huge data, but are single-pass — exhausted after one iteration.
- A for-loop calls iter() once, then next() until StopIteration; iterables are reusable, iterators are not.
- __repr__, __eq__, __len__ are the dunders to know cold; super() reuses parent init; deepcopy clones the whole object graph, copy only the outer shell.
Next in this series: Tricky Python Questions Interviewers Love.
Comments
Post a Comment