Python Internals for Interviews: Decorators, Generators, GIL & OOP

Part 2 of the Python Interview Prep track. Last updated: October 2026.

Algorithms questions test how you think; Python internals questions test how well you know the tool you're building with. Interviewers use them as a fast filter: they can't be faked with a memorized pattern, and wrong answers reveal experience gaps instantly. Below are the six internals topics that come up most often, each framed as the interviewer actually asks it, with runnable code and the one-line answer to memorize.

1. The mutable default argument gotcha

What the interviewer asks: "What's wrong with def f(x, cache=[])?" This is Python's most famous trap, and it still catches people. Run it:

def add_to_cart(item, cart=[]):
    cart.append(item)
    return cart

print(add_to_cart("apple"))   # first call
print(add_to_cart("banana"))  # second call — surprise!

# Output:
# ['apple']
# ['apple', 'banana']

The second call should start from an empty cart, but the banana lands in the same list as the apple. The fix is the None sentinel:

def add_to_cart(item, cart=None):
    if cart is None:
        cart = []
    cart.append(item)
    return cart

print(add_to_cart("apple"))    # ['apple']
print(add_to_cart("banana"))  # ['banana']  -- fresh list every time

Why it happens: default argument values are evaluated once, when the def statement runs — not each time the function is called. So every call shares the same list object. You can prove it: add_to_cart.__defaults__ shows the single list stored on the function.

Memorize: "Defaults are evaluated once at def time, so a mutable default is shared across all calls — use None as the default and create the object inside the function."

2. The GIL, simply

What the interviewer asks: "What is the GIL, and when does it matter?" The Global Interpreter Lock means only one thread executes Python bytecode at a time. For CPU-bound work (math, data crunching), threads give you no speedup at all:

import threading, multiprocessing, time

def busy(n):          # pure CPU work
    total = 0
    for i in range(n):
        total += i * i
    return total

N = 5_000_000

t0 = time.perf_counter()
for _ in range(4):
    busy(N)                       # sequential
print(f"sequential: {time.perf_counter() - t0:.2f}s")

t0 = time.perf_counter()
threads = [threading.Thread(target=busy, args=(N,)) for _ in range(4)]
[t.start() for t in threads]
[t.join() for t in threads]
print(f"4 threads:  {time.perf_counter() - t0:.2f}s")

t0 = time.perf_counter()
with multiprocessing.Pool(4) as pool:   # separate processes, separate GILs
    pool.map(busy, [N] * 4)
print(f"4 processes: {time.perf_counter() - t0:.2f}s")

# Output (typical):
# sequential: 1.45s
# 4 threads:  1.46s   -- zero speedup: the GIL
# 4 processes: 1.36s  -- real parallelism, each process has its own GIL

Notice the headline result: 4 threads take the same time as 4 sequential runs. The GIL serialized them. Processes bypass the lock because each gets its own Python interpreter — at the cost of heavier startup and inter-process communication.

But threads aren't useless: for I/O-bound work (network requests, disk reads), a thread releases the GIL while waiting, so threads shine there.

Memorize: "Threads for I/O-bound, processes for CPU-bound — the GIL lets only one thread run Python bytecode at a time."

3. Decorators with arguments + functools.wraps

What the interviewer asks: "Write a decorator that takes arguments" — and then, "what breaks if you skip functools.wraps?" Here's a practical @retry(times=3) decorator:

import functools, time

def retry(times):
    def decorator(func):                 # real decorator, built by the factory
        @functools.wraps(func)           # preserves name, docstring, signature
        def wrapper(*args, **kwargs):
            for attempt in range(1, times + 1):
                try:
                    return func(*args, **kwargs)
                except Exception as e:
                    print(f"attempt {attempt}/{times} failed: {e}")
                    time.sleep(0.1)
            raise
        return wrapper
    return decorator

attempts = {"n": 0}

@retry(times=3)
def flaky():
    """Fetch data from an unreliable API."""
    attempts["n"] += 1
    if attempts["n"] < 3:
        raise ConnectionError("server hiccup")
    return "data received"

print("result:", flaky())
print("name:", flaky.__name__)   # 'flaky' -- thanks to wraps
print("doc:", flaky.__doc__)     # 'Fetch data from an unreliable API.'

# Output:
# attempt 1/3 failed: server hiccup
# attempt 2/3 failed: server hiccup
# result: data received
# name: flaky
# doc: Fetch data from an unreliable API.

The pattern is three nested layers: factory(times) → decorator(func) → wrapper(*args, **kwargs). @retry(times=3) calls the factory, which returns the actual decorator. Now watch what happens without wraps:

def plain_decorator(func):
    def wrapper(*args, **kwargs):
        return func(*args, **kwargs)
    return wrapper

@plain_decorator
def real_name():
    """My real docstring."""
    pass

print("without wraps name:", real_name.__name__)  # 'wrapper' -- wrong!
print("without wraps doc:", real_name.__doc__)    # None -- docstring lost!

Why it matters: without @functools.wraps, the decorated function reports the wrapper's __name__ and loses its docstring — which breaks logging, debuggers, and tools like Sphinx that inspect metadata.

Memorize: "A decorator with arguments is a factory returning the decorator, and @functools.wraps(func) on the wrapper preserves the original function's name and docstring."

4. Generators vs lists

What the interviewer asks: "How would you process a 10 GB file on a machine with 4 GB of RAM?" The answer is lazy evaluation. Compare the memory:

import sys

N = 100_000
lst = list(range(N))
gen = (x for x in range(N))          # generator expression, lazy
print(f"list of {N}: {sys.getsizeof(lst):,} bytes")   # 800,056 bytes
print(f"generator:   {sys.getsizeof(gen):,} bytes")   # 192 bytes

A list stores all 100,000 items (~800 KB); the generator stores almost nothing (192 bytes) because it computes values one at a time. Chain generators into a lazy pipeline and memory stays flat no matter how big the input is:

def read_lines(n):          # pretend these come from a huge file
    for i in range(n):
        yield f"row-{i},data"

def clean(lines):
    for line in lines:
        yield line.replace("-", "_").upper()

def take(pipeline, k):
    for i, item in enumerate(pipeline):
        if i >= k:
            break
        print(item)

take(clean(read_lines(10_000_000)), 3)   # 10M rows, constant memory
# Output:
# ROW_0,DATA
# ROW_1,DATA
# ROW_2,DATA

Only 3 rows are ever materialized, even though the source pretends to have 10 million. Now the classic trap:

g = (x * 2 for x in range(3))
print("first pass:", list(g))    # [0, 2, 4]
print("second pass:", list(g))   # [] -- exhausted!

Why: a generator is a one-shot iterator — once consumed, it's empty forever. If you need two passes, rebuild it or convert to a list.

Memorize: "Generators yield values lazily with O(1) memory, perfect for pipelines over huge data — but they're single-pass: exhausted after one iteration."

5. The iterator protocol: __iter__ and __next__

What the interviewer asks: "How does a for-loop actually work?" It doesn't index — it uses the iterator protocol. Any object with __iter__ (returns an iterator) and __next__ (returns the next item or raises StopIteration) works in a for-loop:

class Countdown:
    def __init__(self, start):
        self.current = start

    def __iter__(self):
        return self                    # the object IS its own iterator

    def __next__(self):
        if self.current <= 0:
            raise StopIteration         # the signal that ends the loop
        self.current -= 1
        return self.current + 1

for n in Countdown(3):
    print("T-minus", n)
# Output:
# T-minus 3
# T-minus 2
# T-minus 1

Here's exactly what the for-loop does under the hood:

it = iter(Countdown(2))     # for-loop calls iter() once
print("manual next:", next(it))   # 2
print("manual next:", next(it))   # 1
try:
    next(it)                      # no more items...
except StopIteration:
    print("loop would end here")  # ...so the for-loop stops

Why it matters: iterable vs iterator is a favorite follow-up. A list is an iterable (it has __iter__, which hands you a fresh iterator each time — that's why you can loop a list twice). A generator is an iterator (its __iter__ returns itself — that's why it's single-pass, linking back to the exhaustion trap above).

Memorize: "A for-loop calls iter() once, then next() until StopIteration. An iterable gives you a fresh iterator each time; an iterator is single-pass."

6. OOP for interviews: dunders, super(), and copy semantics

What the interviewer asks: "What are dunder methods?" — plus the perennial shallow vs deep copy question. One compact class shows all three dunders interviewers love:

class Point:
    def __init__(self, x, y):
        self.x = x
        self.y = y

    def __repr__(self):        # how the object prints and debugs
        return f"Point({self.x}, {self.y})"

    def __eq__(self, other):   # value equality, not identity
        return isinstance(other, Point) and (self.x, self.y) == (other.x, other.y)

    def __len__(self):          # a Point "contains" two coordinates
        return 2

class Point3D(Point):
    def __init__(self, x, y, z):
        super().__init__(x, y)  # reuse the parent's __init__
        self.z = z

    def __repr__(self):
        return f"Point3D({self.x}, {self.y}, {self.z})"

p = Point(1, 2)
print(p)                 # Point(1, 2)
print(p == Point(1, 2))  # True  (without __eq__, this would be False)
print(len(p))            # 2
print(Point3D(1, 2, 3))  # Point3D(1, 2, 3)

__repr__ controls what you see in the debugger, __eq__ lets == compare values instead of object identity (note the isinstance guard — a sloppy __eq__ that assumes the type is another interview red flag), and __len__ makes len() work. super().__init__(x, y) reuses the parent's setup instead of duplicating it.

Now the copy question — the aliasing surprise:

import copy

original = [[1, 2], [3, 4]]
shallow = copy.copy(original)        # copies the OUTER list only
deep = copy.deepcopy(original)       # copies everything, recursively

shallow[0].append(99)                # mutates the shared inner list!
print("original:", original)  # [[1, 2, 99], [3, 4]] -- changed too!
print("shallow: ", shallow)   # [[1, 2, 99], [3, 4]]

deep[1].append(100)
print("original:", original)  # [[1, 2, 99], [3, 4]] -- untouched
print("deep:    ", deep)      # [[1, 2], [3, 4, 100]]

Why: copy.copy duplicates only the outer container, so the inner lists are still shared — mutating through one alias mutates the other. copy.deepcopy clones the entire object graph. Bonus: plain assignment (b = a) doesn't copy anything at all.

Memorize: "copy.copy duplicates the outer container but shares nested objects; copy.deepcopy clones everything recursively."

Key takeaways

  • Mutable default arguments are evaluated once at def time — use a None sentinel and build the object inside the function.
  • The GIL lets one thread run Python bytecode at a time: threads for I/O-bound, processes for CPU-bound.
  • A decorator with arguments is a factory returning the decorator; always use @functools.wraps to preserve name and docstring.
  • Generators give O(1) memory pipelines over huge data, but are single-pass — exhausted after one iteration.
  • A for-loop calls iter() once, then next() until StopIteration; iterables are reusable, iterators are not.
  • __repr__, __eq__, __len__ are the dunders to know cold; super() reuses parent init; deepcopy clones the whole object graph, copy only the outer shell.

Next in this series: Tricky Python Questions Interviewers Love.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number