Take-Home Build: End-to-End Client Integration Project
Part 7 of the Python for FDE track. Last updated: October 2026.
Time to put the whole track to work. FDE interviews are decided by take-home builds: a realistic brief, a tight time box, and a rubric that rewards judgment. This post is a complete practice take-home — the brief, the grading rubric, and a reference solution sketch using everything from Posts 1–6. Do it start to finish before your next interview.
The brief
Scenario: You are the FDE embedded at "Northwind Retail". They want a price-monitoring demo by end of day.
- Integrate — Pull the product catalog from the supplier API (https://api.supplier.example/v1/products, bearer-token auth, cursor pagination, 100/page). Handle 429s with backoff. Save raw responses to raw/.
- Wrangle — Load Northwind's client_prices.csv (messy: mixed-case SKUs, "$1,240.00" price strings, some missing rows). Clean it and merge with the API catalog on SKU. Report how many SKUs failed to match.
- AI feature — Add an "Email summary" button that uses an LLM to draft a client-ready summary of the week's price changes. API key from environment only.
- Dashboard — A Streamlit app a non-technical stakeholder can use: metrics up top, a filterable table, the summary button.
- Ship it — Dockerfile + README so the hiring manager runs it with two commands. No secrets in the repo.
Constraints: 4 hours. Python only for the app code. You may use any PyPI libraries. Assume the reviewer will run it cold.
The grading rubric
This is what hiring managers actually score. Encode it as a checklist and grade yourself honestly:
RUBRIC = {
"API integration (auth, pagination, retry/backoff, raw saved)": 25,
"Data cleaning (dtypes, missing values, normalized merge keys)": 20,
"LLM feature (structured prompt, defensive parsing, env secrets)": 20,
"Dashboard usability (a non-technical stakeholder can click it)": 20,
"Ship-ready (Dockerfile, README, two-command run, no leaked secrets)": 15,
}
print("Total:", sum(RUBRIC.values()), "points")
# Total: 100 points
The 4-hour plan
Time-box ruthlessly. A finished, slightly rough demo beats a polished half-demo every time:
PLAN = [
(30, "Read the spec; get the API returning data; save raw/"),
(60, "Cleaning + merge pipeline working end to end"),
(45, "LLM summary feature with defensive parsing"),
(45, "Streamlit dashboard a stakeholder can click"),
(30, "Dockerfile + README; two-command cold run"),
(30, "Buffer: error handling, re-read rubric, record a walkthrough"),
]
for minutes, task in PLAN:
print(f"{minutes:>3} min — {task}")
# 30 min — Read the spec; get the API returning data; save raw/
# 60 min — Cleaning + merge pipeline working end to end
# 45 min — LLM summary feature with defensive parsing
# 45 min — Streamlit dashboard a stakeholder can click
# 30 min — Dockerfile + README; two-command cold run
# 30 min — Buffer: error handling, re-read rubric, record a walkthrough
Reference solution: the pipeline
The graded core is pipeline.py — fetch, clean, merge, in that order, each step a testable function (Posts 2 and 4):
# pipeline.py — fetch -> clean -> merge
import requests
import pandas as pd
def fetch_products(api_base: str, token: str) -> list:
session = requests.Session()
session.headers.update({"Authorization": f"Bearer {token}"})
items, cursor = [], None
while True:
params = {"limit": 100, **({"cursor": cursor} if cursor else {})}
r = session.get(f"{api_base}/products", params=params, timeout=15)
r.raise_for_status()
page = r.json()
items.extend(page["data"])
cursor = page.get("next_cursor")
if not cursor:
return items
def clean_client_csv(path: str) -> pd.DataFrame:
df = pd.read_csv(path, dtype={"sku": str})
df["sku"] = df["sku"].str.strip().str.upper()
df["price"] = df["price"].replace(r"[\$,]", "", regex=True).astype(float)
return df.dropna(subset=["sku"])
def merge(api_rows: list, client: pd.DataFrame) -> pd.DataFrame:
api_df = pd.DataFrame(api_rows)
api_df["sku"] = api_df["sku"].astype(str).str.strip().str.upper()
merged = client.merge(api_df, on="sku", how="left", indicator=True)
unmatched = (merged["_merge"] == "left_only").sum()
print(f"Merged {len(merged)} rows; {unmatched} SKUs unmatched")
return merged.drop(columns="_merge")
Reference solution: the LLM summary
The AI feature from Post 5, wired to the merged data — prompt with the actual diff, parse defensively, cap the cost:
# summarize.py
import os
import requests
def summarize_changes(merged) -> str:
diffs = merged[merged["price"] != merged["list_price"]].head(20)
prompt = ("Draft a short client email summarizing these price changes "
"(SKU, our price, supplier list price):\n"
+ diffs[["sku", "price", "list_price"]].to_csv(index=False))
r = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={"Authorization": "Bearer " + os.environ["OPENAI_API_KEY"]},
json={"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": prompt}],
"temperature": 0.2, "max_tokens": 250},
timeout=60,
)
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"]
Reference solution: the dashboard
The Streamlit app from Post 3 — metrics, filters, table, and the summary button. Note it reads the pre-built merged CSV so the demo never waits on the API:
# app.py — run with: streamlit run app.py
import streamlit as st
import pandas as pd
st.title("Northwind — Price Monitor")
merged = pd.read_csv("data/merged.csv") # built by pipeline.py, cached below
st.metric("SKUs tracked", f"{len(merged):,}")
category = st.sidebar.selectbox("Category", ["All"] + sorted(merged["category"].dropna().unique()))
view = merged if category == "All" else merged[merged["category"] == category]
st.dataframe(view.head(50))
if st.button("Generate client summary"):
with st.spinner("Drafting summary..."):
st.write(summarize_changes(merged))
Reference solution: ship it
The Post 6 packaging — a Dockerfile generated from the project, plus the two commands the reviewer will run. Reviewers will run it cold; test that path yourself:
from pathlib import Path
Path("Dockerfile").write_text("""\
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
EXPOSE 8501
CMD ["streamlit", "run", "app.py", "--server.port=8501", "--server.address=0.0.0.0"]
""")
print("Dockerfile written")
# docker build -t northwind-demo .
# docker run -p 8501:8501 --env-file .env northwind-demo
# Dockerfile written
What "great" looks like vs "passes"
- Passes: all five brief items work when the reviewer follows the README.
- Great: the README explains decisions (why cursor pagination, why left merge, what you'd do with 8 hours); errors produce helpful messages instead of tracebacks; the dashboard has a one-line "how to use this" at the top.
- Instant fail: secrets in the repo, the demo only works on your laptop, or the LLM feature hallucinates numbers with no grounding.
Key takeaways
- FDE take-homes score working software + judgment + communication — the rubric above is the real one.
- Time-box ruthlessly: a finished rough demo beats a polished half-demo.
- Structure the solution as fetch → clean → merge → summarize → dashboard → ship — one function per step.
- The demo must survive a cold run by a stranger: Dockerfile, README, env-based secrets, pre-built data.
- Write the README as decisions, not just commands — that's what separates "great" from "passes".
You've completed the Python for FDE track! Browse every tutorial on the Python topic page. Next: Python for Java Developers, then the Python Learning Roadmap 2026 hub post tying all eight tracks into one guided path.
Comments
Post a Comment