LLM Integration in Python: Adding AI Features to Client Solutions
Part 5 of the Python for FDE track. Last updated: October 2026.
Half of today's FDE engagements include the sentence "and can it do something with AI?" The good news: adding LLM features to a client solution is mostly API integration — the same requests skills from Post 2, pointed at a chat endpoint. This post covers the chat API, structured JSON output, and a complete RAG pattern (retrieval-augmented generation) built with nothing but requests and numpy — no heavy frameworks, so you can explain every line to the client.
Calling a chat API from Python
A chat completion is a POST with a messages list. The system message sets the role and rules; user messages carry the actual request. Keep the helper tiny and reuse it everywhere:
import os
import requests
OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]
BASE = "https://api.openai.com/v1"
def chat(messages: list, model: str = "gpt-4o-mini", **kwargs) -> str:
r = requests.post(
f"{BASE}/chat/completions",
headers={"Authorization": f"Bearer {OPENAI_API_KEY}"},
json={"model": model, "messages": messages, **kwargs},
timeout=60,
)
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"]
reply = chat([
{"role": "system", "content": "You are a support assistant for Acme Corp."},
{"role": "user", "content": "Summarize the refund policy in two sentences."},
])
print(reply)
Structured output: JSON mode
Client features need data, not prose. JSON mode forces the model to reply with parseable JSON — then validate it, because models still misbehave:
import json
raw = chat(
[{"role": "system", "content": "You output valid JSON and nothing else."},
{"role": "user", "content": (
"Extract order details as JSON with keys: order_id (string), "
"total (number), items (array of strings).\n\n"
"Email: 'Order ORD-1042 shipped: 2x Widget Pro, total $84.50'")}],
response_format={"type": "json_object"},
)
order = json.loads(raw) # raises on malformed output — catch it upstream
print(order["order_id"], order["total"], order["items"])
# ORD-1042 84.5 ['Widget Pro', 'Widget Pro']
Defensive parsing: salvage, don't crash
In a live demo, a JSON parse error is a dead screen. Wrap parsing with a last-resort salvage that extracts the first {...} block — ugly, but it keeps the demo alive:
import json
def safe_json(raw: str) -> dict:
try:
return json.loads(raw)
except json.JSONDecodeError:
start = raw.find("{") # last-resort salvage
end = raw.rfind("}") + 1
return json.loads(raw[start:end])
print(safe_json('Sure! {"order_id": "ORD-9", "total": 12.0}'))
# {'order_id': 'ORD-9', 'total': 12.0}
RAG part 1: embeddings and chunking
RAG grounds the model in the client's own documents: embed doc chunks into vectors, retrieve the chunks most similar to the question, and answer from those. First the two primitives — embeddings via the API, chunking in pure Python:
import numpy as np
def embed(texts: list) -> np.ndarray:
r = requests.post(
f"{BASE}/embeddings",
headers={"Authorization": f"Bearer {OPENAI_API_KEY}"},
json={"model": "text-embedding-3-small", "input": texts},
timeout=60,
)
r.raise_for_status()
return np.array([d["embedding"] for d in r.json()["data"]])
def chunk(text: str, size: int = 120) -> list:
words = text.split()
return [" ".join(words[i:i + size]) for i in range(0, len(words), size)]
RAG part 2: cosine-similarity retrieval
Retrieval is one line of linear algebra: cosine similarity between the question vector and every chunk vector, then take the top-k. No vector database needed at demo scale:
def retrieve(query: str, chunks: list, vecs: np.ndarray, k: int = 3) -> list:
q = embed([query])[0]
scores = vecs @ q / (np.linalg.norm(vecs, axis=1) * np.linalg.norm(q))
top = np.argsort(scores)[-k:][::-1]
return [chunks[i] for i in top]
docs = ["Acme's refund policy allows returns within 30 days with receipt. "
"Refunds are issued to the original payment method within 5 business days."]
chunks = [c for d in docs for c in chunk(d)]
print(retrieve("how long do refunds take?", chunks, embed(chunks)))
# ["Acme's refund policy allows returns within 30 days with receipt. ..."]
RAG part 3: answer with context
Assemble the pipeline: retrieve the relevant chunks, stuff them into the prompt, and instruct the model to answer only from the context. That "I don't know" instruction is what stops hallucinations in front of clients:
def ask_with_docs(question: str, docs: list) -> str:
chunks = [c for d in docs for c in chunk(d)]
context = "\n\n".join(retrieve(question, chunks, embed(chunks)))
return chat(
[{"role": "system", "content": (
"Answer using ONLY the context below. "
"If the answer is not in the context, say 'I don't know'.")},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}],
temperature=0.2, # low temperature: factual, not creative
max_tokens=300, # keep demos fast and cheap
)
print(ask_with_docs("How long do refunds take?", docs))
# Refunds are issued to the original payment method within 5 business days.
Key takeaways
- LLM features are API integration: POST a messages list, read back the content — Post 2's skills transfer directly.
- The system prompt sets role and rules; keep temperature low and max_tokens capped for demos.
- Client features need structured output — use JSON mode and validate/parse defensively.
- RAG = embed chunks → cosine-similarity retrieval → answer from context; numpy handles demo-scale retrieval with no extra infrastructure.
- The "say I don't know" instruction is what makes RAG trustworthy in front of a client.
Next in this series: Shipping the Demo: Docker and One-Click Deploys for Client Demos.
Comments
Post a Comment