Machine Learning with Python: Roadmap and scikit-learn Setup
Part 1 of the Python for AI/ML track. Last updated: September 2026.
You finished the Python core track and the Data Science track — you can wrangle data with pandas and summarize it with statistics. Now the fun part: teaching the machine to learn patterns from data by itself. That is machine learning, and Python is its home turf.
What machine learning actually is
Traditional programming: you write the rules, data goes in, answers come out. Machine learning flips it: data and answers go in, rules come out. You show the algorithm thousands of examples ("this house sold for $650k"), and it figures out the pattern (price grows ~$300 per square foot). Then it applies that pattern to houses it has never seen.
Every ML project follows the same workflow:
- Collect data — rows of examples with known outcomes.
- Pick features — the columns the model gets to look at.
- Train — the algorithm adjusts itself to fit the examples.
- Evaluate — test it on examples it never saw during training.
- Predict — use it on brand-new data.
Supervised vs unsupervised learning
The two big families:
- Supervised learning — the data comes with answers (labels). Predicting house prices, spam detection, image captions. This is 90% of practical ML, and most of this track.
- Unsupervised learning — no labels, the algorithm finds structure: customer segments (clustering), anomaly detection, dimensionality reduction.
Within supervised learning there are two jobs: regression (predict a number — price, temperature) and classification (predict a category — spam or not, which species). Parts 2 and 3 cover each.
Setting up scikit-learn
scikit-learn is the standard ML library for tabular data — clean API, great docs, every classical algorithm. Install it in your virtual environment:
pip install scikit-learn pandas numpy
Your first model in ten lines
The classic Iris dataset: 150 flowers, 4 measurements each, 3 species. Train a decision tree to identify the species:
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
iris = load_iris()
X_train, X_test, y_train, y_test = train_test_split(
iris.data, iris.target, test_size=0.2, random_state=42)
model = DecisionTreeClassifier(random_state=42)
model.fit(X_train, y_train) # train: learn patterns from examples
print(model.score(X_test, y_test)) # 1.0 — perfect on unseen flowers
# Predict one new flower (sepal/petal measurements)
print(model.predict([X_test[0]])) # [1] → versicolor
print(iris.target_names[1]) # 'versicolor'
Notice the pattern you will repeat for the entire track: fit → predict → score.
The universal scikit-learn API
Almost every scikit-learn estimator works the same way:
model = SomeModel(hyperparameters...) # 1. choose the algorithm model.fit(X_train, y_train) # 2. train on labeled data model.predict(X_new) # 3. predict on new data model.score(X_test, y_test) # 4. evaluate (higher = better)
Learn this once and you can swap in logistic regression, random forests, or gradient boosting with a one-line change. That consistency is why scikit-learn is the best place to start.
Roadmap of this track
- Part 2 — Linear Regression: predicting numbers
- Part 3 — Classification: logistic regression and decision trees
- Part 4 — Evaluation done right: train-test split and cross-validation
- Part 5 — Feature engineering: scaling, encoding, pipelines
- Part 6 — Neural networks crash course with PyTorch
- Part 7 — End-to-end project: raw data to saved model
Key takeaways
- ML learns rules from labeled examples instead of you hand-coding them.
- Supervised = labeled data (regression for numbers, classification for categories); unsupervised = find structure without labels.
- scikit-learn's universal API is fit → predict → score — learn it once, reuse everywhere.
- Always evaluate on data the model never saw during training.
Next in this series: Linear Regression with scikit-learn: Predicting Numbers.
Comments
Post a Comment