Classification Models
Predicting Categories
Most ML problems in business are classification problems: will a loan default, will a customer churn, is this transaction fraud? This lesson covers the three models you will reach for most often, logistic regression, decision trees and random forests, plus the confusion matrix, the tool for understanding what a classifier gets right and wrong.
Our example is a loan default model of the kind that digital lenders such as Carbon, FairMoney and Renmoney build, and that banks like Access Bank use for credit scoring.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(3)
n = 3000
loans = pd.DataFrame({
"monthly_income": np.round(rng.lognormal(12.2, 0.6, n), -3),
"loan_amount": np.round(rng.uniform(20000, 500000, n), -3),
"tenure_months": rng.integers(1, 60, n),
"past_late_payments": rng.poisson(0.8, n),
"has_salary_account": rng.integers(0, 2, n),
})
risk = (loans["loan_amount"] / loans["monthly_income"]) * 0.9 + loans["past_late_payments"] * 0.8 \
- loans["tenure_months"] * 0.03 - loans["has_salary_account"] * 0.7 + rng.normal(0, 1, n)
loans["defaulted"] = (risk > np.quantile(risk, 0.8)).astype(int) # about 20% default
X = loans.drop(columns="defaulted")
y = loans["defaulted"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)
stratify=y keeps the same share of defaulters in the training and test sets. Always use it for classification, especially when one class is rare.
1. Logistic Regression
Despite its name, logistic regression is a classification model. It computes a weighted sum of the features (like linear regression) and passes it through the sigmoid function, which squashes any number into a probability between 0 and 1.
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
logreg = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
logreg.fit(X_train, y_train)
proba = logreg.predict_proba(X_test)[:, 1] # probability of default
pred = logreg.predict(X_test) # 0/1 using a 0.5 threshold
Strengths: fast, stable, produces well-calibrated probabilities, and its coefficients are explainable. That is why it remains the standard for regulated credit scoring. Weaknesses: it can only draw a straight (linear) boundary between the classes unless you engineer features.
StandardScaler puts every feature on the same scale, which logistic regression needs to train well. Wrapping both in a pipeline ensures the scaling learned on the training data is applied identically to new data.
2. Decision Trees
A decision tree learns a flowchart of yes/no questions:
Is past_late_payments > 2?
├── Yes → Is has_salary_account = 1?
│ ├── Yes → Low risk
│ └── No → HIGH RISK
└── No → Is loan_amount / income > 3?
├── Yes → Medium risk
└── No → Low risk
from sklearn.tree import DecisionTreeClassifier, plot_tree
import matplotlib.pyplot as plt
tree = DecisionTreeClassifier(max_depth=3, random_state=0)
tree.fit(X_train, y_train)
plt.figure(figsize=(16, 7))
plot_tree(tree, feature_names=X.columns, class_names=["Repaid", "Default"], filled=True)
plt.show()
Strengths: easy to explain to non-technical colleagues, captures non-linear patterns and interactions, and needs no scaling.
Weaknesses: a deep tree memorises the training data (overfits). max_depth and min_samples_leaf keep it in check.
3. Random Forests
A random forest trains hundreds of decision trees, each on a random sample of the rows and features, and lets them vote. Individual trees make different mistakes, and averaging them cancels much of the noise. This is called an ensemble.
from sklearn.ensemble import RandomForestClassifier
forest = RandomForestClassifier(n_estimators=300, max_depth=8, random_state=0, n_jobs=-1)
forest.fit(X_train, y_train)
importances = pd.Series(forest.feature_importances_, index=X.columns).sort_values(ascending=False)
importances.plot(kind="barh", title="Feature importance")
Strengths: strong accuracy with little tuning, robust to outliers, and provides feature importance. Weaknesses: slower, larger, and harder to explain than a single tree.
Gradient boosting models (XGBoost, LightGBM, CatBoost) build trees one after another, each fixing the previous ones' errors. They often win Kaggle and Zindi competitions on tabular data, and they are a natural next step once you are comfortable with random forests.
The Confusion Matrix
Accuracy alone hides what kind of mistakes a model makes. The confusion matrix shows all four outcomes:
| Predicted: Repaid | Predicted: Default | |
|---|---|---|
| Actual: Repaid | True Negative (TN) | False Positive (FP): a good customer rejected |
| Actual: Default | False Negative (FN): a bad loan approved | True Positive (TP) |
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay, classification_report
for name, m in [("Logistic", logreg), ("Tree", tree), ("Forest", forest)]:
print(name)
print(classification_report(y_test, m.predict(X_test), target_names=["Repaid", "Default"]))
ConfusionMatrixDisplay.from_estimator(forest, X_test, y_test, display_labels=["Repaid", "Default"])
plt.show()
For a lender, a false negative (approving someone who defaults) loses the whole loan. A false positive (rejecting a good customer) loses the interest they would have paid. The two errors have very different costs, which is why the next lesson goes beyond accuracy.
Which Model Should You Choose?
| Situation | Start with |
|---|---|
| You need to explain every decision to a regulator | Logistic regression |
| You need to explain the logic to business stakeholders | A shallow decision tree |
| You want strong accuracy on tabular data quickly | Random forest |
| You are chasing the best possible score | Gradient boosting (XGBoost, LightGBM) |
In practice: start with logistic regression as a baseline, then check whether a forest or boosting model beats it by enough to justify the extra complexity.
Try it yourself
Key Takeaways
- Logistic regression outputs probabilities through the sigmoid function; it is fast and explainable, and it is the standard baseline for credit scoring.
- Decision trees learn yes/no rules that are easy to explain, but deep trees overfit unless you limit max_depth.
- Random forests average many randomised trees for strong, robust accuracy, and gradient boosting (XGBoost, LightGBM) often performs best on tabular data.
- Use stratify=y when splitting, and use pipelines with StandardScaler for models that are sensitive to feature scale.
- The confusion matrix separates true and false positives and negatives, because different errors carry very different business costs.
Quick Quiz
1.What does logistic regression output before a threshold is applied?
2.Why does a random forest usually generalise better than a single deep decision tree?
3.In a loan default model, what is a false negative?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx