Model Evaluation
The 99% Accurate Model That Catches Nothing
Imagine a bank where 1 in every 100 card transactions is fraudulent. A model that simply predicts "not fraud" for every transaction is 99% accurate, and completely useless: it never catches a single fraudster.
This is the accuracy paradox, and it is why data scientists rarely judge a classifier by accuracy alone. Fraud, loan default, churn and disease are all imbalanced problems where the class you care about is rare. You need metrics that focus on it.
Back to the Confusion Matrix
Every metric in this lesson comes from four counts:
| Predicted Negative | Predicted Positive | |
|---|---|---|
| Actual Negative | TN: correctly ignored | FP: false alarm |
| Actual Positive | FN: missed it | TP: correctly caught |
Take a fraud model tested on 10,000 transactions, of which 100 are actually fraudulent. It flags 120 transactions, and 80 of those are real fraud.
- TP = 80 (fraud caught)
- FP = 40 (legitimate transactions flagged)
- FN = 20 (fraud missed)
- TN = 9,860 (legitimate, correctly ignored)
The Core Metrics
Accuracy
(TP + TN) / total = (80 + 9,860) / 10,000 = 99.4% Share of all predictions that were right. It looks impressive, but it is dominated by the huge number of easy negatives.
Precision
TP / (TP + FP) = 80 / 120 = 66.7% When the model raises an alarm, how often is it right? Low precision means many false alarms, which annoys customers whose cards get blocked and wastes the fraud team's time.
Recall (sensitivity, true positive rate)
TP / (TP + FN) = 80 / 100 = 80% Of all the actual fraud, how much did we catch? Low recall means fraud slips through.
F1 score
2 x (precision x recall) / (precision + recall) = 72.7% The harmonic mean of precision and recall. It is only high when both are high, so it is a good single number for imbalanced problems.
Specificity
TN / (TN + FP) = 99.6% Of all the legitimate transactions, how many did we correctly leave alone?
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, classification_report
print(classification_report(y_test, y_pred))
The Precision-Recall Trade-off
A classifier outputs a probability, and you choose the threshold above which you predict "positive".
- Lower the threshold → flag more cases → recall rises, precision falls
- Raise the threshold → flag fewer cases → precision rises, recall falls
proba = model.predict_proba(X_test)[:, 1]
for t in [0.2, 0.35, 0.5, 0.65]:
pred = (proba >= t).astype(int)
print(t, round(precision_score(y_test, pred), 2), round(recall_score(y_test, pred), 2))
Which matters more depends on the cost of each error:
| Problem | Worse error | Prioritise |
|---|---|---|
| Card fraud at a bank | Missing fraud (FN) | Recall, with a human review queue |
| Cancer screening | Missing a disease (FN) | Recall |
| Spam filter | Hiding a real email (FP) | Precision |
| Loan approval at a digital lender | Both are costly | Tune the threshold to expected profit |
| Churn campaign with a limited budget | Wasting offers on loyal users (FP) | Precision at the top of the list |
ROC Curve and AUC
The ROC curve plots the true positive rate (recall) against the false positive rate at every threshold. AUC (area under the curve) summarises it in one number: 0.5 is random guessing and 1.0 is perfect. AUC measures how well the model ranks positives above negatives, regardless of the threshold you pick.
from sklearn.metrics import roc_auc_score, RocCurveDisplay
print("AUC:", roc_auc_score(y_test, proba))
RocCurveDisplay.from_predictions(y_test, proba)
For heavily imbalanced data, the precision-recall curve and average precision are often more informative than ROC AUC.
Overfitting and Underfitting
- Underfitting: the model is too simple to capture the pattern. Both training and test scores are poor.
- Overfitting: the model memorises the training data, including its noise. The training score is excellent, but the test score is much worse.
print("Train F1:", f1_score(y_train, model.predict(X_train)))
print("Test F1: ", f1_score(y_test, model.predict(X_test)))
A big gap between the two is the classic sign of overfitting. Fixes include simpler models (lower max_depth), more data, fewer or better features, and regularisation.
Cross-validation
A single train/test split can be lucky or unlucky. K-fold cross-validation trains and tests the model k times on different folds and averages the results:
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5, scoring="f1")
print(scores.mean(), scores.std())
Evaluation Checklist
- What is the business cost of a false positive versus a false negative?
- Choose the metric that reflects that cost (recall, precision, F1 or AUC)
- Compare against a baseline, such as predicting the majority class
- Check training versus test scores for overfitting
- Use cross-validation for a stable estimate
- Tune the threshold, not just the model
Try it: The lab below lets you adjust TP, FP, TN and FN directly. Recreate the "99% accurate but useless" model and watch recall collapse to zero.
Try it yourself
Key Takeaways
- Accuracy is misleading on imbalanced problems; a model that always predicts the majority class can be 99% accurate and useless.
- Precision asks how many alerts were right, recall asks how many real positives were caught, and F1 balances the two.
- The decision threshold trades precision against recall, so choose it based on the business cost of each type of error.
- ROC AUC measures ranking quality across all thresholds; for heavily imbalanced data, precision-recall curves are often more informative.
- A large gap between training and test scores signals overfitting; cross-validation gives a more reliable performance estimate.
Quick Quiz
1.Out of 100 actual fraud cases, a model catches 60. It also raises 40 false alarms. What are its recall and precision?
2.A churn model scores 98% F1 on the training data and 61% F1 on the test data. What is the most likely problem?
3.You lower a fraud model's decision threshold from 0.5 to 0.2. What usually happens?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx