The Data Science Workflow
From Problem to Prediction
Data science is not about running random algorithms on data. It is a structured, iterative process that starts with a business problem and ends with a system that delivers measurable value.
Understanding the full workflow before diving into individual techniques gives you the mental model to see how everything fits together. Let us walk through each stage.
Stage 1: Problem Definition
Every data science project begins with a question -- not a technical question, but a business one.
Bad framing: "Let us apply machine learning to our data." Good framing: "We are losing 15% of users within the first 30 days. Can we predict which users are likely to churn so we can intervene before they leave?"
The problem definition stage answers:
- What is the business problem?
- What would a good solution look like in concrete terms?
- What data might be available?
- Is this a prediction problem, a clustering problem, or something else?
- What would success look like? (Define the metric before you start)
Stage 2: Data Collection
Once you know what you are trying to predict, you collect the data that might help you predict it.
Data sources in a typical company:
- Databases: User tables, event logs, transaction records
- Third-party APIs: Weather, demographic, or economic data
- Web scraping: Publicly available data from websites
- Surveys and forms: User-reported information
- Sensors and devices: Internet of Things (IoT) data, mobile accelerometers
In practice, data scientists spend a lot of time negotiating access to data, understanding what data exists, and working with data engineers to get it into a usable form.
Stage 3: Exploratory Data Analysis (EDA)
Before building any model, explore the data to understand it deeply.
EDA involves:
- Checking the shape (rows and columns)
- Identifying missing values and deciding how to handle them
- Looking at distributions (is the data normally distributed? Skewed? Bimodal?)
- Identifying outliers
- Looking at correlations between features
- Plotting charts to see patterns visually
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
df = pd.read_csv('user_data.csv')
print(df.describe())
print(df.isnull().sum())
# Distribution of key feature
df['days_active'].hist(bins=30)
plt.title('Distribution of Days Active Before Churn')
plt.show()
# Correlation heatmap
sns.heatmap(df.corr(), annot=True, cmap='coolwarm')
plt.show()
Stage 4: Feature Engineering
Raw data is rarely in the right shape for a model. Feature engineering is the process of creating, transforming, and selecting the variables (features) that a model will learn from.
Examples of feature engineering:
- From a raw timestamp, extract: day of week, hour, days since signup
- From raw revenue, create: revenue percentile, revenue growth rate
- From text, create: word count, sentiment score, presence of specific keywords
- Normalise numerical features so they are on the same scale (0 to 1)
- Encode categorical variables as numbers (Lagos = 0, Abuja = 1, Kano = 2)
Feature engineering often has more impact on model performance than the choice of algorithm. Good features built from domain knowledge can dramatically improve predictions.
Stage 5: Model Selection and Training
Now you choose a model and train it. "Training" means fitting the model's parameters to your data so that it can make accurate predictions.
Common algorithms:
| Problem Type | Common Algorithms |
|---|---|
| Classification (predict a category) | Logistic Regression, Random Forest, XGBoost, Neural Networks |
| Regression (predict a number) | Linear Regression, Gradient Boosting, Neural Networks |
| Clustering (find natural groups) | K-Means, DBSCAN, Hierarchical Clustering |
| Recommendation | Collaborative Filtering, Matrix Factorisation |
You typically do not just pick one model -- you try several and compare them.
Training and Test Splits
To evaluate a model fairly, you split your data:
- Training set (typically 70-80%): The model learns from this
- Test set (typically 20-30%): The model is evaluated on this
The test set simulates unseen data. If the model performs well on training data but poorly on test data, it has overfit -- memorised the training examples rather than learning general patterns.
Stage 6: Model Evaluation
Choosing the right evaluation metric depends on your problem:
| Problem | Metric | What It Means |
|---|---|---|
| Binary classification | Accuracy, Precision, Recall, F1, AUC-ROC | How well does the model classify? |
| Regression | MAE, RMSE, R squared | How close are predictions to actual values? |
| Clustering | Silhouette score | How well-separated are the clusters? |
For a churn prediction model, you would care more about recall (catching as many churners as possible) than precision, because the cost of missing a churner (who then leaves) is higher than the cost of a false alarm.
Stage 7: Deployment
A model that never leaves a Jupyter Notebook creates no value. Deployment means putting the model into production where it can make real predictions.
Common deployment approaches:
- REST API: Wrap the model in a Flask or FastAPI server. The frontend or other services call the API with input data and get a prediction back.
- Batch processing: Run predictions overnight for all users and store results in a database.
- Embedded in product: A recommendation model running directly in an app.
Stage 8: Monitoring and Retraining
Models degrade over time. This is called model drift: the patterns in real-world data change but the model was trained on old data.
Monitoring involves:
- Tracking prediction accuracy over time
- Alerting when performance drops below a threshold
- Scheduling regular retraining with fresh data
A well-maintained model is a living system, not a one-time project.
The Iterative Nature of Data Science
These stages are not a strict sequence. Real data science is deeply iterative:
- EDA might reveal that you need more data (back to Stage 2)
- Model evaluation might show that different features are needed (back to Stage 4)
- Stakeholder feedback might reframe the problem (back to Stage 1)
This iteration is not a sign of failure -- it is how good data science works.
Try it yourself
Key Takeaways
- The data science workflow has 8 stages: problem definition, data collection, EDA, feature engineering, model training, evaluation, deployment, and monitoring.
- Problem definition is the most important stage -- a poorly defined problem produces useless results regardless of technique.
- Feature engineering often contributes more to model performance than the choice of algorithm.
- Train and test splits let you evaluate models on unseen data and detect overfitting.
- Deployed models require ongoing monitoring and periodic retraining as real-world data patterns change over time.
Quick Quiz
1.Why do data scientists split data into training and test sets?
2.What is feature engineering?
3.What is model drift?
4.For a churn prediction model, why would you prioritise recall over precision?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx