Types of Data and Data Sources
Not All Data is the Same
Before you can analyse data, you need to understand what kind of data you are working with. Different types of data require different analytical approaches, different tools, and different statistical methods.
This lesson covers the main types of data and where data comes from. Both are foundational knowledge that every data analyst needs.
Quantitative vs Qualitative Data
The most fundamental distinction is between data you can measure numerically and data that describes qualities.
Quantitative Data (Numerical)
Quantitative data is data that can be counted or measured and expressed as a number.
Examples:
- A product's price: 12,500 NGN
- A customer's age: 34 years
- Monthly website visitors: 85,000
- Average delivery time: 2.3 days
Quantitative data can be further split:
Discrete: Countable, whole numbers. You cannot have 2.7 customers -- it is always a whole number. Examples: number of orders, number of products, number of employees.
Continuous: Can take any value within a range, including decimals. Examples: height, temperature, revenue (NGN 15,432.67).
Qualitative Data (Categorical)
Qualitative data describes qualities or categories that cannot be meaningfully measured as numbers.
Examples:
- A customer's gender: Male, Female, Other
- Product category: Electronics, Clothing, Food
- Customer satisfaction: Satisfied, Neutral, Dissatisfied
- Country of origin: Nigeria, Ghana, Kenya
Qualitative data is further divided:
Nominal: Categories with no meaningful order. Country, product colour, payment method.
Ordinal: Categories with a meaningful order but unequal gaps between them. Satisfaction ratings (Low, Medium, High), education level (Primary, Secondary, University).
Structured vs Unstructured Data
Another key distinction is whether data is organised in a predictable format.
Structured Data
Structured data is organised in rows and columns -- like a spreadsheet or database table. It is easy to query with SQL and analyse with spreadsheets.
Example: A table of customer transactions:
| Customer ID | Date | Amount | Category |
|---|---|---|---|
| C001 | 2024-01-15 | 8500 | Electronics |
| C002 | 2024-01-16 | 2200 | Clothing |
Semi-Structured Data
Has some structure but does not fit neatly into a table. JSON (JavaScript Object Notation) and XML are common examples. APIs typically return semi-structured data.
{
"customer": "Amara Okafor",
"orders": [
{ "id": "ORD-001", "amount": 8500 },
{ "id": "ORD-002", "amount": 3200 }
]
}
Unstructured Data
No predefined structure. This is the hardest to analyse but makes up about 80% of all data in the world.
Examples:
- Customer reviews and comments
- Social media posts
- Emails
- Images, audio, and video
- PDFs and documents
Analysing unstructured data often requires techniques from Natural Language Processing (NLP) and machine learning.
Primary vs Secondary Data Sources
Primary Data
Data you collect yourself, directly from the source.
| Method | Example |
|---|---|
| Surveys | Google Forms questionnaire sent to customers |
| Interviews | One-on-one conversations with users |
| Experiments | A/B test comparing two versions of a webpage |
| Observations | Watching how users navigate your app |
Advantages: Tailored to your exact needs, higher trust in accuracy. Disadvantages: Takes time and money to collect.
Secondary Data
Data collected by someone else that you access and use.
| Source | Example |
|---|---|
| Government data | Nigeria's National Bureau of Statistics reports |
| Industry reports | McKinsey, Deloitte, PwC research |
| Public datasets | Kaggle, Google Dataset Search, World Bank Open Data |
| Internal company data | CRM records, accounting software, website analytics |
Advantages: Already collected, often free or cheap, covers large populations. Disadvantages: May not match your exact question, may be outdated.
Internal vs External Data
Internal data lives within your organisation:
- Sales records from your CRM (Customer Relationship Management system)
- Website traffic from Google Analytics
- Financial records from your accounting software
- Customer support tickets from Zendesk
External data comes from outside your organisation:
- Market research reports
- Social media data
- Economic indicators
- Competitor pricing (scraped from websites)
Most analytics projects use a combination of both.
Common Data Formats You Will Encounter
| Format | Description | Common Use |
|---|---|---|
| CSV | Comma-Separated Values, plain text rows | Spreadsheet exports, database dumps |
| Excel (.xlsx) | Spreadsheet format | Business reports, finance data |
| JSON | JavaScript Object Notation | API responses, web data |
| SQL Database | Tables stored in a database | Production business data |
| Parquet | Efficient columnar format | Big data and data warehouses |
As a data analyst, you will spend a surprising amount of time converting between these formats and cleaning the data within them.
Try it yourself
Key Takeaways
- Quantitative data is numerical (discrete or continuous); qualitative data describes categories (nominal or ordinal).
- Structured data fits neatly into tables; unstructured data (like text or images) requires more advanced processing.
- Primary data is collected by you; secondary data is collected by others and reused for your analysis.
- About 80% of the world's data is unstructured, making text and NLP skills increasingly valuable.
- Common data formats include CSV, Excel, JSON, and SQL tables -- analysts regularly move between them.
Quick Quiz
1.A dataset column contains customer satisfaction ratings: 'Very Dissatisfied', 'Dissatisfied', 'Neutral', 'Satisfied', 'Very Satisfied'. What type of data is this?
2.What is the main advantage of primary data over secondary data?
3.Which of the following is an example of unstructured data?
4.A dataset stores the number of transactions per customer per month. Is this discrete or continuous?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx