Where Data Comes From
No Data, No Data Science
Every model starts with data, and in the real world nobody hands you a clean CSV. Part of your job is knowing where data lives, how to get access to it, and what problems each source usually brings. A data scientist at Paystack might combine transaction logs from a database, exchange rates from an API and survey results from a spreadsheet, all in one project.
There are four main sources you will work with.
1. Files: CSV, Excel and JSON
Flat files are the most common starting point, especially in Nigerian organisations where a lot of reporting still happens in Excel.
import pandas as pd
df = pd.read_csv("sales.csv")
df = pd.read_excel("branch_report.xlsx", sheet_name="Q1")
df = pd.read_json("orders.json")
Watch out for:
- Different delimiters (
;instead of,): usesep=";" - Numbers stored as text with commas, such as "1,250,000": use
thousands="," - Dates in the wrong format (DD/MM versus MM/DD): use
dayfirst=True - Merged cells and multi-row headers in Excel reports: use
skiprowsandheader
2. Databases
Most company data lives in relational databases such as PostgreSQL, MySQL or SQL Server. Banks like Access Bank and GTBank run core banking systems on large relational databases, and analysts query copies of them in a data warehouse (BigQuery, Snowflake, Redshift).
import sqlite3
import pandas as pd
conn = sqlite3.connect("company.db")
df = pd.read_sql("""
SELECT customer_id, SUM(amount) AS total_spend
FROM transactions
WHERE txn_date >= '2025-01-01'
GROUP BY customer_id
""", conn)
For PostgreSQL or MySQL you use the same pd.read_sql with a SQLAlchemy connection. SQL is not optional for data scientists: almost every job interview includes a SQL test.
3. APIs
An API (Application Programming Interface) lets you request data from another system over the internet. You send a request to a URL and get back structured data, usually JSON.
Examples:
- The World Bank API returns economic indicators for Nigeria and every other country
- Exchange rate APIs return the latest Naira rates
- The Twitter/X, GitHub and Spotify APIs return social, code and music data
- Payment providers such as Paystack and Flutterwave expose transaction data to merchants through their APIs
APIs are great because data is fresh and structured. The trade-offs are rate limits, authentication keys and pagination. You will call a real API in the next lesson.
4. Web Scraping
When data is shown on a website but there is no API, you can scrape it: download the HTML and extract the values.
import pandas as pd
# read_html pulls every <table> on a page into a list of DataFrames
tables = pd.read_html("https://en.wikipedia.org/wiki/List_of_Nigerian_states_by_population")
states = tables[0]
For more complex pages, the requests and BeautifulSoup libraries let you target specific elements.
Scraping ethics and rules:
- Check the site's
robots.txtand terms of service - Do not overload servers: add delays between requests
- Never scrape personal data without a lawful basis. The Nigeria Data Protection Act (2023) and Europe's GDPR both apply to personal data
- Prefer an official API or download when one exists
Other Sources Worth Knowing
| Source | Examples |
|---|---|
| Open data portals | Nigerian Bureau of Statistics (NBS), CBN statistics, data.gov.uk, data.gov |
| Competition platforms | Kaggle, Zindi (an Africa-focused data science competition platform) |
| Surveys and forms | Google Forms, KoboToolbox (popular with NGOs in Nigeria) |
| Logs and events | App analytics (Mixpanel, Google Analytics), server logs |
| Sensors / IoT | Smart meters, GPS trackers on logistics fleets |
Primary vs Secondary Data
- Primary data is collected by you for your specific question, such as a survey of Lagos commuters. It is expensive but targeted.
- Secondary data was collected by someone else for another purpose, such as NBS inflation figures. It is cheap and fast, but you must understand how it was collected and what its limitations are.
Before using any dataset, ask: who collected it, how, when, and why? The answers tell you what biases and gaps to expect.
Try it yourself
Key Takeaways
- The four main data sources are files (CSV, Excel, JSON), databases queried with SQL, APIs that return JSON, and web scraping.
- Each source has typical pitfalls: wrong delimiters and date formats in files, access rights in databases, rate limits and keys in APIs, and legal limits on scraping.
- SQL is a core data science skill because most company data lives in relational databases and warehouses.
- Scrape responsibly: check robots.txt, avoid personal data, and respect laws such as the Nigeria Data Protection Act 2023 and GDPR.
- Before using any dataset, ask who collected it, how, when and why, so you understand its biases and gaps.
Quick Quiz
1.A public website shows a table of data but offers no API or download. What is the most appropriate first step before scraping it?
2.What is the main difference between primary and secondary data?
3.An Excel export shows amounts like "1,250,000" and pandas reads the column as text. Which read_csv argument fixes this?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx