Python Basics for Data Analysts
Why Python for Data Analysis?
Python has become the dominant language for data analysis, and for good reason. It combines the readability of a general-purpose language with a powerful ecosystem of data-specific libraries. Where Excel struggles with more than a few hundred thousand rows, Python handles millions. Where SQL excels at querying databases, Python excels at transformation, modelling, and visualisation.
For data analysts, Python adds capabilities that Excel and SQL alone cannot provide:
- Statistical analysis and hypothesis testing
- Machine learning and predictive modelling
- Automated data pipelines and scheduled reports
- Advanced visualisations and interactive dashboards
- Text analysis and natural language processing
- Connecting to APIs and web scraping
You do not need to become a software engineer to benefit from Python. Learning the fundamentals and the pandas library will already transform what you can do as an analyst.
Setting Up Your Python Environment
Anaconda (Recommended for Beginners)
Anaconda is a Python distribution that includes Python, Jupyter Notebook, and most data science libraries pre-installed.
- Download Anaconda from anaconda.com
- Install it (accept default settings)
- Open Anaconda Navigator
- Launch Jupyter Notebook or JupyterLab
Python + pip (More Control)
Alternatively: install Python from python.org, then install libraries via pip:
pip install pandas numpy matplotlib seaborn jupyter
Google Colab (No Installation)
Google Colaboratory (colab.research.google.com) provides a free Jupyter Notebook environment in the browser with most data science libraries pre-installed. Excellent for learning without setup overhead.
Python Data Types Relevant to Analysis
# Integers and floats
rows = 1500 # integer
revenue = 84250.75 # float
# Strings
country = "Nigeria"
column_name = 'total_amount'
# Booleans
is_active = True
has_orders = False
# None (represents missing/null values)
discount = None
# Check type
type(revenue) # <class 'float'>
Lists and Dictionaries
Lists (ordered, mutable sequences)
countries = ["Nigeria", "Ghana", "Kenya", "South Africa"]
amounts = [125.50, 340.00, 89.99, 210.75]
# Accessing elements (0-indexed)
countries[0] # "Nigeria"
countries[-1] # "South Africa" (last element)
countries[1:3] # ["Ghana", "Kenya"] (slice)
# List operations
countries.append("Egypt")
len(countries) # 5
sorted(amounts) # sorted copy (does not modify original)
Dictionaries (key-value pairs)
order = {
"order_id": "ORD-001",
"customer": "Emeka Okafor",
"amount": 450.00,
"status": "completed"
}
# Accessing values
order["amount"] # 450.00
order.get("notes") # None (safer than order["notes"] which raises KeyError)
# Updating
order["status"] = "shipped"
Control Flow
Conditionals
amount = 750
if amount >= 1000:
tier = "VIP"
elif amount >= 500:
tier = "Gold"
elif amount >= 100:
tier = "Standard"
else:
tier = "Low Value"
print(tier) # "Gold"
Loops
# For loop over a list
for country in countries:
print(f"Processing: {country}")
# For loop with range
for i in range(5):
print(i) # 0, 1, 2, 3, 4
# While loop
count = 0
while count < 3:
print(count)
count += 1
List Comprehensions (Pythonic shorthand)
# Generate a list using a compact expression
amounts = [120, 50, 340, 780, 90]
# Equivalent to: for each amount, check if it is above 100
above_threshold = [a for a in amounts if a > 100]
# [120, 340, 780]
# Transform: calculate tax for each amount
with_tax = [a * 1.2 for a in amounts]
# [144.0, 60.0, 408.0, 936.0, 108.0]
Functions
def categorise_order(amount):
"""Assign a size category to an order based on amount."""
if amount >= 1000:
return "Large"
elif amount >= 200:
return "Medium"
else:
return "Small"
# Call the function
categorise_order(750) # "Medium"
categorise_order(50) # "Small"
Python for Data Analysis: The Library Ecosystem
The power of Python for analysis comes from its libraries:
pandas: Data manipulation and analysis. Works with tabular data (DataFrames). The most important library for analysts.
NumPy: Numerical computing, array operations, mathematical functions. Used under the hood by pandas.
Matplotlib: Base plotting library. Full control over chart appearance.
Seaborn: Statistical visualisation built on matplotlib. Easier to use for common analytical charts.
Plotly: Interactive charts. Excellent for dashboards.
scikit-learn: Machine learning. Regression, classification, clustering.
Import libraries at the top of your script or notebook:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
The aliases (pd, np, plt, sns) are universal conventions -- always use them.
Key Takeaways
- Python extends data analysis beyond what Excel and SQL can do: automation, statistical modelling, advanced visualisation, and API integration.
- Anaconda (anaconda.com) or Google Colab are the easiest ways to get started with a full Python data analysis environment.
- Python's core data types for analysts: integers, floats, strings, booleans, None, lists, and dictionaries.
- Control flow (if/elif/else, for loops, while loops) and list comprehensions are fundamental to writing analytical scripts.
- The pandas library is the foundation of Python data analysis -- it provides the DataFrame, which is to Python what the spreadsheet is to Excel.
Practice Exercise
Open a Jupyter Notebook (via Anaconda or Google Colab) and write Python code to:
# 1. Create a list of 5 sales amounts
amounts = [120, 450, 89, 730, 200]
# 2. Calculate total and average manually
total = sum(amounts)
average = total / len(amounts)
print(f"Total: {total}, Average: {average:.2f}")
# 3. Use a list comprehension to filter amounts over 200
high_value = [a for a in amounts if a > 200]
print(f"High value orders: {high_value}")
# 4. Write a function to categorise each amount
def categorise(amount):
if amount >= 500:
return "High"
elif amount >= 200:
return "Medium"
else:
return "Low"
# 5. Apply the function to each amount
categories = [categorise(a) for a in amounts]
print(categories)
Try it yourself
Key Takeaways
- Python extends analytical capabilities beyond Excel and SQL: millions of rows, automation, statistical modelling, and advanced visualisation.
- Anaconda or Google Colab provide the easiest way to start a Python data analysis environment with all necessary libraries pre-installed.
- Core Python data structures for analysts: lists (ordered sequences) and dictionaries (key-value pairs like spreadsheet rows).
- List comprehensions provide a concise, Pythonic way to filter and transform lists: [expression for item in list if condition].
- The pandas library (imported as pd) is the foundation of Python data analysis -- learning it is the priority after Python basics.
Quick Quiz
1.What is the primary advantage of Python over Excel for data analysis?
2.What does the following Python list comprehension return? [x * 2 for x in [1, 2, 3, 4, 5] if x > 2]
3.What is the universal import alias convention for the pandas library?
4.What does dict.get(key, default) do differently from dict[key]?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx