Scaling Your Backend Application
Scaling Your Backend Application
When your API grows from handling 10 requests per second to 10,000, the architecture that worked before may buckle. Scaling is the practice of ensuring your system remains fast, reliable, and available as demand grows.
Vertical vs Horizontal Scaling
Vertical scaling (scaling up) means giving your server more resources — more RAM, more CPU cores, a faster disk. It is simple but has limits. The biggest server you can buy still has a ceiling.
Horizontal scaling (scaling out) means adding more server instances and distributing load between them. This is how every major platform (Netflix, Paystack, Twitter) handles millions of users. It has no hard ceiling — you add more machines as demand grows.
Node.js and the Cluster Module
Node.js runs on a single thread by default, which means it cannot use more than one CPU core. The built-in cluster module spawns multiple worker processes, each running on a separate core:
const cluster = require('cluster');
const os = require('os');
if (cluster.isPrimary) {
const numCPUs = os.cpus().length;
console.log(`Primary process spawning ${numCPUs} workers`);
for (let i = 0; i < numCPUs; i++) {
cluster.fork(); // Start a worker for each CPU core
}
cluster.on('exit', (worker) => {
console.log(`Worker ${worker.process.pid} died — restarting`);
cluster.fork(); // Auto-restart crashed workers
});
} else {
// Each worker runs the Express server independently
const app = require('./app');
app.listen(process.env.PORT || 3000, () => {
console.log(`Worker ${process.pid} started`);
});
}
PM2 — Process Manager for Production
PM2 (Process Manager 2) is the standard way to run Node.js in production. It handles clustering, auto-restart on crash, log management, and zero-downtime reloads.
npm install -g pm2
# Start app using all CPU cores
pm2 start index.js -i max
# View running processes
pm2 list
# View logs
pm2 logs
# Zero-downtime restart (for deployments)
pm2 reload all
# Auto-start on server reboot
pm2 startup
pm2 save
Stateless API Design
For horizontal scaling to work, your API must be stateless — each request must contain all information needed to process it, and the server must not store session data in memory.
// WRONG: Storing session data in memory (breaks with multiple instances)
const sessions = {}; // Each server instance has its own copy
app.post('/login', (req, res) => {
sessions[req.body.email] = { loggedIn: true }; // Only this instance knows
res.json({ ok: true });
});
// RIGHT: Use JWTs (stateless — no server storage needed)
app.post('/login', async (req, res) => {
const token = jwt.sign({ userId: user._id }, process.env.JWT_SECRET, { expiresIn: '24h' });
res.json({ token }); // Client holds all state in the token
});
Database Connection Pooling
Opening and closing a database connection for every request is slow. Connection pooling maintains a pool of open connections and reuses them.
// Mongoose manages connection pooling automatically
mongoose.connect(process.env.MONGODB_URI, {
maxPoolSize: 10, // Max 10 concurrent connections
minPoolSize: 2, // Keep at least 2 connections open
serverSelectionTimeoutMS: 5000,
});
Caching to Reduce Database Load
When the same data is requested frequently, cache it so the database is not queried every time:
const Redis = require('ioredis');
const redis = new Redis(process.env.REDIS_URL);
app.get('/api/products', async (req, res) => {
const cached = await redis.get('products');
if (cached) return res.json(JSON.parse(cached));
const products = await Product.find().sort({ createdAt: -1 }).limit(100);
await redis.setex('products', 300, JSON.stringify(products)); // Cache 5 min
res.json(products);
});
Load Balancers
A load balancer sits in front of multiple server instances and distributes incoming requests between them. Heroku, Render, and AWS handle this automatically when you scale to multiple dynos or instances.
Client → Load Balancer → Instance 1 (port 3000)
→ Instance 2 (port 3001)
→ Instance 3 (port 3002)
Key Takeaways
- Vertical scaling adds resources to one server; horizontal scaling adds more server instances.
- The Node.js cluster module uses all CPU cores by running multiple worker processes.
- PM2 manages Node.js processes in production — clustering, auto-restart, and log management.
- Stateless API design (JWT over sessions) is required for horizontal scaling.
- Connection pooling and Redis caching reduce database load as traffic grows.
Practice Exercise
- Use the cluster module to spawn one worker per CPU core in your
index.js. - Install PM2 globally and start your app with
pm2 start index.js -i max. - Verify your API is stateless — ensure no in-memory session storage exists.
- Add connection pooling options to your Mongoose connection with maxPoolSize: 10.
Try it yourself
Key Takeaways
- Horizontal scaling (adding more instances) is how major platforms handle high traffic — it has no hard ceiling.
- The Node.js cluster module spawns one worker process per CPU core, making full use of modern multi-core hardware.
- PM2 manages Node.js in production: clustering, auto-restart, log management, and zero-downtime reloads.
- Stateless API design using JWTs is a prerequisite for horizontal scaling — in-memory sessions break across instances.
- Connection pooling and Redis caching reduce database load as the number of concurrent requests grows.
Quick Quiz
1.What is the difference between vertical and horizontal scaling?
2.Why must an API be stateless to support horizontal scaling?
3.What does the Node.js cluster module do?
4.What is database connection pooling?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx