Rate Limiting APIs: Algorithms and Implementations
Your API is under attack. Not by sophisticated hackers, but by a simple script that hammers your endpoints thousands of times per second, consuming resources and degrading service for legitimate users. Without rate limiting, a single misbehaving client can bring down your entire application. This guide explains the core algorithms and shows practical implementations in Node.js and Nginx to protect your API.
Why Rate Limiting Matters
Rate limiting controls how many requests a client can make in a given time window. It's your first line of defense against:
- Brute-force attacks on login or API key endpoints
- Denial-of-Service (DoS) from excessive requests
- Resource exhaustion from expensive operations like database queries or file processing
- Abuse of free tiers by scrapers or bots
Beyond security, rate limiting ensures fair usage and helps you enforce business rules, such as tiered pricing plans.
Rate Limiting Algorithms
Four algorithms dominate API rate limiting. Each has trade-offs in accuracy, memory usage, and burst handling.
1. Fixed Window
Count requests in fixed time intervals (e.g., 100 requests per minute). When the window resets, the counter resets.
- Pros: Simple, low memory (one counter per client).
- Cons: Burst at window boundaries. A client can send 100 requests at 12:00:59 and another 100 at 12:01:00, effectively 200 requests in two seconds.
2. Sliding Window
Track timestamps of each request and count how many fall within the last N seconds. This smooths out bursts.
- Pros: Accurate, no boundary spikes.
- Cons: Higher memory usage (store timestamps) or use a sliding window counter with weighted average.
3. Token Bucket
A bucket holds tokens. Tokens are added at a fixed rate. Each request consumes one token. If the bucket is empty, the request is denied.
- Pros: Allows bursts up to bucket size, then enforces average rate.
- Cons: Requires storing token count and last refill time per client.
4. Leaky Bucket
Requests enter a queue (bucket) and are processed at a constant rate. If the queue is full, requests are dropped.
- Pros: Smooths traffic to a steady rate, ideal for protecting downstream services.
- Cons: Adds latency; not suitable for real-time APIs.
| Algorithm | Burst Handling | Memory | Use Case |
|---|---|---|---|
| Fixed Window | Poor | Low | Simple APIs |
| Sliding Window | Good | Medium | General purpose |
| Token Bucket | Excellent | Medium | APIs with burst tolerance |
| Leaky Bucket | None | Medium | Smoothing traffic |
Implementing Rate Limiting in Node.js
We'll implement a token bucket rate limiter using Express and Redis for distributed state. Redis is essential when you have multiple server instances.
Step 1: Install dependencies
npm install express redis
Step 2: Create the rate limiter middleware
const redis = require('redis');
const client = redis.createClient();
async function tokenBucketLimiter(req, res, next) {
const key = `rate_limit:${req.ip}`;
const capacity = 10; // max tokens
const refillRate = 1; // tokens per second
const now = Date.now();
const data = await client.hGetAll(key);
let tokens = data.tokens ? parseFloat(data.tokens) : capacity;
let lastRefill = data.lastRefill ? parseInt(data.lastRefill) : now;
// Refill tokens based on elapsed time
const elapsed = (now - lastRefill) / 1000;
tokens = Math.min(capacity, tokens + elapsed * refillRate);
if (tokens < 1) {
return res.status(429).json({ error: 'Too many requests' });
}
tokens -= 1;
await client.hSet(key, {
tokens: tokens.toString(),
lastRefill: now.toString()
});
await client.expire(key, 60); // auto-cleanup
next();
}
app.use(tokenBucketLimiter);
This middleware checks and updates the token count atomically. For production, use Redis transactions or Lua scripts to avoid race conditions.
Implementing Rate Limiting in Nginx
Nginx offers built-in rate limiting with the limit_req module. It uses a leaky bucket algorithm.
Step 1: Define a rate limit zone
In http block of nginx.conf:
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
This creates a 10MB zone named api that allows 10 requests per second per IP.
Step 2: Apply the limit
In your location block:
location /api/ {
limit_req zone=api burst=20 nodelay;
proxy_pass http://backend;
}
burst=20 allows short bursts of up to 20 requests. nodelay processes burst requests immediately instead of queuing them.
Best Practices for API Rate Limiting
- Return proper status codes: Use
429 Too Many Requestsand includeRetry-Afterheader. - Identify clients correctly: Use API keys or user IDs instead of IP addresses when possible, as IPs can be shared (NAT) or spoofed.
- Distribute state: Use Redis or a similar store for multi-instance deployments.
- Log and monitor: Track rate limit hits to detect attacks and adjust thresholds. Tools like the Nginx Log Analyzer can help you analyze 429 responses and identify abusive IPs.
- Communicate limits: Document rate limits in your API docs and include headers like
X-RateLimit-Limit,X-RateLimit-Remaining.
FAQ
What is the difference between rate limiting and throttling?
Rate limiting blocks requests beyond a threshold, while throttling slows them down (e.g., by queuing or delaying). Rate limiting is binary; throttling is gradual.
Which rate limiting algorithm should I use?
For most APIs, token bucket offers a good balance: it allows bursts but enforces an average rate. If you need strict smoothing, use leaky bucket. For simplicity, fixed window works for low-traffic APIs.
How do I handle rate limiting for authenticated vs. anonymous users?
Apply stricter limits to anonymous users (e.g., by IP) and more generous limits to authenticated users (e.g., by API key). You can also implement tiered limits based on subscription plans.
Ready to analyze your API traffic? Use our Nginx Log Analyzer to parse logs, spot rate limit violations, and optimize your thresholds.