Nginx Rate Limiting
Rate limiting caps how many requests a client can make in a period. It's a first line of defense against brute-force login attempts, credential stuffing, scrapers, abusive API clients, and accidental request floods from buggy integrations. It also protects backends from being overwhelmed. Nginx implements it efficiently at the edge with limit_req (request rate) and limit_conn (concurrent connections), before requests ever reach your application.
Configured well, rate limits are invisible to normal users and stop abuse cold. Configured carelessly, they block legitimate users behind shared IPs, fail to protect anything because the key is wrong, or behave inconsistently across multiple Nginx instances. See rate limiting for the algorithms and strategy in general.
TL;DR
limit_req_zonedefines a key (usually the client IP), a shared memory zone, and a rate (for example10r/s).limit_reqapplies it to a location.burstallows short spikes, andnodelayserves burst requests immediately instead of queueing them.- Nginx uses a leaky bucket: excess requests beyond rate + burst get 429 (configure
limit_req_status 429). - Key by API key or user ID for authenticated APIs, and by IP for anonymous endpoints. Apply stricter limits to login and signup.
limit_conncaps simultaneous connections per key (downloads, WebSockets).- Limits are per Nginx instance. For global limits across a fleet, use a shared store (API gateway, Redis-based limiter).
Quick Example
Core Concepts
The Leaky Bucket
limit_req implements a leaky bucket: requests drain at the configured rate, for example 10r/s, meaning one request every 100 ms. Requests arriving faster fill a queue of size burst:
- Without
burst: any request arriving sooner than the rate allows is rejected immediately. That's very strict: two requests 50 ms apart means the second fails. - With
burst=N: up to N excess requests are queued and released at the configured rate, which adds delay. - With
burst=N nodelay: up to N excess requests are served immediately, while their slots in the bucket still drain at the rate. Further excess is rejected. This suits browsers loading many resources at once. delay=M(two-stage limiting): the first M excess requests are served immediately, then the rest are queued up to the burst.
Rates can be per second (r/s) or per minute (r/m), which is good for sensitive endpoints like login.
Keys
The key decides who is limited:
Use $binary_remote_addr rather than $remote_addr, since it's smaller in memory. Behind proxies, configure the realip module so the key is the actual client IP, not the load balancer's. An empty key value excludes a request from limiting, which is how allowlists work.
Connection Limits
limit_conn_zone + limit_conn cap concurrent connections per key, which is useful for large downloads, streaming, and WebSockets, and against slowloris-style attacks. Combine with limit_rate to cap bandwidth per connection.
Responses and Observability
- Set
limit_req_status 429(Too Many Requests); the default is 503, which misleads clients and monitoring. - Add
Retry-Afterheaders via an error page or location for well-behaved clients. limit_req_log_levelcontrols log severity for rejections and delays, andlimit_req_dry_run onlogs what would be limited without enforcing it, which is perfect for tuning new limits safely.
Multiple Instances
Each Nginx instance keeps its own counters in shared memory, shared across its worker processes but not across servers. With 4 instances behind a load balancer, a client might get roughly 4× the configured rate. Options: divide rates by instance count, use consistent hashing of client IPs at the load balancer, or enforce precise global limits in an API gateway or a Redis-backed limiter.
Best Practices
Layer Limits by Sensitivity
Apply a generous site-wide per-IP limit to stop floods, strict limits on authentication endpoints (login, password reset, OTP verification), and per-key quotas on APIs matching your plans or SLAs.
Start in Dry-Run Mode
Enable limit_req_dry_run on, review logs for legitimate traffic that would be affected, tune rates and bursts, then enforce. Real traffic patterns (page loads, SPA bursts, retries) are burstier than intuition suggests.
Combine With Application-Level Controls
Nginx limits are coarse. For account-level protection (lockouts after failed logins, per-user quotas, CAPTCHA triggers), implement logic in the application or identity provider too. See API security.
Return Helpful Errors
Serve a clear 429 body (JSON for APIs) with Retry-After, and document limits for API consumers, so clients can back off gracefully instead of hammering retries.
Common Mistakes
Limiting the Load Balancer's IP
Without real_ip_header and set_real_ip_from, every request appears to come from the load balancer, so the whole site shares one bucket and legitimate users get 429s during normal traffic.
No Burst for Browser Traffic
Add burst with nodelay for page-serving locations, or exclude static assets from the limit.
Leaving the Default 503 Status
Clients and monitors treat 503 as a server outage. Use 429, so dashboards and SDKs can tell throttling apart from failures.
FAQ
What does burst and nodelay do in Nginx rate limiting?
burst defines how many requests above the steady rate can be accepted at once. Without nodelay they're queued and released at the configured rate; with nodelay they're processed immediately, while the bucket still enforces the long-term average. Requests beyond the burst are rejected.
What status code should rate-limited requests return?
429 Too Many Requests, ideally with a Retry-After header. Set limit_req_status 429; and limit_conn_status 429;, since Nginx defaults to 503.
How do I exclude certain IPs from rate limiting?
Use geo (or map) to mark trusted addresses, and map them to an empty key: requests with an empty key aren't counted. It's commonly used for internal networks, monitoring, and partner integrations.
Are Nginx rate limits shared across multiple servers?
No. Counters are kept per Nginx instance, shared among its worker processes. For consistent global limits across a fleet, use an API gateway or rate limiter backed by a shared store like Redis, or account for the number of instances when setting rates.
Related Topics
- Nginx — The web server overview
- Rate Limiting — Algorithms and strategies in general
- API Security — Protecting APIs from abuse
- API Gateway — Centralized quotas and limits
- Nginx Reverse Proxy — Recovering real client IPs
- Security — Defense in depth