Nginx Rate Limiting

Rate limiting caps how many requests a client can make in a period. It's a first line of defense against brute-force login attempts, credential stuffing, scrapers, abusive API clients, and accidental request floods from buggy integrations. It also protects backends from being overwhelmed. Nginx implements it efficiently at the edge with limit_req (request rate) and limit_conn (concurrent connections), before requests ever reach your application.

Configured well, rate limits are invisible to normal users and stop abuse cold. Configured carelessly, they block legitimate users behind shared IPs, fail to protect anything because the key is wrong, or behave inconsistently across multiple Nginx instances. See rate limiting for the algorithms and strategy in general.

TL;DR

Quick Example

Core Concepts

The Leaky Bucket

limit_req implements a leaky bucket: requests drain at the configured rate, for example 10r/s, meaning one request every 100 ms. Requests arriving faster fill a queue of size burst:

Rates can be per second (r/s) or per minute (r/m), which is good for sensitive endpoints like login.

Keys

The key decides who is limited:

Use $binary_remote_addr rather than $remote_addr, since it's smaller in memory. Behind proxies, configure the realip module so the key is the actual client IP, not the load balancer's. An empty key value excludes a request from limiting, which is how allowlists work.

Connection Limits

limit_conn_zone + limit_conn cap concurrent connections per key, which is useful for large downloads, streaming, and WebSockets, and against slowloris-style attacks. Combine with limit_rate to cap bandwidth per connection.

Responses and Observability

Multiple Instances

Each Nginx instance keeps its own counters in shared memory, shared across its worker processes but not across servers. With 4 instances behind a load balancer, a client might get roughly 4× the configured rate. Options: divide rates by instance count, use consistent hashing of client IPs at the load balancer, or enforce precise global limits in an API gateway or a Redis-backed limiter.

Best Practices

Layer Limits by Sensitivity

Apply a generous site-wide per-IP limit to stop floods, strict limits on authentication endpoints (login, password reset, OTP verification), and per-key quotas on APIs matching your plans or SLAs.

Start in Dry-Run Mode

Enable limit_req_dry_run on, review logs for legitimate traffic that would be affected, tune rates and bursts, then enforce. Real traffic patterns (page loads, SPA bursts, retries) are burstier than intuition suggests.

Combine With Application-Level Controls

Nginx limits are coarse. For account-level protection (lockouts after failed logins, per-user quotas, CAPTCHA triggers), implement logic in the application or identity provider too. See API security.

Return Helpful Errors

Serve a clear 429 body (JSON for APIs) with Retry-After, and document limits for API consumers, so clients can back off gracefully instead of hammering retries.

Common Mistakes

Limiting the Load Balancer's IP

Without real_ip_header and set_real_ip_from, every request appears to come from the load balancer, so the whole site shares one bucket and legitimate users get 429s during normal traffic.

No Burst for Browser Traffic

Add burst with nodelay for page-serving locations, or exclude static assets from the limit.

Leaving the Default 503 Status

Clients and monitors treat 503 as a server outage. Use 429, so dashboards and SDKs can tell throttling apart from failures.

FAQ

What does burst and nodelay do in Nginx rate limiting?

burst defines how many requests above the steady rate can be accepted at once. Without nodelay they're queued and released at the configured rate; with nodelay they're processed immediately, while the bucket still enforces the long-term average. Requests beyond the burst are rejected.

What status code should rate-limited requests return?

429 Too Many Requests, ideally with a Retry-After header. Set limit_req_status 429; and limit_conn_status 429;, since Nginx defaults to 503.

How do I exclude certain IPs from rate limiting?

Use geo (or map) to mark trusted addresses, and map them to an empty key: requests with an empty key aren't counted. It's commonly used for internal networks, monitoring, and partner integrations.

Are Nginx rate limits shared across multiple servers?

No. Counters are kept per Nginx instance, shared among its worker processes. For consistent global limits across a fleet, use an API gateway or rate limiter backed by a shared store like Redis, or account for the number of instances when setting rates.

Related Topics

References