In 2002, a Russian system administrator named Igor Sysoev started working on a problem engineers had a name for but no tidy answer to. They called it the C10K problem: how do you get one server to hold 10,000 open connections at once without running out of memory? Most web servers of that era gave every visitor a dedicated process or thread. That works when traffic is light. When it grows, the server spends more effort keeping track of visitors than serving them.
Sysoev's answer was Nginx, first released to the public in 2004. Over twenty years later, the Nginx web server still sits in front of a large share of the internet, from personal blogs to some of the busiest apps in the world.
Here is the part that catches teams off guard. A fresh Nginx install is set up to be safe on a small machine, not fast on a big one. The built-in default lets each worker process hold just 512 connections, and many Linux packages only raise that to 768 or 1024 in the config file they ship. Those numbers are fine for a side project. For an app that just got featured on a news site, they can become the reason visitors see an error page while your server's CPU sits mostly idle.
This guide walks through Nginx performance tuning in the order that tends to pay off, explains each term as it comes up, and spends real time on the messy parts: missing data, contradictory signals, and how the system behaves when pushed hard.
First, a quick look at how Nginx handles traffic
When Nginx starts, it launches one master process and a handful of worker processes. The master reads the configuration and manages the workers. The workers do the real work: accepting connections, reading requests, and sending responses.
Each worker runs what's called an event loop. Think of a waiter in a busy restaurant. A slow waiter takes one table's order and stands in the kitchen until the food is ready. A good waiter drops off the order and serves three other tables while the food cooks. An Nginx worker behaves like the good waiter, so one worker can juggle thousands of connections.
Most high-traffic setups use Nginx as a reverse proxy. That means Nginx sits in front of your application (written in Node.js, Python, PHP, Java, or anything else) and acts as the front door. Visitors talk to Nginx, and Nginx passes their requests to your app servers (called "upstream" servers in Nginx's documentation), then sends the answers back.
Once you have more than one app server behind that front door, Nginx also handles load balancing, which simply means deciding which server gets each request so no single machine gets buried.
How many sites run on Nginx? It depends who you ask
Search for Nginx's market share and you'll find numbers that disagree. The sources aren't wrong; they measure different things.
Nginx usage figures from two independent surveys
So is Nginx on 20% of sites or 42% of servers? Both. Netcraft counts every hostname it can reach, including parked domains, while W3Techs counts only sites where it could identify the server software.
These surveys also have a blind spot. When a site sits behind a service like Cloudflare, the response often names Cloudflare as the server, hiding the Nginx machine behind it. W3Techs reported in February 2026 that among Nginx sites whose reverse proxy service it could identify, 61.7% used Cloudflare. OpenResty, a popular platform built on top of Nginx, is also counted separately (Netcraft had it at 8.02% of sites in July 2026). The fair reading: Nginx's real footprint is probably larger than any headline figure, even as its visible share slips.
Why a few milliseconds are worth the effort
MARKET STATISTICS
The Deloitte study, Milliseconds Make Millions, looked at mobile data from 37 brands. One of its four metrics was server latency, measured as time to first byte (TTFB): the delay before the first piece of a reply arrives. That's the number Nginx tuning affects most.
One honest caveat: the gains appeared only when every page a shopper visited got faster on all four metrics together. Server tuning alone won't hand you an 8.4% lift. It removes one of four brakes.
The ITIC figures are self-reported by more than 1,000 firms, and 41% of enterprises put an hour of downtime between $1 million and over $5 million. Your numbers will differ, but a server that falls over during a spike costs far more than one that's a little slow.
The tuning ladder: seven steps in order of payoff
Cheap, high-impact changes come first. Jumping to clever tricks before fixing the basics is the most common way teams waste a weekend.
Step 1: Measure before you change anything
The default Nginx access log records who visited, what they asked for, and the status code. It doesn't record how long anything took. Without timing, you can't tell whether Nginx is slow, your app is slow, or the visitor's phone has a weak signal.
Add a log format that captures timing:
The request time (rt) is the total time Nginx spent on a request. The upstream connect time (uct) is how long opening a connection to your app took. The upstream header time (uht) is how long your app took to start answering, and the upstream response time (urt) is how long it took to finish. The cache field shows whether the answer came from Nginx's cache.
Also turn on the built-in status page (the stub_status module). It reports active connections plus two counters, "accepts" and "handled." If they drift apart, Nginx is turning connections away because it has hit a limit.
Step 2: Size your workers and connections
Three settings decide how many visitors Nginx can hold at once.
▪ The worker_processes setting tells Nginx how many workers to start. "Auto" gives one per CPU core, right for nearly everyone.
▪ The worker_connections setting is the ceiling per worker. It counts every connection, including the ones Nginx opens to your app.
▪ The worker_rlimit_nofile setting raises the number of open files each worker may have. On Linux, every network connection counts as an open file, and many systems default to a limit of 1024.
Now apply the two-connections rule. A 4-core server with the shipped setting of 1024 can hold 4 × 1024 = 4,096 connections, or roughly 2,048 visitors at once. That sounds like plenty until you remember that browsers often open several connections each, and slow networks hold them open longer.
Watch the error log for the line "worker_connections are not enough." If it appears, raise the setting and the open-file limit together. Raising one alone just moves the error somewhere less obvious.
Step 3: Reuse connections on both sides
Opening a connection costs time. For HTTPS, both sides first exchange several messages (a TLS handshake) before any page data moves. Reusing an open connection, which Nginx calls keepalive, skips all of that.
On the visitor side, Nginx keeps idle connections open for 75 seconds by default (keepalive_timeout) and allows 1,000 requests per connection (keepalive_requests). On a very busy server, shortening the timeout frees slots that idle visitors are holding.
The app side is where the gains have usually hidden. Until March 2026, Nginx talked to upstream servers over HTTP/1.0 by default, closing the connection after every request. Teams had to fix it by hand, and many never knew.
Nginx 1.29.7, released on 24 March 2026, made HTTP/1.1 with keepalive the upstream default, caching up to 32 idle connections per worker. The change carried into the 1.30 stable series.
The new default has an edge case. Without a keepalive line of your own, Nginx uses "local" mode, so cached connections aren't shared between location blocks. Write a keepalive line without the word "local" and they are shared, as in older versions.
How do you know keepalive isn't working? If the upstream connect time in your log stays above zero, connections are being opened fresh. Thousands of sockets stuck in a state called TIME_WAIT is another clue. Widening Linux's outgoing port range buys time but doesn't fix the cause.
Step 4: Cache what you can, even for one second
The fastest request your app ever handles is the one it never sees. Nginx can store copies of responses and hand them out directly.
Micro-caching works surprisingly well for busy pages that change often. If your homepage gets 1,000 requests per second and Nginx caches it for one second, your app builds the page about once per second instead of 1,000 times. Nobody notices content that's a second old.
Three of those lines handle situations that only show up at scale.
▪ The proxy_cache_lock line prevents a "cache stampede." Without it, when a popular page expires, every visitor arriving in that instant hits your app at once. With it, one request rebuilds the page while the others wait briefly.
▪ The proxy_cache_use_stale line serves the slightly old copy if your app is down, timing out, or refreshing. During an outage, visitors still see a page.
▪ The bypass and no-cache lines skip caching for logged-in visitors, identified here by a session cookie. Change the cookie name to match your app.
The edge case that bites people is personal data. Cache a page showing "Hello, Priya" carelessly and the next visitor sees Priya's name. Nginx won't cache a response that sets a cookie, but pages that only read one get no such protection, so always bypass the cache for requests carrying a login cookie.
Step 5: Get buffering right, especially for streaming
Buffering means Nginx reads your app's whole response into memory quickly, frees the app for its next request, then feeds the visitor at whatever speed their connection allows. It's on by default, and it's one of the biggest reasons a reverse proxy in front of a slow app helps so much. Without it, a visitor on shaky mobile data could tie up your app for ten seconds.
The exception is streaming. Live chat, AI assistants that type answers word by word, and server-sent event feeds need small pieces of data to arrive immediately. With buffering on, those pieces arrive in clumps, and users say the app "freezes, then dumps text." Turn buffering off for those routes only, with proxy_buffering off or by having your app send the header X-Accel-Buffering: no.
Uploads have a related issue. By default, Nginx rejects request bodies over 1 megabyte with a 413 error (client_max_body_size), and bodies larger than a small memory buffer get written to disk. For large uploads, raise the limit on that route only and keep temporary files on fast storage.
Step 6: Compress text and keep secure connections cheap
Compression (gzip) shrinks text files like HTML, CSS, and JavaScript. The default level is 1, the lightest. Levels around 4 to 6 usually shrink files noticeably more, while higher levels burn CPU for tiny extra savings. Skip images and videos, which are already compressed.
For HTTPS, store recent handshake results so returning visitors can skip part of the setup. Nginx's documentation says one megabyte of session cache holds about 4,000 sessions.
The sendfile setting lets the operating system copy files straight to the network, which helps static files a lot.
Step 7: Use rate limits to protect performance
Rate limiting caps how many requests one visitor can make per second. It's usually filed under security, yet it's also a performance tool: one badly written script hammering your search page can slow the app for everyone.
The burst value lets a visitor briefly exceed the limit, since one page load can fire a dozen requests.
Two edge cases deserve attention. Many people can share one public IP address, such as an office, a university, or a mobile carrier's customers, so a limit that's generous for one person can block a whole building. And if Nginx sits behind a CDN or cloud load balancer, every request appears to come from that service unless you configure Nginx's real IP module. Get this wrong and you rate limit your own CDN.
Choosing a load balancing method
With several app servers behind Nginx, the load balancing method decides who gets each request. Open-source Nginx offers several, and since version 1.29.6 (March 2026) it includes "sticky" sessions, which used to require the paid NGINX Plus.
How the methods compare
The least_conn note is easy to miss. Each worker keeps its own connection count unless you add a zone line to the upstream block. Without it, four workers can each make a sensible choice while together piling work onto one server.
The same thing happens across machines. Several Nginx servers in front of one app pool each see only their own traffic, so they can all pick the same "least busy" server at the same moment. The random two option adds just enough chance to break that up.
How Nginx decides a server is unhealthy
Load balancing in open-source Nginx relies on passive health checks, meaning it only notices a problem when a real request fails. By default, one failure within 10 seconds (max_fails=1, fail_timeout=10s) marks a server as unavailable for the next 10 seconds.
That real-time decision has consequences. With the defaults, one slow request during a brief app pause can pull a healthy server out of rotation, pushing its traffic onto the others. Set the threshold too high and a broken server keeps getting visitors. A max_fails of around 3 is a common middle ground, but test it against your own failures.
When a request fails with an error or timeout, Nginx tries the next server. Since version 1.9.13 it won't automatically retry requests like POST that can change data, since charging a card twice is worse than an error. You can cap retries with proxy_next_upstream_tries.
A few edge cases from the official documentation:
▪ If an upstream group has only one server, the failure settings are ignored and that server is never marked unavailable.
▪ Host names in an upstream block are normally looked up once, at startup. If your servers' IP addresses change, Nginx keeps using the old ones. The resolve option, in open-source Nginx since version 1.27.3, fixes this when paired with a resolver line.
▪ Active health checks, which send test requests on a schedule, remain a paid NGINX Plus feature, as does the least_time method.
When signals contradict each other
The hardest part of Nginx performance tuning is reading the evidence correctly when different numbers point in different directions. Here are the situations that confuse teams most often.
Reading mixed signals from Nginx
The 499 code isn't a standard HTTP code. Nginx uses it to log clients that hung up before getting an answer. A rise in 499s usually means the problem sits behind Nginx: the app got slow, visitors gave up, and Nginx recorded it. Tuning Nginx to "fix" 499s changes nothing.
Averages are another trap. A 200 millisecond average can hide 2% of visitors waiting five seconds, so check the 95th and 99th percentile values too.
What actually happens during a traffic spike
Suppose a product launch sends ten times your normal traffic within a few minutes. Here's how the overload usually unfolds.
Stage 1. Your app servers fill up first, and new requests wait in line.
Stage 2. Every waiting visitor ties up two Nginx connections, so the connection count climbs much faster than real throughput.
Stage 3. Requests time out, and Nginx retries them on other servers that are also overloaded. Engineers call this a retry storm.
Stage 4. Servers get marked as failed, pushing traffic onto fewer machines, which fail faster.
Stage 5. Nginx hits its connection ceiling, and new visitors can't get in the door.
The tuning ladder breaks this chain at several points. Stale cache serves visitors without touching the app, rate limits trim abusive traffic, capped retries stop the storm, and higher connection limits buy breathing room. An error_page with a friendly "we're busy" message beats a raw error.
One more limit sits below Nginx. Connections waiting to be accepted sit in a Linux queue capped by a kernel setting called somaxconn. Nginx asks for 511 by default, but kernels before version 5.4 capped it at 128, silently shrinking the queue on older systems.
Where Nginx reaches its limits
Nginx doesn't scale forever. The most detailed public account of its limits comes from Cloudflare, which ran Nginx at the core of its network for years before announcing a replacement called Pingora in 2022.
Cloudflare's engineers described two design problems. Each request is tied to one worker, so a CPU-heavy or disk-blocked request slows everything else on that worker. And because each worker keeps its own pool of reusable connections, adding workers lowers the chance of finding one ready to reuse.
Their numbers are striking. Cloudflare reported that Pingora used about 70% less CPU and 67% less memory than its old Nginx-based service for the same traffic. For one major customer, connection reuse went from 87.1% to 99.92%, which cut new connections to that customer's servers by 160 times.
Keep that in proportion. Cloudflare works at a scale almost no other company sees. For a startup or a mid-size business, a sensibly configured Nginx web server is very rarely the slowest part of the stack. Your database or app code will usually give out first.
A note for teams running Kubernetes
If your app runs on Kubernetes, check your "ingress controller," the component that routes outside traffic into your cluster. The community project Ingress NGINX was retired in March 2026. In January 2026, the Kubernetes Steering and Security Response Committees called it critical infrastructure for about half of cloud native environments and warned that staying on it leaves users open to attack.
Existing installs keep working without security fixes. The retirement covers that one project, not Nginx itself, and F5 maintains a separate NGINX-based controller. Kubernetes points users toward the newer Gateway API or another maintained controller. Re-check your tuning after switching, since each controller exposes settings differently.
Wrapping up
Good Nginx performance tuning mostly comes down to understanding what the server is doing. Measure first, size connections with the real math, reuse connections, cache carefully, and decide in advance how your setup should behave when things fail.
Most changes here take an afternoon. The harder part is the habit afterward: reading your logs after launches and incidents, and adjusting from evidence. If your team is building something that has to hold up under heavy traffic and nobody on it has tuned a production front door before, pairing with an experienced software development partner for the first round can save a painful outage. Either way, the Nginx web server rewards teams who take time to understand it.


