Nginx Performance Tuning for High-Traffic Applications

Nginx Performance Tuning for High-Traffic Applications

In 2002, a Russian system administrator named Igor Sysoev started working on a problem engineers had a name for but no tidy answer to. They called it the C10K problem: how do you get one server to hold 10,000 open connections at once without running out of memory? Most web servers of that era gave every visitor a dedicated process or thread. That works when traffic is light. When it grows, the server spends more effort keeping track of visitors than serving them.

Sysoev's answer was Nginx, first released to the public in 2004. Over twenty years later, the Nginx web server still sits in front of a large share of the internet, from personal blogs to some of the busiest apps in the world.

Here is the part that catches teams off guard. A fresh Nginx install is set up to be safe on a small machine, not fast on a big one. The built-in default lets each worker process hold just 512 connections, and many Linux packages only raise that to 768 or 1024 in the config file they ship. Those numbers are fine for a side project. For an app that just got featured on a news site, they can become the reason visitors see an error page while your server's CPU sits mostly idle.

This guide walks through Nginx performance tuning in the order that tends to pay off, explains each term as it comes up, and spends real time on the messy parts: missing data, contradictory signals, and how the system behaves when pushed hard.

First, a quick look at how Nginx handles traffic

When Nginx starts, it launches one master process and a handful of worker processes. The master reads the configuration and manages the workers. The workers do the real work: accepting connections, reading requests, and sending responses.

Each worker runs what's called an event loop. Think of a waiter in a busy restaurant. A slow waiter takes one table's order and stands in the kitchen until the food is ready. A good waiter drops off the order and serves three other tables while the food cooks. An Nginx worker behaves like the good waiter, so one worker can juggle thousands of connections.

Most high-traffic setups use Nginx as a reverse proxy. That means Nginx sits in front of your application (written in Node.js, Python, PHP, Java, or anything else) and acts as the front door. Visitors talk to Nginx, and Nginx passes their requests to your app servers (called "upstream" servers in Nginx's documentation), then sends the answers back.

Once you have more than one app server behind that front door, Nginx also handles load balancing, which simply means deciding which server gets each request so no single machine gets buried.

Field note: Why this matters for tuning

When Nginx acts as a front door, every visitor uses two connections: one from the visitor to Nginx, and one from Nginx to your app. That one detail explains a surprising number of capacity problems later in this guide.

How many sites run on Nginx? It depends who you ask

Search for Nginx's market share and you'll find numbers that disagree. The sources aren't wrong; they measure different things.

Nginx usage figures from two independent surveys

Source and date

What was counted

Nginx share

W3Techs, 2 September 2026

Websites whose web server could be identified

31.3%

W3Techs, 1 January 2019

Same method, at its yearly peak

40.7%

Netcraft, July 2026

All sites that answered the survey

20.4%

Netcraft, March 2026

The top million busiest sites

20.05%

Netcraft, March 2026

Web-facing computers (physical or virtual machines)

42.36%

So is Nginx on 20% of sites or 42% of servers? Both. Netcraft counts every hostname it can reach, including parked domains, while W3Techs counts only sites where it could identify the server software.

These surveys also have a blind spot. When a site sits behind a service like Cloudflare, the response often names Cloudflare as the server, hiding the Nginx machine behind it. W3Techs reported in February 2026 that among Nginx sites whose reverse proxy service it could identify, 61.7% used Cloudflare. OpenResty, a popular platform built on top of Nginx, is also counted separately (Netcraft had it at 8.02% of sites in July 2026). The fair reading: Nginx's real footprint is probably larger than any headline figure, even as its visible share slips.

Why a few milliseconds are worth the effort

MARKET STATISTICS

8.4%

Rise in retail conversions after a 0.1 second mobile speed improvement

Deloitte study commissioned by Google, 2020

10.1%

Rise in travel conversions from the same 0.1 second improvement

Deloitte study commissioned by Google, 2020

90%+

Mid-size and large firms where one hour of downtime costs over $300,000

ITIC Hourly Cost of Downtime Survey, 2024

The Deloitte study, Milliseconds Make Millions, looked at mobile data from 37 brands. One of its four metrics was server latency, measured as time to first byte (TTFB): the delay before the first piece of a reply arrives. That's the number Nginx tuning affects most.

One honest caveat: the gains appeared only when every page a shopper visited got faster on all four metrics together. Server tuning alone won't hand you an 8.4% lift. It removes one of four brakes.

The ITIC figures are self-reported by more than 1,000 firms, and 41% of enterprises put an hour of downtime between $1 million and over $5 million. Your numbers will differ, but a server that falls over during a spike costs far more than one that's a little slow.

The tuning ladder: seven steps in order of payoff

Cheap, high-impact changes come first. Jumping to clever tricks before fixing the basics is the most common way teams waste a weekend.

Step 1: Measure before you change anything

The default Nginx access log records who visited, what they asked for, and the status code. It doesn't record how long anything took. Without timing, you can't tell whether Nginx is slow, your app is slow, or the visitor's phone has a weak signal.

Add a log format that captures timing:

log_format timed 'remoteaddr [time_local] "$request" $status '

                 'rt=requesttime uct=upstream_connect_time '

                 'uht=upstreamheadertime urt=upstream_response_time '

                 'cache=$upstream_cache_status';

access_log /var/log/nginx/access.log timed;

The request time (rt) is the total time Nginx spent on a request. The upstream connect time (uct) is how long opening a connection to your app took. The upstream header time (uht) is how long your app took to start answering, and the upstream response time (urt) is how long it took to finish. The cache field shows whether the answer came from Nginx's cache.

Also turn on the built-in status page (the stub_status module). It reports active connections plus two counters, "accepts" and "handled." If they drift apart, Nginx is turning connections away because it has hit a limit.

PRO TIP

Keep the status page private. Allow it only from your monitoring server's address or from localhost, since it reveals how busy you are.

Step 2: Size your workers and connections

Three settings decide how many visitors Nginx can hold at once.

worker_processes auto;

worker_rlimit_nofile 65535;

events {

    worker_connections 8192;

}

▪        The worker_processes setting tells Nginx how many workers to start. "Auto" gives one per CPU core, right for nearly everyone.

▪        The worker_connections setting is the ceiling per worker. It counts every connection, including the ones Nginx opens to your app.

▪        The worker_rlimit_nofile setting raises the number of open files each worker may have. On Linux, every network connection counts as an open file, and many systems default to a limit of 1024.

Now apply the two-connections rule. A 4-core server with the shipped setting of 1024 can hold 4 × 1024 = 4,096 connections, or roughly 2,048 visitors at once. That sounds like plenty until you remember that browsers often open several connections each, and slow networks hold them open longer.

Watch the error log for the line "worker_connections are not enough." If it appears, raise the setting and the open-file limit together. Raising one alone just moves the error somewhere less obvious.

Step 3: Reuse connections on both sides

Opening a connection costs time. For HTTPS, both sides first exchange several messages (a TLS handshake) before any page data moves. Reusing an open connection, which Nginx calls keepalive, skips all of that.

On the visitor side, Nginx keeps idle connections open for 75 seconds by default (keepalive_timeout) and allows 1,000 requests per connection (keepalive_requests). On a very busy server, shortening the timeout frees slots that idle visitors are holding.

The app side is where the gains have usually hidden. Until March 2026, Nginx talked to upstream servers over HTTP/1.0 by default, closing the connection after every request. Teams had to fix it by hand, and many never knew.

Nginx 1.29.7, released on 24 March 2026, made HTTP/1.1 with keepalive the upstream default, caching up to 32 idle connections per worker. The change carried into the 1.30 stable series.

upstream app_servers {

    server 10.0.0.11:3000;

    server 10.0.0.12:3000;

    keepalive 64;

}

# Only needed on versions older than 1.29.7:

# proxy_http_version 1.1;

# proxy_set_header Connection "";

The new default has an edge case. Without a keepalive line of your own, Nginx uses "local" mode, so cached connections aren't shared between location blocks. Write a keepalive line without the word "local" and they are shared, as in older versions.

PRO TIP

Run nginx -v before copying any config from a blog post, this one included. Long-term-support Linux releases often ship Nginx versions that are years old, so the hand-written keepalive lines may still be required on your server.

How do you know keepalive isn't working? If the upstream connect time in your log stays above zero, connections are being opened fresh. Thousands of sockets stuck in a state called TIME_WAIT is another clue. Widening Linux's outgoing port range buys time but doesn't fix the cause.

Step 4: Cache what you can, even for one second

The fastest request your app ever handles is the one it never sees. Nginx can store copies of responses and hand them out directly.

Micro-caching works surprisingly well for busy pages that change often. If your homepage gets 1,000 requests per second and Nginx caches it for one second, your app builds the page about once per second instead of 1,000 times. Nobody notices content that's a second old.

proxy_cache_path /var/cache/nginx levels=1:2 keys_zone=pages:50m

                 max_size=2g inactive=10m use_temp_path=off;

location / {

    proxy_cache pages;

    proxy_cache_valid 200 1s;

    proxy_cache_use_stale error timeout updating;

    proxy_cache_lock on;

    proxy_cache_background_update on;

    proxy_cache_bypass $cookie_sessionid;

    proxy_no_cache $cookie_sessionid;

    proxy_pass http://app_servers;

}

Three of those lines handle situations that only show up at scale.

▪        The proxy_cache_lock line prevents a "cache stampede." Without it, when a popular page expires, every visitor arriving in that instant hits your app at once. With it, one request rebuilds the page while the others wait briefly.

▪        The proxy_cache_use_stale line serves the slightly old copy if your app is down, timing out, or refreshing. During an outage, visitors still see a page.

▪        The bypass and no-cache lines skip caching for logged-in visitors, identified here by a session cookie. Change the cookie name to match your app.

The edge case that bites people is personal data. Cache a page showing "Hello, Priya" carelessly and the next visitor sees Priya's name. Nginx won't cache a response that sets a cookie, but pages that only read one get no such protection, so always bypass the cache for requests carrying a login cookie.

Step 5: Get buffering right, especially for streaming

Buffering means Nginx reads your app's whole response into memory quickly, frees the app for its next request, then feeds the visitor at whatever speed their connection allows. It's on by default, and it's one of the biggest reasons a reverse proxy in front of a slow app helps so much. Without it, a visitor on shaky mobile data could tie up your app for ten seconds.

The exception is streaming. Live chat, AI assistants that type answers word by word, and server-sent event feeds need small pieces of data to arrive immediately. With buffering on, those pieces arrive in clumps, and users say the app "freezes, then dumps text." Turn buffering off for those routes only, with proxy_buffering off or by having your app send the header X-Accel-Buffering: no.

Uploads have a related issue. By default, Nginx rejects request bodies over 1 megabyte with a 413 error (client_max_body_size), and bodies larger than a small memory buffer get written to disk. For large uploads, raise the limit on that route only and keep temporary files on fast storage.

Step 6: Compress text and keep secure connections cheap

Compression (gzip) shrinks text files like HTML, CSS, and JavaScript. The default level is 1, the lightest. Levels around 4 to 6 usually shrink files noticeably more, while higher levels burn CPU for tiny extra savings. Skip images and videos, which are already compressed.

For HTTPS, store recent handshake results so returning visitors can skip part of the setup. Nginx's documentation says one megabyte of session cache holds about 4,000 sessions.

gzip on;

gzip_comp_level 5;

gzip_types text/css application/javascript application/json image/svg+xml;

ssl_session_cache shared:SSL:20m;

ssl_session_timeout 1h;

sendfile on;

tcp_nopush on;

 

The sendfile setting lets the operating system copy files straight to the network, which helps static files a lot.

Step 7: Use rate limits to protect performance

Rate limiting caps how many requests one visitor can make per second. It's usually filed under security, yet it's also a performance tool: one badly written script hammering your search page can slow the app for everyone.

limit_req_zone $binary_remote_addr zone=perip:10m rate=10r/s;

location /search {

    limit_req zone=perip burst=20 nodelay;

    proxy_pass http://app_servers;

}

The burst value lets a visitor briefly exceed the limit, since one page load can fire a dozen requests.

Two edge cases deserve attention. Many people can share one public IP address, such as an office, a university, or a mobile carrier's customers, so a limit that's generous for one person can block a whole building. And if Nginx sits behind a CDN or cloud load balancer, every request appears to come from that service unless you configure Nginx's real IP module. Get this wrong and you rate limit your own CDN.

Choosing a load balancing method

With several app servers behind Nginx, the load balancing method decides who gets each request. Open-source Nginx offers several, and since version 1.29.6 (March 2026) it includes "sticky" sessions, which used to require the paid NGINX Plus.

How the methods compare

Method

How it picks a server

Works well for

Watch out for

Round robin (default)

Takes turns, adjusted by any weights you set

Similar servers handling short, even requests

Ignores how busy each server actually is

least_conn

Server with the fewest active connections

Requests that vary a lot in length, like reports or uploads

Each worker counts on its own unless you add a shared zone

ip_hash

Same visitor IP always goes to the same server

Simple session stickiness

Lopsided load when many users share one IP

hash with consistent

Any key you choose, such as the URL

Spreading cached content across cache servers

Uneven traffic if a few keys are very popular

random two least_conn

Picks two servers at random, then the less busy one

Several Nginx machines sharing one pool of app servers

Less familiar to most teams

sticky (1.29.6 and later)

A cookie or route ties a visitor to one server

Apps that keep login sessions in server memory

An overloaded server keeps its visitors

The least_conn note is easy to miss. Each worker keeps its own connection count unless you add a zone line to the upstream block. Without it, four workers can each make a sensible choice while together piling work onto one server.

The same thing happens across machines. Several Nginx servers in front of one app pool each see only their own traffic, so they can all pick the same "least busy" server at the same moment. The random two option adds just enough chance to break that up.

How Nginx decides a server is unhealthy

Load balancing in open-source Nginx relies on passive health checks, meaning it only notices a problem when a real request fails. By default, one failure within 10 seconds (max_fails=1, fail_timeout=10s) marks a server as unavailable for the next 10 seconds.

That real-time decision has consequences. With the defaults, one slow request during a brief app pause can pull a healthy server out of rotation, pushing its traffic onto the others. Set the threshold too high and a broken server keeps getting visitors. A max_fails of around 3 is a common middle ground, but test it against your own failures.

When a request fails with an error or timeout, Nginx tries the next server. Since version 1.9.13 it won't automatically retry requests like POST that can change data, since charging a card twice is worse than an error. You can cap retries with proxy_next_upstream_tries.

A few edge cases from the official documentation:

▪        If an upstream group has only one server, the failure settings are ignored and that server is never marked unavailable.

▪        Host names in an upstream block are normally looked up once, at startup. If your servers' IP addresses change, Nginx keeps using the old ones. The resolve option, in open-source Nginx since version 1.27.3, fixes this when paired with a resolver line.

▪        Active health checks, which send test requests on a schedule, remain a paid NGINX Plus feature, as does the least_time method.

When signals contradict each other

The hardest part of Nginx performance tuning is reading the evidence correctly when different numbers point in different directions. Here are the situations that confuse teams most often.

Reading mixed signals from Nginx

What you see

What it often means

Where to look

CPU is low, but pages are slow

Nginx is waiting on your app or the disk, not working hard

Compare rt with urt in your timing log

rt is high, urt is low

Your app was fast; the visitor's network was slow

Usually normal for mobile users; buffering is doing its job

502 Bad Gateway

Your app refused the connection, crashed, or closed it early

Error log lines mentioning "connect() failed" or "prematurely closed"

504 Gateway Timeout

Your app accepted the request but took too long

proxy_read_timeout (default 60 seconds) and uht values

Many 499 status codes

Visitors or another proxy gave up and closed the connection

request_time on those lines, plus timeouts on any proxy in front

"accepts" and "handled" differ

Nginx is refusing connections at its limit

worker_connections and open-file limits

The 499 code isn't a standard HTTP code. Nginx uses it to log clients that hung up before getting an answer. A rise in 499s usually means the problem sits behind Nginx: the app got slow, visitors gave up, and Nginx recorded it. Tuning Nginx to "fix" 499s changes nothing.

Averages are another trap. A 200 millisecond average can hide 2% of visitors waiting five seconds, so check the 95th and 99th percentile values too.

What actually happens during a traffic spike

Suppose a product launch sends ten times your normal traffic within a few minutes. Here's how the overload usually unfolds.

Stage 1.          Your app servers fill up first, and new requests wait in line.

Stage 2.          Every waiting visitor ties up two Nginx connections, so the connection count climbs much faster than real throughput.

Stage 3.          Requests time out, and Nginx retries them on other servers that are also overloaded. Engineers call this a retry storm.

Stage 4.          Servers get marked as failed, pushing traffic onto fewer machines, which fail faster.

Stage 5.          Nginx hits its connection ceiling, and new visitors can't get in the door.

The tuning ladder breaks this chain at several points. Stale cache serves visitors without touching the app, rate limits trim abusive traffic, capped retries stop the storm, and higher connection limits buy breathing room. An error_page with a friendly "we're busy" message beats a raw error.

One more limit sits below Nginx. Connections waiting to be accepted sit in a Linux queue capped by a kernel setting called somaxconn. Nginx asks for 511 by default, but kernels before version 5.4 capped it at 128, silently shrinking the queue on older systems.

Field note: Reloads under load

A reload starts fresh workers and lets old ones finish their requests. Old workers holding WebSockets can live for hours, so frequent reloads on a busy day can leave several generations of workers eating memory. The worker_shutdown_timeout setting caps how long they linger.

Where Nginx reaches its limits

Nginx doesn't scale forever. The most detailed public account of its limits comes from Cloudflare, which ran Nginx at the core of its network for years before announcing a replacement called Pingora in 2022.

Cloudflare's engineers described two design problems. Each request is tied to one worker, so a CPU-heavy or disk-blocked request slows everything else on that worker. And because each worker keeps its own pool of reusable connections, adding workers lowers the chance of finding one ready to reuse.

Their numbers are striking. Cloudflare reported that Pingora used about 70% less CPU and 67% less memory than its old Nginx-based service for the same traffic. For one major customer, connection reuse went from 87.1% to 99.92%, which cut new connections to that customer's servers by 160 times.

Keep that in proportion. Cloudflare works at a scale almost no other company sees. For a startup or a mid-size business, a sensibly configured Nginx web server is very rarely the slowest part of the stack. Your database or app code will usually give out first.

A note for teams running Kubernetes

If your app runs on Kubernetes, check your "ingress controller," the component that routes outside traffic into your cluster. The community project Ingress NGINX was retired in March 2026. In January 2026, the Kubernetes Steering and Security Response Committees called it critical infrastructure for about half of cloud native environments and warned that staying on it leaves users open to attack.

Existing installs keep working without security fixes. The retirement covers that one project, not Nginx itself, and F5 maintains a separate NGINX-based controller. Kubernetes points users toward the newer Gateway API or another maintained controller. Re-check your tuning after switching, since each controller exposes settings differently.

Key takeaways

»   Add timing to your logs before touching any other setting. Without it, every change is a guess.

»   Remember the two-connections rule when you size worker_connections, and raise the open-file limit at the same time.

»   Check your Nginx version. Upstream keepalive became the default only in March 2026.

»   Cache for even one second on busy pages, and always skip the cache for logged-in visitors.

»   Turn buffering off for streaming routes and leave it on everywhere else.

»   Read percentiles, not averages, and treat a rise in 499s as a sign the problem is probably behind Nginx.

Wrapping up

Good Nginx performance tuning mostly comes down to understanding what the server is doing. Measure first, size connections with the real math, reuse connections, cache carefully, and decide in advance how your setup should behave when things fail.

Most changes here take an afternoon. The harder part is the habit afterward: reading your logs after launches and incidents, and adjusting from evidence. If your team is building something that has to hold up under heavy traffic and nobody on it has tuned a production front door before, pairing with an experienced software development partner for the first round can save a painful outage. Either way, the Nginx web server rewards teams who take time to understand it.

Nainesh Pandya

Nainesh Pandya

Nainesh is the marketing expert helping our clients and customers achieve success in terms of outreach and visibility. From understanding the complexities of value-chain and the impact of future technologies, Nainesh’s incredible understanding of digital marketing and online outreach helps create high-impact strategies.

Build Your Agile Team

We provide you with a top-performing extended team for all your development needs in any technology.

Hourly
$20
It Includes
Duration
Hourly Basis
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
25 Hours (MIN)
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Monthly
$2600
It Includes
Duration
160 Hours
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Team
$13200
It Includes
Team Members
1 (PM), 1 (QA), 4 (Developers)
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile

Frequently Asked Questions

Is Nginx faster than Apache?
For static files and many simultaneous connections, Nginx usually uses less memory thanks to its event loop. Apache's newer "event" mode narrowed the gap, and for dynamic apps the difference is often small because the app does most of the work.
How many visitors can one Nginx server handle?
It depends on hardware, response sizes, encryption, and your app's speed. A quick ceiling is CPU cores times worker_connections, divided by two when Nginx works as a reverse proxy. Confirm the real figure with a load test that mirrors your actual traffic.
Do I still need Nginx if my site uses a CDN like Cloudflare?
Usually, yes. A CDN caches content near visitors, but uncacheable requests still reach your servers. Nginx handles those with routing, load balancing across app servers, upstream keepalive, and protection from slow connections.
What's the difference between open-source Nginx and NGINX Plus?
NGINX Plus is F5's paid version. Its main performance extras are active health checks, the least_time method, and a live dashboard with detailed upstream statistics. Sticky sessions and DNS-based server lookups, once paid-only, are now open source.
How often should I review my Nginx settings?
After big traffic changes, after every upgrade, and after any incident. Upgrades matter more than people expect, as the March 2026 keepalive change showed. A short quarterly look at your timing and error logs is a sensible habit even when things seem fine.