Loading...

Basics · 11 min read

What is nginx? The web server in front of half the internet, explained

nginx (pronounced "engine-x") is open-source software that accepts HTTP connections and decides what to do with each request: send a file from disk, pass it to an application, spread it across several servers, or answer from its own cache. A few worker processes, each running an event loop, hold thousands of connections at once, which is why nginx sits in front of so many sites, applications and CDNs.

Updated

What is nginx? The web server in front of half the internet, explained

What nginx is, and the five jobs it does

nginx (pronounced "engine-x") is an open-source HTTP server and proxy. Igor Sysoev wrote it to keep ten thousand simultaneous connections open on one machine, at a time when that brought most servers down, and released it in 2004. Today it is one of the two most widely deployed web servers, alongside Apache, and the engine inside many load balancers, Kubernetes ingress controllers and CDNs. The open-source version lives at nginx.org; F5, which bought the company behind it in 2019, sells a commercial edition called NGINX Plus.

The name covers five roles, and one configuration can combine all of them:

Web server. It reads files from disk and sends them: HTML, CSS, JavaScript, images, downloads. This is what it does fastest, with sendfile handing the copying to the kernel.

Reverse proxy. It accepts the request and passes it to an application that should not face the internet itself: PHP-FPM, Node.js, Python, Go, Java, a container. The full topic, including the headers that must be forwarded, is in our reverse proxy guide.

Load balancer. It spreads requests over several copies of that application and stops sending traffic to one that fails.

Cache. It stores upstream responses on disk and answers repeat requests without asking the application again.

TLS terminator. It holds the certificates, speaks HTTPS, HTTP/2 and HTTP/3 to the browser, and talks plain HTTP to the application on a private address.

What it is not: an application server. nginx does not execute your PHP, Python or Ruby code. It hands those requests to a process that does, and sends the answer back.

Event-driven: why nginx holds so many connections

When nginx starts, a master process reads the configuration, opens the listening ports and starts a few worker processes, usually one per CPU core (worker_processes auto). The master never serves traffic. It manages the workers, and it is what makes a reload graceful.

Each worker is single-threaded and runs an event loop. It asks the kernel, through epoll on Linux or kqueue on BSD and macOS, which of its connections have something ready: a new request, a client that can take more bytes, an upstream that has answered. It does a short piece of work on each ready connection and moves on; nothing sits waiting. A keep-alive connection that is idle between requests, or a slow mobile client trickling in its request, costs the worker a few kilobytes of memory and no CPU.

Apache's classic model is the opposite. The prefork MPM gives every connection its own process and the worker MPM its own thread, and that process or thread is tied up for as long as the connection lasts, busy or idle. Ten thousand open keep-alive connections mean ten thousand processes or threads, with the memory and context switching that implies. Apache's newer event MPM parks idle keep-alive connections with a listener thread, which narrows the gap, but each active request still holds a thread.

The rule that follows: a worker must never block. The work that genuinely takes time, like running PHP or querying a database, happens in another process, and nginx treats the answer as one more event. The capacity sum is worker_processes × worker_connections, and when nginx proxies, each client uses two of those slots: one to the browser, one to the upstream.

The top of nginx.conf: one master, a worker per core, an event loop in each

user www-data;
worker_processes auto;            # one worker per CPU core
error_log /var/log/nginx/error.log warn;

events {
    worker_connections 1024;      # per worker: clients + upstream connections
}

http {
    include      mime.types;
    sendfile     on;
    keepalive_timeout 65s;

    include /etc/nginx/conf.d/*.conf;   # one file per site
}

What nginx is used for

In practice nginx sits in one of these positions, often several at once.

Serving a static site or a front-end build. A React, Vue or Astro build, documentation, a landing page: files on disk and nginx in front, nothing else.

In front of PHP. WordPress, Laravel and most PHP applications run as nginx plus PHP-FPM. nginx serves images, CSS and JavaScript itself and passes .php requests to FPM over a Unix socket with fastcgi_pass.

In front of an application server. Node.js, Python (Gunicorn, Uvicorn), Ruby (Puma) and Go services listen on a local port. nginx terminates HTTPS, absorbs slow clients by buffering the request, then forwards it with proxy_pass, so the application only ever sees fast, complete requests.

Load balancing across several application instances, covered briefly below.

TLS and modern protocols for an application that has none: certificates (most people get them with certbot and Let's Encrypt), HTTP/2 with http2 on; and HTTP/3 with listen 443 quic; on nginx 1.25 and later. What the protocols change is in HTTP/2 vs HTTP/3.

Rules at the door: redirects, rate limiting with limit_req, IP allow and deny lists, gzip compression and response headers. Which headers to send, and why, is in HTTP security headers.

A minimal server block, line by line

A server block is one site. nginx chooses it by the port the request arrived on and by the Host header, compared against server_name. Inside it, location blocks match URL paths. The block below serves a static site and passes /api/ to an application on port 3000.

listen is the port, once for IPv4 and once for IPv6.

server_name lists the hostnames this block answers. A request whose Host matches no block goes to the port's default server: the block marked default_server, or else the first one nginx read. That is why an unknown hostname pointed at your IP shows some other site. A catch-all block that answers return 444; (close the connection) stops it.

root is where files are looked up: /about.html becomes /var/www/example/about.html.

try_files tries the file, then a directory, then the fallback. For a single-page app, replace =404 with /index.html so client-side routes load the app.

location matching has an order that surprises people. An exact match (=) wins outright. Otherwise nginx remembers the longest matching prefix, then tries the regular expressions (~, ~*) in file order and takes the first that matches; only if none matches is the remembered prefix used. A ^~ prefix skips the regex step. In this example /api/logo.png is served by the image regex, not by /api/, which is the usual answer to "why is my location ignored".

expires sets Cache-Control: max-age and Expires on fingerprinted assets. Choose the values with the Cache-Control guide.

This block is plain HTTP. Add the certificate with certbot, which edits the block for you, and redirect port 80 to 443.

/etc/nginx/conf.d/example.conf: a static site with an API behind it

server {
    listen 80;
    listen [::]:80;
    server_name example.com www.example.com;

    root  /var/www/example;
    index index.html;

    access_log /var/log/nginx/example.access.log;
    error_log  /var/log/nginx/example.error.log;

    location / {
        try_files $uri $uri/ =404;
    }

    location ~* \.(?:css|js|woff2|png|jpg|webp|avif|svg)$ {
        expires 30d;
        access_log off;
    }

    location /api/ {
        proxy_pass http://127.0.0.1:3000;
        # plus the forwarded headers from the reverse proxy guide
    }
}

Install, test, reload: the commands you use every day

Distribution packages put the main file at /etc/nginx/nginx.conf. Debian and Ubuntu keep sites in /etc/nginx/sites-available/ and enable them with a symlink into sites-enabled/; RHEL, Rocky, Alpine and the nginx.org packages read /etc/nginx/conf.d/*.conf. Logs go to /var/log/nginx/.

Two habits prevent most self-inflicted outages. Run nginx -t before every reload: it parses the whole configuration and names the file and line of any mistake. And reload rather than restart. On a reload the master starts new workers with the new configuration and lets the old ones finish their open requests, so no connection is dropped; if the new configuration does not load, the master keeps the old workers running and logs why.

nginx -T prints the complete configuration as nginx sees it, every include expanded. When a setting seems to have no effect, it is usually overridden in a file you forgot, and -T shows you which one.

Package install, safe reload and the effective config
# Debian / Ubuntu
sudo apt install nginx
nginx -v                                   # version; -V adds build options and modules

# check the syntax, then apply without dropping connections
sudo nginx -t && sudo systemctl reload nginx

# the whole effective configuration, includes expanded
sudo nginx -T | less

# which names and ports are configured
sudo nginx -T | grep -E '^\s*(server_name|listen)'

# watch errors while you test
sudo tail -f /var/log/nginx/error.log

nginx vs Apache

Both are mature, free, and fast enough for nearly any site. The differences that actually decide between them:

Connection model. nginx's event loop holds idle and slow connections almost for free; Apache ties a process or thread to each active request and, outside the event MPM, to each idle connection as well. With many concurrent keep-alive connections nginx uses far less memory.

Configuration. With AllowOverride on, Apache reads .htaccess files from every directory on the path of each request, so a user can change rules without touching the server configuration. nginx has no equivalent: every rule lives in the central configuration and is parsed once, at load. That is faster and easier to audit, but a WordPress or Laravel .htaccess has to be rewritten as location and try_files rules when you move.

PHP. Apache can run PHP inside its own processes with mod_php; nginx always talks to a separate PHP-FPM pool. Current Apache setups often use PHP-FPM too, so this is now mostly a difference in defaults.

Modules. Apache loads modules at runtime from a very large catalogue. nginx supports dynamic modules, but each must be built against the exact nginx version you run, so many extras mean compiling.

Static files and proxying are where nginx is strongest. That is why a common hybrid puts nginx in front, serving files and terminating TLS, with Apache behind it running an application that depends on .htaccess.

The honest summary: starting fresh, nginx is the default choice for most teams. Replacing a working Apache will not make a slow site fast, though. The slow part is almost always the application or the distance to the visitor, not the web server.

Reverse proxy, load balancer and cache: the short version

Proxying is one directive, proxy_pass, plus the headers that tell the application who the client really is. Those headers, WebSocket support and proxy_cache are covered in the reverse proxy guide, so this page does not repeat them.

Load balancing adds an upstream block that names the servers. The default is round robin. least_conn sends each request to the server with the fewest active connections, which suits requests of uneven length, and ip_hash or hash keeps a client on the same server. Health checking in open-source nginx is passive: after max_fails failures within fail_timeout, a server is skipped for that long, and proxy_next_upstream retries the request on another one. Active checks that probe a URL on a schedule are an NGINX Plus feature. For the wider picture, see what a load balancer does.

Caching works well, with one gap to know up front: open-source nginx has no command to purge a cached URL. You wait for the entry to expire or delete files from the cache directory on every server. That gap is one of the most common reasons teams move the public cache to a CDN.

Three application servers behind one nginx

upstream app {
    least_conn;
    server 10.0.0.11:3000 max_fails=3 fail_timeout=10s;
    server 10.0.0.12:3000 max_fails=3 fail_timeout=10s;
    server 10.0.0.13:3000 backup;     # used only when the others are down
    keepalive 32;                     # reuse connections to the app
}

server {
    listen 80;
    server_name app.example.com;

    location / {
        proxy_pass http://app;
        proxy_http_version 1.1;
        proxy_set_header Connection "";          # required for upstream keepalive
        proxy_next_upstream error timeout http_502;
        # plus the forwarded headers from the reverse proxy guide
    }
}

502, 504 and 413: what nginx is telling you

The status code the visitor sees is the summary; the error log is the explanation. Each of these errors leaves a distinctive line there.

502 Bad Gateway means nginx got no usable answer from the upstream. connect() failed (111: Connection refused) says nothing is listening at the address in proxy_pass: the application is down or on another port. connect() to unix:/run/php/php8.3-fpm.sock failed (2: No such file or directory) is a PHP-FPM socket path that does not match the PHP version installed, and (13: Permission denied) on the same line is a socket nginx's user may not open. upstream prematurely closed connection means the application crashed or was killed mid-request; read its own log and the kernel's out-of-memory messages. upstream sent too big header means the response headers, often a pile of cookies, do not fit the buffer: raise proxy_buffer_size, or fastcgi_buffer_size for PHP. The full diagnosis, including the case with a CDN in front, is in 502 Bad Gateway.

504 Gateway Timeout means the upstream accepted the request and did not answer within proxy_read_timeout (or fastcgi_read_timeout), 60 seconds by default. The log says upstream timed out (110: Connection timed out) while reading response header from upstream. A longer timeout is right for a known slow export or report; for ordinary pages it only hides a slow query or a full worker pool. Every proxy in the chain has its own limit, so a load balancer or CDN in front may give up before nginx does.

413 Request Entity Too Large means the request body is larger than client_max_body_size, which is 1 MB by default: the classic failed image or backup upload. The log says client intended to send too large body. Raise the limit in the server or location that receives uploads rather than globally, and for PHP raise upload_max_filesize and post_max_size to match, or PHP rejects what nginx let through.

A 503 from nginx is usually its own limit_req or limit_conn turning a request away; that and the other causes are in 503 Service Unavailable. And directory index of "/var/www/..." is forbidden is a 403 for a directory without an index file, covered in 403 Forbidden.

Error log lines and the directive each one points to

# /var/log/nginx/error.log (abridged)
connect() failed (111: Connection refused) while connecting to upstream        -> 502
upstream prematurely closed connection while reading response header          -> 502
upstream sent too big header while reading response header from upstream       -> 502
upstream timed out (110: Connection timed out) while reading response header   -> 504
client intended to send too large body: 52428800 bytes                         -> 413

# the matching fixes, in the server or location that needs them
client_max_body_size 64m;     # PHP: raise upload_max_filesize and post_max_size too
proxy_read_timeout   300s;    # only for a known slow endpoint
proxy_buffer_size    16k;     # large response headers and cookies
proxy_buffers        8 16k;

When to put a CDN in front of nginx

One nginx serves a lot of traffic. What it cannot change is where it is. A visitor in another country waits on every round trip to your one datacentre; a traffic spike made mostly of repeat requests still lands on your one machine; an attack aimed at your IP reaches the server that runs your site. Those are the moments to add a CDN: your audience is far from the server, cacheable traffic is a large share of the load, you need cache purge across more than one machine, or the origin should stop being directly reachable.

You keep nginx; its job changes. The CDN becomes the public front door and nginx becomes the origin, serving files and routing to the application while the edge handles distance, caching and filtering. Three things to adjust on the nginx side. Send correct Cache-Control headers, because the edge follows them. Restore the visitor's IP with the realip module (set_real_ip_from for the CDN's addresses, real_ip_header for the header it sends), or every log line and every limit_req zone sees the edge instead of the visitor. And allow only the CDN to reach the origin, so attacks cannot go around it.

The CDN.com.tr edge runs on nginx itself: its configuration applies your package's file size as client_max_body_size, so an upload that passes through the CDN needs that limit to be large enough as well as your own. In front of your origin, the edge network, with servers in Turkey and abroad, caches by per-path delivery rules and purges instantly by exact path or folder from the panel, the cdnctl command line or the REST API. A WAF, DDoS protection, bot protection with a JavaScript challenge, country and IP rules and rate limiting run before a request reaches your server, and Let's Encrypt certificates are issued and renewed for every connected domain.

Two details make the switch easier to verify. Pages already in the edge cache keep serving while your origin is down, and the X-Proxy-Cache-MT response header shows HIT or MISS, so you can tell whether a request reached your nginx at all. The edge also sends SNI to the origin, so an nginx with several HTTPS server blocks presents the right certificate.

nginx FAQ

How do you pronounce nginx?

"Engine-x". Both nginx and NGINX are used; the lowercase form is the name of the open-source project and its binary.

Is nginx free?

The open-source nginx from nginx.org is free under a two-clause BSD licence, for commercial use too. NGINX Plus is a paid subscription from F5 that adds active health checks, an API for purging the cache and changing upstreams at runtime, a live status dashboard and support.

Is nginx a web server or a reverse proxy?

Both, often in the same configuration. A location with root serves files from disk, one with proxy_pass or fastcgi_pass hands the request to an application. Most sites do both: static assets directly, everything dynamic through the proxy.

Is nginx better than Apache?

For static files, proxying and large numbers of concurrent connections it uses less memory, and it is the usual choice for new setups. Apache is the better fit when you depend on .htaccess or an Apache-only module. For a typical site, neither web server is the bottleneck; the application and the distance to visitors are.

Why does my domain show a different site on nginx?

No server_name matched the request's Host header, so nginx used the default server for that port: the block marked default_server, or the first one it loaded. Check the names with nginx -T | grep server_name, and add a catch-all block that returns 444 so unknown hostnames get nothing.

Do I still need nginx if I use a CDN?

Usually yes, as the origin: something still has to serve files and route requests to your application, and the CDN pulls from it. On CDN.com.tr container apps you can skip it for publishing, because the edge reaches each app through an internal platform route with no reverse-proxy container in between.