What a load balancer does
A load balancer sits between clients and a group of servers that can all answer the same request, and decides which server gets each one. Clients connect to one address; the load balancer picks a healthy backend, forwards the request, and passes the answer back.
It does three jobs. It spreads load, so three servers can carry roughly three times the traffic of one. It removes failed servers: health checks notice a backend that stopped answering and send nobody to it until it recovers. And it makes changes invisible: you can take a server out, patch it, deploy to it and put it back without visitors seeing an error.
The second job often matters most: two servers behind a load balancer are less about capacity than about one of them being allowed to die. A layer 7 load balancer is a kind of reverse proxy whose job is choosing between identical backends; HAProxy and nginx do both.
Layer 4 vs layer 7 load balancing
The layer is how much of the traffic the load balancer reads before it decides.
A layer 4 load balancer works with TCP connections and UDP packets. It sees source and destination addresses and ports, picks a backend when the connection opens, and copies bytes in both directions. It cannot see a URL, a header or a cookie, so every request on that connection goes to the same server. In exchange it is fast, protocol-agnostic (databases, MQTT, SMTP, game servers) and can pass TLS straight through, so the certificate stays on the backends. HAProxy's mode tcp and nginx's stream {} block work here.
A layer 7 load balancer speaks HTTP. It terminates the connection, reads each request, and can route by host, path, header or cookie: /api to one pool, / to another, a canary cookie to the new version. It can retry a failed request on another server and spread the requests of one HTTP/2 connection over several backends (HAProxy mode http, nginx http {}).
To read HTTP, a layer 7 balancer must decrypt it, so TLS termination comes with the territory. The certificate lives on the load balancer and backends get plain HTTP on a private network, or a second TLS connection if the network is not trusted. The cost is that the backend no longer sees the client: it sees the load balancer. Layer 7 balancers fix that with X-Forwarded-For and X-Forwarded-Proto; layer 4 balancers use the PROXY protocol, which the backend must be configured to accept.
Balancing algorithms: round robin, least connections, hashing, weights
Round robin hands requests to each server in turn. It is the default nearly everywhere and is right when requests cost about the same and servers are identical.
Weighted round robin gives a bigger server a bigger share: with weights 2, 2 and 1, the third server gets a fifth of the traffic. Weights are also how you run a canary: put the new version in the pool with a small weight, watch its error rate, then raise it.
Least connections sends the next request to the server with the fewest active connections. Use it when request cost varies a lot (reports next to page views, uploads next to API calls) or connections are long-lived, as with WebSockets.
IP hash (HAProxy balance source, nginx ip_hash) maps each client address to the same server every time. It gives you affinity without cookies, but distribution follows the client population: one office or one mobile carrier behind a shared address can land thousands of users on one backend.
Consistent hashing on a key (the URI, a user ID) keeps the same key on the same server, and when a server is added or removed only a small share of keys move. That is what you want in front of caches: hash by URI and each object lives in one cache instead of all of them. nginx writes it as hash $request_uri consistent;, HAProxy as balance uri with hash-type consistent.
Whatever you choose, the algorithm matters less than health checks that tell the truth.
Health checks and failover
A load balancer is only as good as its idea of which servers are alive.
Active health checks probe each backend on a timer: open a TCP connection, or request a URL such as /healthz and expect a 200. After fall failures in a row the server is marked down and receives nothing; after rise successes it comes back. HAProxy does this in the open-source version. Open-source nginx does not: max_fails and fail_timeout are passive checks, counting failed real requests and resting the server for a while. Active checks in nginx are part of the commercial NGINX Plus. Timing is a trade-off: a check every two seconds with fall 3 removes a dead server in about six seconds; every 30 seconds means a minute and a half of errors.
Failover is what happens next. Traffic shifts to the remaining servers, so they must have room for it: three servers each running at 80% cannot absorb the loss of one. A backup server receives traffic only when every primary is down, which suits a cold spare or a maintenance page.
Make the health endpoint answer the question the load balancer asks: *should this server get traffic right now?* That means checking what is local to the process (it started, it can reach its own dependencies, it is not shutting down) and returning quickly. During a deploy, fail the check first, let in-flight requests finish (connection draining), then stop the process.
Session persistence (sticky sessions)
If an application keeps a logged-in user's session in the memory of one server, the next request must reach that server or the user is logged out. Session persistence makes the load balancer remember who goes where.
The usual ways are an inserted cookie (the load balancer adds a cookie naming the server, as HAProxy does with cookie SRV insert indirect nocache plus a cookie value on each server line), an application cookie it learns from (PHPSESSID, JSESSIONID), or source IP hashing. Open-source nginx has only the hashing methods (ip_hash or hash $cookie_name); the sticky directive belongs to NGINX Plus.
Stickiness has a price. Load stops being even, because a few heavy users stay pinned to one server. Draining a server takes as long as its longest session, and when a server dies its users lose their sessions anyway.
The durable fix is to make the servers interchangeable: keep sessions in a shared store such as Redis or the database, or in a signed cookie, and let any server answer any request. Then a failed server costs nobody their login.
A working HAProxy configuration
An HAProxy configuration reads top to bottom: global for the process, defaults inherited by every section, a frontend that accepts connections and a backend that holds the pool.
The example terminates TLS on 443 (the .pem file holds the certificate chain and the private key together), redirects plain HTTP, adds the forwarding headers, and balances with least connections. The health check is a real HTTP request with a Host header, because many applications answer 404 or a redirect to a request without one, and a check that expects 200 then fails on a healthy server. http-check send needs HAProxy 2.2 or newer; older versions put the request on the option httpchk line.
default-server sets the check timing once: a probe every two seconds, three failures to mark a server down, two successes to bring it back. app3 is a smaller machine and takes half the share of the others; spare gets traffic only when all three are down. To make sessions sticky, add cookie SRV insert indirect nocache to the backend and cookie app1 (and so on) to each server line.
Check the file with haproxy -c -f /etc/haproxy/haproxy.cfg before every reload.
/etc/haproxy/haproxy.cfg: TLS termination, least connections, HTTP health checks
global
log /dev/log local0
maxconn 20000
defaults
mode http
log global
option httplog
timeout connect 5s
timeout client 60s
timeout server 60s
frontend fe_web
bind :80
bind :443 ssl crt /etc/haproxy/certs/example.com.pem alpn h2,http/1.1
http-request redirect scheme https unless { ssl_fc }
option forwardfor
http-request set-header X-Forwarded-Proto https
default_backend be_app
backend be_app
balance leastconn
option httpchk
http-check send meth GET uri /healthz ver HTTP/1.1 hdr Host app.example.com
http-check expect status 200
default-server inter 2s fall 3 rise 2
server app1 10.0.0.11:8080 check weight 100
server app2 10.0.0.12:8080 check weight 100
server app3 10.0.0.13:8080 check weight 50
server spare 10.0.0.20:8080 check backup
Load balancing with nginx upstream
In nginx, the pool is an upstream block and proxy_pass points at it by name. With no algorithm line, nginx uses weighted round robin; least_conn, ip_hash, hash … consistent and random change that.
A few details in the example are easy to miss. keepalive 32 keeps idle connections to the backends open for reuse, but only works together with proxy_http_version 1.1 and an empty Connection header; without those two lines nginx opens a fresh connection for every request. max_fails=3 fail_timeout=10s takes a server out for ten seconds after three failed requests within ten seconds: that is nginx's passive health check. backup only works with the round robin, weighted and least_conn methods, not with the hashing ones.
proxy_next_upstream decides when a failed request is retried on the next server. By default nginx retries on connection errors and timeouts, and since 1.9.13 it never retries a POST, LOCK or PATCH unless you add non_idempotent. Keep it that way: retrying a payment because the first server timed out after charging the card is worse than an error. proxy_next_upstream_tries 2 stops one slow request from walking the whole pool. The headers are the ones any reverse proxy needs; the nginx guide covers the rest.
nginx: an upstream pool with weights, passive checks, a backup and keep-alive
upstream app {
least_conn;
server 10.0.0.11:8080 weight=2 max_fails=3 fail_timeout=10s;
server 10.0.0.12:8080 weight=2 max_fails=3 fail_timeout=10s;
server 10.0.0.13:8080 max_fails=3 fail_timeout=10s;
server 10.0.0.20:8080 backup;
keepalive 32;
}
server {
listen 443 ssl;
server_name example.com;
ssl_certificate /etc/ssl/example.com/fullchain.pem;
ssl_certificate_key /etc/ssl/example.com/privkey.pem;
location / {
proxy_pass http://app;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_next_upstream error timeout http_502 http_503;
proxy_next_upstream_tries 2;
proxy_connect_timeout 3s;
}
}
Global vs local: DNS, GeoDNS, anycast and cloud load balancers
Everything above is local load balancing: one site, one pool, a proxy in the request path. Global load balancing decides which site, region or data centre a visitor reaches in the first place, and it is usually done before any proxy sees the request.
DNS round robin is the oldest form: publish several A records and resolvers hand them out in rotating order. It costs nothing and spreads traffic roughly, but DNS knows nothing about health. A dead server's address keeps being handed out until someone removes it, and resolvers and browsers cache the old answer for the record's TTL and sometimes longer.
GeoDNS answers differently depending on where the query comes from, so visitors in one country get the nearby site's address. Combined with health checks on the DNS side, it becomes DNS failover: a region that fails its checks is withdrawn from answers. The limits are the same TTL caching and the fact that the location is that of the resolver, not the visitor, unless the resolver passes on part of the client's address (EDNS Client Subnet). The DNS guide covers TTLs and resolvers in detail.
Anycast announces the same IP address from many locations over BGP, and the internet's routing delivers each packet to the nearest one. There is no TTL to wait for: when a location stops announcing, routes move within seconds to minutes. It is how large DNS services and CDNs reach their nodes. The layers stack: DNS or anycast picks the location, a proxy inside it picks the server.
Cloud load balancers package the same ideas as a managed service. AWS has the Application Load Balancer (layer 7) and the Network Load Balancer (layer 4); Google Cloud Load Balancing offers global and regional, application and network variants; Azure separates Azure Load Balancer (layer 4) from Application Gateway (layer 7). In Kubernetes, a Service spreads connections over pods and an Ingress controller does layer 7 routing. The algorithms, health checks and pitfalls in this guide apply to them unchanged.
Common mistakes that cause outages
The load balancer is the single point of failure. Three application servers behind one HAProxy box means the box is now the thing that takes you down. Run two and move a floating IP between them with VRRP (keepalived is the usual tool), or put a managed or DNS-level layer in front. Then test it by switching off the active one.
The health check lies, in either direction. A /healthz that always returns 200 keeps a server whose database connection pool is exhausted in rotation, so a share of requests fail while the dashboard shows everything green. The opposite is just as bad: a check that queries the shared database marks *every* server down during a five-second database hiccup, and a short slowdown becomes a full outage. Check the server, not the systems all servers share.
Mismatched timeouts. If the backend closes idle keep-alive connections sooner than the load balancer expects, the balancer occasionally sends a request down a connection the backend has just closed, and the client gets a sporadic 502. Keep the backend's keep-alive timeout longer than the load balancer's idle timeout. More causes are in the 502 Bad Gateway guide.
The client's address disappears. Without X-Forwarded-For (layer 7) or the PROXY protocol (layer 4), the application logs, rate-limits and geolocates the load balancer's IP.
Servers that are not identical. Different builds, different config, or different file timestamps that change the default ETag, so a browser's revalidation returns a full 200 instead of a 304 whenever the other server answers. Deploy one artefact everywhere.
Testing is simple. Expose the backend's name in a response header during testing and loop over requests; watch the distribution, then stop one backend and watch the requests move.
# which backend answered? (expose the server name in a debug header first)
for i in $(seq 1 10); do
curl -s -o /dev/null -D - https://example.com/ | grep -i '^x-served-by'
done
# HAProxy: live state of every server through the runtime socket
# (needs "stats socket /run/haproxy/admin.sock mode 660 level admin" in global)
echo "show servers state be_app" | socat stdio /run/haproxy/admin.sock
# take one server out gracefully before a deploy, then put it back
echo "set server be_app/app1 state drain" | socat stdio /run/haproxy/admin.sock
echo "set server be_app/app1 state ready" | socat stdio /run/haproxy/admin.sock
Where a CDN fits, and what CDN.com.tr does
A CDN is global load balancing you do not have to build. Visitors are routed to an edge server near them, the edges answer from cache whatever they can, and only misses travel to your origin. CDN.com.tr runs edge servers in Turkey and abroad and spreads visitors across them, so spreading visitors across locations is already done before a request reaches anything you operate.
What stays yours is the origin. For a Pull CDN site, the edge fetches from the Source Address you set in the panel, an IP or a domain. If you run several origin servers, the load balancer that chooses between them sits behind that address, and everything in this guide applies to it. Lock it down so it accepts connections only from the edge: the panel's own guidance is to keep the origin's address out of public DNS so the only path runs through the edge.
The edge also softens origin failures. Content already in the edge cache keeps being served during its TTL while the origin is down, and the Traffic quality report marks a cached copy served while the origin was slow or failing as STALE. The same report shows origin health: requests to origin, retry share, connection errors (502/504) and the origin's average time to first byte. Uncached and expired objects still need a working origin, which is why the origin still needs its own redundancy for dynamic pages.
If you would rather not run that layer at all, container apps on CDN.com.tr take a replica count and a health check (HTTP path, TCP or none) per app. Traffic enters through the edge; with several replicas, if one container is restarting or unhealthy the others keep serving, and rollouts are health-gated, so deploying a new version does not open a window where the app is unreachable. Sessions go into managed Redis, so a user stays logged in whichever replica serves the next request: the stateless design recommended above, without the sticky cookie.
Load balancer FAQ
HAProxy or nginx for load balancing?
Both are fast and reliable. HAProxy has active health checks, cookie stickiness, a runtime API for draining servers and very detailed statistics in the free version, which makes it the more complete pure load balancer. nginx is the better fit when the same box also serves files, caches or runs the routing for several sites, but its free version has only passive health checks.
What is the difference between layer 4 and layer 7 load balancing?
Layer 4 picks a server per TCP connection or UDP flow and never reads the content, so it works for any protocol and can pass TLS through untouched. Layer 7 reads each HTTP request, so it can route by URL, header or cookie, retry on another server and add forwarding headers, at the cost of terminating TLS.
Which load balancing algorithm should I use?
Start with round robin for identical servers and similar requests. Switch to least connections when request durations vary widely or connections stay open, as with WebSockets. Use consistent hashing when the same key should reach the same server, such as URLs in front of a cache layer.
Is DNS round robin a real load balancer?
It spreads traffic, but it does not check health and its answers are cached for the record's TTL, so clients keep trying a dead address. It is fine for coarse distribution between sites that each have their own load balancer, and poor as the only failover mechanism.
Do I need a load balancer if I use a CDN?
The CDN spreads visitors across its own edge servers and absorbs most of the traffic from cache. If your origin is one server, you do not need a balancer for capacity, though the origin is still a single point of failure for anything uncached. Once you run two or more origin servers, put a load balancer behind the origin address the CDN fetches from.