Load Balancing definition
Load balancing is the distribution of incoming network traffic across multiple servers or instances so that no single machine is overwhelmed. A load balancer sits in front of a group of servers, sends each request to a healthy one using an algorithm such as round robin or least connections, and removes failed servers from rotation.
How does load balancing work?
Clients connect to a single address, such as a DNS name or virtual IP, that belongs to the load balancer. The load balancer keeps a list of backend servers and runs health checks against them, for example requesting a health endpoint every few seconds. Each incoming request or connection goes to a healthy server chosen by the balancing algorithm. Servers that fail checks are removed from rotation until they recover, and servers being retired finish their active requests first.
Load balancers work at different layers. Layer 4 balancers route TCP or UDP connections using addresses and ports, which is fast and protocol-agnostic. Layer 7 balancers understand HTTP, so they can route by host name, path or header, terminate TLS, add security headers and send traffic for /api to one group of services and everything else to another.
Load balancing algorithms
- Round robin: each server takes the next request in turn.
- Weighted round robin: more powerful servers receive proportionally more requests.
- Least connections: new requests go to the server with the fewest active connections.
- Least response time: favors servers that are currently answering fastest.
- IP hash or consistent hashing: the same client or key goes to the same server.
- Power of two choices: pick two servers at random and send to the less busy one.
Types of load balancers
Hardware appliances from vendors such as F5 still run in many enterprise data centers. Software balancers such as NGINX, HAProxy and Envoy run on ordinary servers and inside container platforms. In the cloud, managed services handle scaling and redundancy: AWS Application and Network Load Balancers, Azure Load Balancer and Application Gateway, and Google Cloud Load Balancing. Global load balancing through DNS or anycast spreads users across regions, and Kubernetes provides Services and ingress controllers for traffic inside clusters.
Load balancing and high availability
Load balancing is the foundation of high availability. Placing servers in several availability zones behind a load balancer means a failed machine, or even a whole zone, removes only part of the capacity while traffic continues. The load balancer itself must be redundant, which managed cloud balancers handle automatically. Paired with auto scaling, it lets capacity grow and shrink while users always reach a working server.
Applications should be stateless for this to work well. Sticky sessions, which pin a user to one server, are sometimes necessary for legacy software but reduce resilience and make scaling uneven. Storing sessions in a shared cache such as Redis is the usual alternative.
Example: an ecommerce API
An ecommerce company runs its API on containers across three availability zones behind an application load balancer. Requests to /checkout go to a dedicated service with extra capacity, while browsing traffic goes to a general pool. Health checks remove a container within seconds if it stops responding, and deployments drain old containers before stopping them. Nexzem designs load balancing and failover like this and tests it by deliberately removing instances before launch.