Load Testing definition
Load testing is a type of performance testing that simulates expected numbers of concurrent users or requests against an application to measure response times, throughput, error rates and resource use. It shows whether a system can handle normal and peak traffic, where it starts to slow down, and which component becomes the bottleneck first.
Types of performance tests
Load testing is one member of a family of performance tests. They use similar tools but answer different questions, so a good test plan states which question each run is meant to answer, and which environment and data it uses:
- Load test: expected normal and peak traffic, to confirm performance targets are met
- Stress test: increasing load beyond peak until something breaks, to find limits and failure behavior
- Spike test: a sudden jump in traffic, such as a sale launch or a push notification
- Soak or endurance test: sustained load for hours to reveal memory leaks and resource exhaustion
- Scalability test: measuring how added servers or instances change capacity
How to plan a realistic load test
A load test is only as useful as its realism. Base user journeys and traffic mix on analytics or production logs: how many people browse versus search versus check out, how long they pause between actions, and which API calls each step triggers. Use production-like data volumes, because a query that is instant on a hundred rows may crawl on ten million, and test an environment that matches production sizing.
Define pass criteria before the run, ideally from your SLOs: for example, at 2,000 concurrent users, 95 percent of requests complete under 400 ms with an error rate below 0.5 percent. Without targets agreed in advance, results become a debate rather than a decision. Agree the criteria with product owners, so passing the test means something to the business and not only to engineers.
Metrics to watch
Measure from the user's side: response time percentiles (p50, p95, p99), throughput in requests per second and error rate. Then correlate with the server side: CPU, memory, database connections, queue depths, garbage collection pauses and the latency of downstream services. The first resource to saturate is the bottleneck, and tracing shows where requests spend their time.
Look for the knee in the curve, the load level at which response times start rising sharply. Comfortable capacity sits well below that point, leaving headroom for unexpected spikes. Results also feed scalability planning and auto-scaling thresholds, so the system adds capacity before users feel the knee.
Tools and common pitfalls
Popular open-source tools include k6, Apache JMeter, Gatling, Locust and Artillery, while managed options such as Grafana Cloud k6, Azure Load Testing and BlazeMeter generate load from many regions. Scripts should live in version control and run in CI for critical journeys, so performance regressions are caught before release rather than after.
Common pitfalls include generating load from a single underpowered machine that becomes the bottleneck itself, caches that are warm in tests but cold after deploys, third-party APIs that rate-limit or bill you, and testing only the happy path. Nexzem runs load tests as part of performance testing before launches and major sales events.