Latency definition
Latency is the time delay between a request and its response, or between an action and its visible effect, usually measured in milliseconds. In software it includes network travel time, queuing, processing and data access. Low latency makes apps feel instant, while high or unpredictable latency causes slow pages, laggy interfaces and timeouts between services.
What causes latency?
Latency is the sum of several delays, and the biggest one is often not where teams expect it, so measure before optimizing. A single page load or API call typically accumulates time from the following sources.
Each layer can be measured separately. Distributed tracing tools such as OpenTelemetry with Jaeger, Datadog or Grafana Tempo break a request into spans, showing exactly how many milliseconds went to the database, an external API or rendering, which turns latency work from guesswork into a ranked list of fixes. The usual sources are:
- Propagation: signals in fiber take time, so a round trip from Sydney to a US data center adds a noticeable delay before any work happens
- Connection setup: DNS lookups and TCP and TLS handshakes add round trips to the first request
- Queuing: requests waiting for a free thread, database connection or worker under load
- Processing: application code, serialization and server-side rendering
- Data access: database queries, disk reads and calls to other services or third-party APIs
- Client side: JavaScript execution and rendering on slower devices
Latency vs bandwidth vs throughput
Bandwidth is how much data a link can carry per second; latency is how long each piece takes to arrive; throughput is how much work the system actually completes per second. A highway analogy helps: bandwidth is the number of lanes, latency is the drive time and throughput is cars arriving per hour. Adding lanes does not shorten a trip that is slow because of distance.
For most web apps, latency dominates user experience, because pages need many small requests, each paying the round-trip cost. That is why moving content closer to users with a content delivery network or an in-region server often helps more than buying a faster server.
Measuring latency: averages, p95 and p99
Averages hide pain. If 99 requests take 50 ms and one takes 5 seconds, the average looks fine while one user in a hundred waits. Teams track percentiles instead: p50 is the median, while p95 and p99 show the slow tail. When one page triggers dozens of backend calls, tail latency compounds, so a rare slow call becomes a common slow page.
Measure where users are, not only in your data center: real user monitoring, synthetic checks from several regions and browser metrics such as Core Web Vitals. Set latency targets as SLOs, for example 99 percent of API requests under 300 ms, and alert when the error budget starts burning.
How to reduce latency
Start with the largest measured delay. Common fixes include serving static and cacheable content from a CDN, hosting in a region near users, caching expensive queries, adding database indexes, running independent calls in parallel instead of in sequence, reusing connections, compressing responses and moving slow work such as email or report generation into background jobs.
For AI features, stream model output so users see the first words quickly, and choose smaller models where quality allows. Nexzem treats latency budgets as part of design for web, mobile and AI products, profiling real traces before optimizing so effort goes where it actually helps users.