Skip to content

What is Site Reliability Engineering (SRE)?

DevOps & Reliability, explained by the engineers who build it. Definition, how it works, use cases and common questions.

SRE definition

Site Reliability Engineering (SRE) is a discipline, created at Google, that applies software engineering to operations work in order to run large systems reliably. SRE teams define reliability targets called service level objectives, use error budgets to balance new features against stability, automate repetitive operational tasks and lead incident response and blameless postmortems.

How SRE works

Google started SRE in 2003, when Ben Treynor Sloss was asked to run a production team and staffed it with software engineers instead of traditional system administrators. Google later published its practices in the book Site Reliability Engineering in 2016, which spread the approach widely. The central idea is to treat reliability as an engineering problem, measured with data and improved with code and automation.

Reliability is measured with service level indicators (SLIs), such as the share of requests that succeed or complete within 300 milliseconds. A service level objective (SLO) sets a target for an SLI, such as 99.9 percent of requests succeeding over 30 days. A service level agreement (SLA) is a contractual promise to customers, usually looser than the internal SLO, with penalties if it is missed.

Key SRE concepts

  • Error budget: the amount of unreliability an SLO allows, spent on releases and experiments.
  • Toil: manual, repetitive operational work that should be automated away.
  • Blameless postmortems: incident reviews that focus on systems and fixes, not blame.
  • On-call: sustainable rotations with clear escalation and alert quality standards.
  • Capacity planning and load testing ahead of demand.
  • Release engineering: safe, gradual rollouts with fast rollback.
  • Graceful degradation, so partial failures keep core features working.
  • Chaos engineering: deliberately injecting failures to test resilience.

Error budgets: a worked example

A 99.9 percent availability SLO over 30 days allows about 43 minutes of failure, since 0.1 percent of 43,200 minutes is 43.2. That allowance is the error budget. While budget remains, the team ships features at full speed. If a bad release and a dependency outage consume the budget, the agreed policy kicks in: feature releases pause, and engineering effort shifts to reliability until the service is back within target.

The budget turns arguments between product and operations into a shared, data-driven decision. It also shows that 100 percent is the wrong target: users cannot tell the difference between very high reliability and perfection, and chasing the last fraction slows everything else down.

SRE vs DevOps

DevOps is a broad culture and set of practices for shared ownership and fast delivery. SRE is a concrete implementation of many DevOps ideas, with specific tools such as SLOs, error budgets and toil limits. Google itself describes SRE as one way to put DevOps into practice. Many companies run DevOps practices across all teams and add dedicated SREs for their most critical, high-traffic services.

How to adopt SRE

Start by defining SLOs for a few user-facing journeys, such as login, search and checkout, based on what users actually notice. Alert on SLO burn rate rather than on every CPU spike, run blameless postmortems for significant incidents, and track toil so it can be automated. Nexzem helps teams introduce SRE practices step by step, beginning with SLOs and observability for the services that matter most to revenue.

SRE: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What does a site reliability engineer do?

An SRE keeps production systems reliable and scalable using engineering. Typical work includes defining SLOs, building monitoring and alerting, automating operational tasks, improving deployment safety, planning capacity, participating in on-call rotations, leading incident response and running postmortems that lead to lasting fixes.

What is the difference between SLI, SLO and SLA?

An SLI is a measurement, such as the percentage of successful requests. An SLO is an internal target for that measurement, such as 99.9 percent over 30 days. An SLA is an external contract with customers that promises a level of service and defines consequences, such as credits, when it is not met.

Is SRE only for large companies?

The full Google model suits large organizations, but core SRE ideas work at any size. Even a small team gains from defining SLOs for key user journeys, alerting on user impact instead of noisy metrics, writing blameless postmortems and steadily automating repetitive operational work.

Keep exploring the devops & reliability glossary

Need SRE in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.