SRE definition
Site Reliability Engineering (SRE) is a discipline, created at Google, that applies software engineering to operations work in order to run large systems reliably. SRE teams define reliability targets called service level objectives, use error budgets to balance new features against stability, automate repetitive operational tasks and lead incident response and blameless postmortems.
How SRE works
Google started SRE in 2003, when Ben Treynor Sloss was asked to run a production team and staffed it with software engineers instead of traditional system administrators. Google later published its practices in the book Site Reliability Engineering in 2016, which spread the approach widely. The central idea is to treat reliability as an engineering problem, measured with data and improved with code and automation.
Reliability is measured with service level indicators (SLIs), such as the share of requests that succeed or complete within 300 milliseconds. A service level objective (SLO) sets a target for an SLI, such as 99.9 percent of requests succeeding over 30 days. A service level agreement (SLA) is a contractual promise to customers, usually looser than the internal SLO, with penalties if it is missed.
Key SRE concepts
- Error budget: the amount of unreliability an SLO allows, spent on releases and experiments.
- Toil: manual, repetitive operational work that should be automated away.
- Blameless postmortems: incident reviews that focus on systems and fixes, not blame.
- On-call: sustainable rotations with clear escalation and alert quality standards.
- Capacity planning and load testing ahead of demand.
- Release engineering: safe, gradual rollouts with fast rollback.
- Graceful degradation, so partial failures keep core features working.
- Chaos engineering: deliberately injecting failures to test resilience.
Error budgets: a worked example
A 99.9 percent availability SLO over 30 days allows about 43 minutes of failure, since 0.1 percent of 43,200 minutes is 43.2. That allowance is the error budget. While budget remains, the team ships features at full speed. If a bad release and a dependency outage consume the budget, the agreed policy kicks in: feature releases pause, and engineering effort shifts to reliability until the service is back within target.
The budget turns arguments between product and operations into a shared, data-driven decision. It also shows that 100 percent is the wrong target: users cannot tell the difference between very high reliability and perfection, and chasing the last fraction slows everything else down.
SRE vs DevOps
DevOps is a broad culture and set of practices for shared ownership and fast delivery. SRE is a concrete implementation of many DevOps ideas, with specific tools such as SLOs, error budgets and toil limits. Google itself describes SRE as one way to put DevOps into practice. Many companies run DevOps practices across all teams and add dedicated SREs for their most critical, high-traffic services.
How to adopt SRE
Start by defining SLOs for a few user-facing journeys, such as login, search and checkout, based on what users actually notice. Alert on SLO burn rate rather than on every CPU spike, run blameless postmortems for significant incidents, and track toil so it can be automated. Nexzem helps teams introduce SRE practices step by step, beginning with SLOs and observability for the services that matter most to revenue.