Skip to content

What is Chaos Engineering?

DevOps & Reliability, explained by the engineers who build it. Definition, how it works, use cases and common questions.

Chaos Engineering definition

Chaos engineering is the practice of deliberately injecting failures into a system, such as killing servers, adding network latency or breaking dependencies, to verify that it keeps working as expected. Run as controlled experiments with a hypothesis and a limited blast radius, it uncovers weaknesses before they cause real outages. Netflix popularized it with its Chaos Monkey tool.

Why break things on purpose?

Modern systems are built from many services, cloud components and third-party APIs, and their failure behavior is hard to predict from design documents. Timeouts may be missing, retries may amplify load, failover may depend on a credential nobody has used in a year. Chaos engineering tests these assumptions directly, in a controlled way and at a time the team chooses, instead of waiting for an incident at 3 a.m. to reveal them.

Netflix built Chaos Monkey to randomly terminate production instances, forcing engineers to design services that survive the loss of any single server. The discipline has since spread well beyond streaming to banks, retailers and cloud providers, often as part of site reliability engineering practice.

How to run a chaos experiment

A chaos experiment is a scientific test rather than random destruction. Each one follows the same basic steps, and it is stopped immediately if user impact goes beyond the limits agreed with stakeholders beforehand. Experiments are written down in advance so anyone can review them.

Good candidates for first experiments are failures that already happened in past incidents or near misses, because the team knows they are realistic and can confirm whether the fixes made afterward actually work. The steps are:

  • Define steady state: measurable signals of normal behavior, such as orders per minute or error rate
  • Form a hypothesis: for example, if one availability zone fails, checkout errors stay below 1 percent
  • Limit the blast radius: start in staging or with a small share of production traffic
  • Inject the failure: terminate instances, add latency, block a dependency or exhaust CPU
  • Observe and compare against steady state using monitoring and traces
  • Fix what you learn, then automate the experiment so it runs regularly

Example failure scenarios

Useful early experiments are simple. Terminate a random pod or instance and confirm traffic shifts without errors. Add 500 ms of latency to a payment provider call and check that timeouts, retries and user messaging behave sensibly. Make the cache unavailable and see whether the database survives the extra load. Fail over the primary database and measure how long writes are interrupted.

Each experiment usually finds something: an alert that never fired, a retry storm, a dashboard that hid the problem. Those findings feed directly into runbooks and observability improvements. Teams also run game days, scheduled sessions where people practice responding to injected failures together.

Tools and where to start

Tools include Chaos Monkey from Netflix, Gremlin, LitmusChaos and Chaos Mesh for Kubernetes, AWS Fault Injection Service and Azure Chaos Studio. Most can target specific instances, containers or network paths and stop automatically when a safety alarm triggers, which keeps experiments within the agreed blast radius.

Start only after basic monitoring, alerting and high availability are in place; chaos experiments on a system with no redundancy simply cause outages. Nexzem introduces chaos experiments gradually for clients running critical platforms, beginning in staging with clear hypotheses and rollback plans.

Chaos Engineering: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

Is chaos engineering safe to run in production?

It can be, when done carefully: start in staging, then move to production with a small blast radius, clear abort conditions, monitoring in place and the team ready to respond. Production matters because staging rarely matches real traffic and configuration, but it should never be the first step.

What is Chaos Monkey?

Chaos Monkey is an open-source tool created by Netflix that randomly terminates virtual machine instances or containers in production during business hours. It pushes teams to build services that tolerate the loss of individual servers, and it gave its name to the broader practice of chaos engineering.

What is a game day?

A game day is a planned exercise where a team deliberately causes a failure, such as a region outage or database failover, and practices detecting, diagnosing and recovering from it. Game days test people, runbooks and communication as well as technology, and usually produce a list of concrete improvements.

Keep exploring the devops & reliability glossary

Need Chaos Engineering in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.