ExpertDistinguishedReliability, Resilience & Recovery75 minutesPro answer

SDV-024

Build a chaos program that produces evidence instead of theater

Create hypothesis-driven resilience experiments with safety controls, production relevance, learning ownership, and portfolio-level risk reduction.

ReliabilityIncident ResponseBlast RadiusChange Management

Interview prompt

Problem context

Leadership mandates chaos engineering after a major outage. Teams run demo-friendly instance kills, but real incidents come from slow dependencies, bad configuration, quota exhaustion, and organizational handoffs. Design a program that safely tests meaningful hypotheses and leads to funded changes.

Skills being evaluated

resilience strategyexperiment designorganizational learningrisk governance

The full reasoning guide is part of Pro

The scenario and evaluation focus above remain public. Pro unlocks the structured answer, trade-off analysis, follow-up probes, common weak answers, rubric, related reasoning, and any architecture diagram.

Sign in to continue

What the full guide covers

Clarify the decision
Establish scale assumptions
Functional and non-functional requirements
High-level architecture
Data model and flow
Consistency and transaction boundaries
Failure modes and recovery
Security and privacy
Observability and SLOs
Capacity and cost
Alternatives and trade-offs
Evolution and migration
What Staff and Principal candidates should emphasize