IntermediateSeniorReliability, Resilience & Recovery50 minutesPro answer

SDV-017

Contain cascading failures across services that share hidden dependencies

Find common pools and retry amplification, then design bulkheads, deadlines, admission, and recovery that contain a dependency slowdown.

Blast RadiusReliabilityLoad SheddingIncident Response

Interview prompt

Problem context

A metadata database slows down for eight minutes. Seemingly unrelated APIs exhaust connection pools, callers retry, background jobs compete with interactive work, and health checks restart healthy instances. The incident becomes a site-wide outage. Redesign the failure boundaries and response.

Skills being evaluated

failure analysisbulkhead designretry controlrecovery planning

The full reasoning guide is part of Pro

The scenario and evaluation focus above remain public. Pro unlocks the structured answer, trade-off analysis, follow-up probes, common weak answers, rubric, related reasoning, and any architecture diagram.

Sign in to continue

What the full guide covers

Clarify the decision
Establish scale assumptions
Functional and non-functional requirements
High-level architecture
Data model and flow
Consistency and transaction boundaries
Failure modes and recovery
Security and privacy
Observability and SLOs
Capacity and cost
Alternatives and trade-offs
Evolution and migration
What Staff and Principal candidates should emphasize