AdvancedStaffReliability, Resilience & Recovery50 minutesPro answer

SDV-020

Eliminate a retry storm during a long partial outage

Coordinate deadlines, retry ownership, budgets, jitter, hedging, and durable deferral so recovery traffic cannot outnumber original demand.

ReliabilityLoad SheddingIncident ResponseIdempotency

Interview prompt

Problem context

A payment dependency returns a mix of timeouts, 429s, and 500s for forty minutes. Mobile clients, gateways, services, and queue workers all retry. Recovery begins, but the retry backlog immediately saturates the provider again. Design a stable policy across layers.

Skills being evaluated

retry semanticscontrol stabilityclient contractspayment correctness

The full reasoning guide is part of Pro

The scenario and evaluation focus above remain public. Pro unlocks the structured answer, trade-off analysis, follow-up probes, common weak answers, rubric, related reasoning, and any architecture diagram.

Sign in to continue

What the full guide covers

Clarify the decision
Establish scale assumptions
Functional and non-functional requirements
High-level architecture
Data model and flow
Consistency and transaction boundaries
Failure modes and recovery
Security and privacy
Observability and SLOs
Capacity and cost
Alternatives and trade-offs
Evolution and migration
What Staff and Principal candidates should emphasize