InterviewsVector
Arc 11
CapstoneAdvanced180 min estimateOriginal publication

AI Architect Capstone: Defend the System

Turn a business decision into a reviewable system boundary, operating envelope, evidence plan, failure model, ownership map, migration path, and executive tradeoff—then defend what remains uncertain.

Authorship
InterviewsVector
Published / updated
2026-09-27 / 2026-09-27
Review status
Artifact tests passing · primary sources recorded

Original InterviewsVector teaching. Executable artifacts are deterministic illustrative audits with focused tests and recorded primary sources; they do not claim legal interpretation, policy compliance, live-system validation, architecture approval, deployment, migration execution, production readiness, or safety certification.

The decision in one pass

A defensible AI architecture is a decision packet, not a diagram. Begin with the user and business decision, non-goals, affected stakeholders, risk and reversibility, workload distribution, quality and safety criteria, latency, availability, privacy, security, cost, compliance, and organizational constraints. Draw the end-to-end request, data, decision, feedback, and control paths and assign an owner to every boundary. Compare at least two credible options—including the operationally simplest one—using the same evidence and operating envelope. Select a design only after documenting assumptions, architectural decisions, consequences, failure modes, capacity and economics, evaluation, governance, migration, rollback, and observability. Present current evidence separately from forecasts, and residual risk separately from mitigations. Invite adversarial cross-functional review, convert gaps into owned actions, and identify the facts that would reverse the decision. The reference packet auditor checks invented completeness evidence; `READY_FOR_DECISION` means ready for accountable human review, not architecture approval or production certification.

Why this matters

AI architecture decisions couple probabilistic behavior, product policy, data and model supply chains, expensive infrastructure, third-party providers, human operations, and evolving regulation. A polished component diagram can conceal an undefined quality bar, a non-reversible state change, a platform team that owns no outcome, or cost that scales faster than value. Senior architecture work makes these constraints, tradeoffs, evidence, owners, and escape paths legible enough for product, engineering, safety, security, finance, legal, and operations leaders to make and revisit a consequential decision.

You will be able to

  • Frame an AI architecture from the decision, stakeholders, workload, operating envelope, risk, reversibility, and organizational ownership rather than from a favored model or vendor.
  • Produce a coherent packet spanning system boundaries, data and model lifecycle, evaluation, safety, security, reliability, infrastructure, cost, governance, migration, observability, and incident response.
  • Compare alternatives with shared criteria, record architectural decisions and consequences, and distinguish measured evidence from assumptions and forecasts.
  • Run an adversarial cross-functional review that exposes failure paths, ownership gaps, stale evidence, high risks, and one-way-door choices before approval.
  • Defend a recommendation to technical and executive audiences while naming residual uncertainty, trigger points, and what evidence would change the decision.

Your Vector Loop for this lab

  1. 01

    Model

    Map the business and user decision, stakeholders, risks, non-goals, workload and growth, quality and safety, latency and availability, data and model supply chains, trust boundaries, cost, policy, ownership, dependencies, and change horizon.

  2. 02

    Derive

    Derive hard constraints, operating-envelope gates, decision criteria, credible alternatives, capacity and cost models, failure budgets, evaluation and governance evidence, migration stages, rollback, and organizational interfaces.

  3. 03

    Build

    Build a versioned architecture packet containing context, diagrams, contracts, ADRs, option matrix, evidence register, risk register, failure and recovery plan, operating model, migration plan, decision request, and action log.

  4. 04

    Stress

    Inject demand growth, long inputs, provider outage, quality drift, data loss, prompt attack, tool misuse, tenant abuse, regional failure, cost shock, policy change, owner loss, rollback failure, and an assumption that proves false.

  5. 05

    Operate

    Observe user outcomes, quality and safety slices, SLOs, workload shape, capacity, cost and unit economics, dependencies, policy decisions, model and data drift, incidents, human operations, and architecture-decision triggers by revision.

  6. 06

    Defend

    Defend the selected option and rejected alternatives, evidence quality, uncertainty, residual risk, ownership, migration, economics, recovery, and the thresholds or new facts that would force a different architecture.

Frame the decision before drawing the system

Write the decision request in one sentence: who must decide what, by when, for which outcome, under which constraints. Then name the user journey, affected stakeholders, misuse and harm pathways, non-goals, time horizon, reversible and irreversible choices, and confidence of each input. Distinguish a product experiment from a platform commitment. A prototype can answer whether an experience creates value; it cannot by itself answer whether the organization can operate that experience safely and economically at the target scale.

InputEvidence formArchitectural consequence
user outcometask and cohort success criteriamodel, retrieval, tool, and human workflow boundary
workloadjoint arrival, sequence, tenant, and growth distributioncapacity, queue, cache, topology, and admission design
riskimpact map, threat model, safety case, obligationsauthority, controls, isolation, evidence, and rollback
economicsunit cost, budget, sensitivity, opportunity costrouting, build/buy, quality tier, and scale threshold
organizationowners, skills, on-call, procurement, change speedplatform boundary, interfaces, staffing, and delivery sequence

Build one packet that keeps the argument connected

  1. 01Context and decisionState outcome, users, stakeholders, constraints, non-goals, current system, requested decision, decision authority, review date, and validity horizon.
  2. 02Architecture and contractsShow request, data, model, tool, control, feedback, and human paths; document interfaces, state ownership, trust and failure boundaries, and exact revision identities.
  3. 03Alternatives and ADRsCompare credible options under the same quality, safety, reliability, latency, cost, control, switching, and organizational criteria; record context, decision, consequences, owner, and supersession path.
  4. 04Evidence and operating planAttach evaluation, capacity, cost, threat, safety, governance, migration, rollback, observability, incident, and human-operation evidence with owners, age, gaps, and revalidation triggers.
  5. 05Decision and action logSeparate blocking gaps from accepted residual risk, assign actions and dates, record dissent, and ask the accountable authority for one explicit bounded decision.

Use diagrams to reveal behavior, not decorate the packet. A context view explains actors and external dependencies; a container or component view shows ownership and contracts; request and data sequences expose latency and authority; a deployment view shows regions, networks, accelerators, and failure domains; and a control-loop view shows evaluation, rollout, monitoring, governance, and rollback. Every arrow should answer who owns it, what contract crosses it, how it fails, and how it is observed.

Defend an operating envelope, not a single benchmark

Architecture tradeoffs are coupled. A larger model may improve one quality slice while increasing queueing, spend, regional scarcity, and recovery time. Caching can improve latency and cost while creating staleness, privacy, and invalidation requirements. A managed provider may accelerate delivery while changing data rights, observability, switching cost, outage control, and roadmap dependence. Put every option through the same representative workload and hard gates before comparing its frontier.

eligible(option) = quality ∧ safety ∧ latency ∧ reliability ∧ security ∧ governance ∧ budget ∧ operability

Weighted scores are useful only after hard constraints pass. A cheap option that violates a critical safety or reliability requirement is not made eligible by winning many low-consequence categories.

TradeoffEvidence to requestDecision trigger
build vs manageddifferentiation, control, staffing, TCO, data rights, exit proofvolume, roadmap, risk, or provider boundary changes
quality vs latencyslice quality, TTFT/E2E tails, fallback behaviorcritical slice or SLO leaves envelope
reliability vs costfailure model, redundancy, recovery tests, capacity headroomblast radius or recovery objective changes
platform vs product ownershipreuse demand, variance, interface stability, support loadstandardization tax exceeds shared leverage
speed vs reversibilitymigration stages, state compatibility, rollback rehearsalone-way door lacks sufficient evidence

Audit packet completeness before asking for a decision

The reference artifact audits invented metadata across six architecture dimensions. Each dimension is bound to the exact packet, system, and architecture revision and carries a named owner, ADR identity, evidence digest and age, status, assumptions, failure scenarios, rollback, and recorded tradeoff. Quality, latency, cost, and capacity results live in a separate provenance-bound operational record with source evidence IDs, observation time, freshness, and owner; packet-level gates also cover alternatives, high risks, and executive tradeoffs. It rejects missing, duplicate, unknown, stale, reused, forged, wrong-type, out-of-scope, and digest-tampered evidence. It does not read the underlying evidence or judge whether the selected design is good.

courses/ai-engineering/reference-impl/architect_capstone/architecture_review.py
1def review_architecture(contract: ArchitectureContract, evidence: ArchitectureEvidence) -> ArchitectureReport:
2 contract = validate_record(contract, ArchitectureContract)
3 evidence = validate_record(evidence, ArchitectureEvidence)
4 if evidence.scope != contract.scope or evidence.contract_content_id != contract.content_id:
5 raise ValueError("evidence belongs to another architecture contract")
6
7 # Dimension and operational evidence are packet/system/revision-bound.
8 # Gate provenance freshness, alternatives, operating results, high risk,
9 # failure coverage, ownership, rollback, and executive tradeoffs.

Expected output

example=illustrative_only
decision=READY_FOR_DECISION
packet=review-packet-2026-09-27;system=support-assistant
dimensions=6/6;failure_scenarios=12
selected_option=bounded-platform-with-managed-models
claim=LOCAL_PACKET_COMPLETENESS_AUDIT_NOT_ARCHITECTURE_APPROVAL

Verify: python3 -m unittest discover -s courses/ai-engineering/reference-impl/architect_capstone -p 'test_*.py' -v

Review the system where disciplines collide

The highest-value review questions cross boundaries: Can the quality evaluator see the cohort the admission controller sheds? Does provider fallback preserve data and safety policy? Can the incident team identify the prompt, retrieval, model, tool, and policy revisions for one harmful response? Does autoscaling retain rollback capacity? Can finance explain unit cost when retries and long outputs rise? Can a product owner change a tool permission without reopening the threat model? Invite the people who own those consequences, not only the authors of the architecture.

Injected changeCross-layer questions
10× demand with longer promptsqueueing, KV and memory, quotas, quality tiering, spend, dependency capacity, degraded mode
primary provider outagerouting, data policy, semantic parity, cache, capacity, cost, observability, customer communication
prompt or retrieval attacktrust boundary, content provenance, tool authority, tenant isolation, detection, containment, evidence
quality drifts on one languageslice coverage, ownership, rollout identity, fallback, human review, suspension, remediation
platform owner leavesrunbook, decision rights, on-call, vendor and model knowledge, secrets, roadmap, succession
new regulation or policyinventory, applicability, control gap, evidence, exception, migration, regional scope, decision expiry

Ask for a decision that leaders can actually make

The executive summary should state the outcome, recommendation, alternatives rejected, decisive evidence, expected value, cost range, major risks, mitigations, residual risk, owners, timeline, decision requested, and what would change the recommendation. Put detail behind it without hiding uncertainty. Use ranges and sensitivity for forecasts. Separate facts, assumptions, and judgments. Name disagreement and one-way doors. Technical leadership is not winning an argument; it is creating the shared model that lets accountable people make, execute, and later revisit a difficult decision.

  • Translate quality, latency, safety, and reliability into user and business consequences without stripping away the technical boundary of the evidence.
  • Pair every material risk with owner, mitigation, detection, response, residual severity, due date, and escalation—not a color alone.
  • Assign operational ownership before build ownership ends: on-call, model and data changes, evaluation, capacity, cost, policy, incidents, vendor management, and retirement.
  • Record the decision and its context as an ADR; supersede it when evidence or constraints change rather than silently rewriting history.
  • Schedule architecture revalidation at meaningful milestones and triggers, not only after incidents or immediately before launch.

Operate at three altitudes

Production lens

  • — Operate the architecture packet as a versioned claim: join user outcomes, quality and safety, workload, SLOs, capacity, cost, dependencies, controls, incidents, ownership, migration, and ADR triggers to the exact architecture and release revision.
  • — Rehearse demand, sequence, provider, region, data, model, retrieval, tool, safety, security, cost, policy, migration, rollback, human-operation, and owner-loss failures across the complete system rather than one service at a time.
  • — Continuously compare observed conditions with the approved operating envelope and reopen the decision when a hard constraint, assumption, dependency, risk, or organizational boundary changes.

Staff lens

  • — Create decision quality across disciplines: make product, engineering, model, data, evaluation, safety, security, privacy, legal, infrastructure, finance, operations, support, and executive owners see the same alternatives, evidence, uncertainty, and actions.
  • — Match review rigor to consequence and reversibility, keep two-way doors lightweight, inspect one-way doors deeply, and preserve dissent and superseding decisions so organizational learning survives team and vendor change.

Interview defense

Defend the architecture for a multi-tenant enterprise AI assistant that uses retrieval and tools, must meet strict quality, latency, privacy, reliability, and cost requirements, and is expected to grow tenfold.

I would begin with the user decisions, tenant and regional context, critical harms, workload distributions, growth, quality and safety rubrics, SLOs, privacy and security constraints, unit economics, and owner map. I would draw request, data, retrieval, model, tool, policy, feedback, and control paths with trust and failure boundaries, then compare at least a managed-provider, self-hosted, and hybrid option under the same hard gates. The packet would include ADRs, model and data contracts, evaluation and release gates, tenant isolation and tool authority, capacity and cost sensitivity, failure and degraded modes, observability and incident evidence, governance decisions, and an expand-migrate-verify-contract plan with rollback. I would load-test representative sequence and tenant slices, exercise provider and region failure, and quantify headroom rather than quote average throughput. The recommendation would state measured evidence, forecasts, residual risks, owners, and triggers such as demand, provider terms, quality drift, or regulatory scope that reopen it. I would ask the review board for a bounded staged decision, not a blanket production certification.

Expect the interviewer to press on

  • — How do you compare a managed-provider option with self-hosting without hiding switching and organizational cost?
  • — Which architecture decisions are one-way doors in this system?
  • — What evidence would make you reverse your recommendation six months later?

Misconceptions to remove

“The most detailed architecture diagram is the strongest design.”

A diagram is useful only when connected to decisions, contracts, evidence, failure behavior, ownership, operations, and tradeoffs. Visual detail can still hide an undefined outcome or recovery path.

“A weighted option score can decide the architecture objectively.”

Weights embed judgment and can let low-consequence wins offset a violated hard constraint. Gate mandatory quality, safety, reliability, security, governance, budget, and operability first, then compare eligible tradeoffs.

“Architecture review is the approval meeting before launch.”

Review should begin while choices are cheap, continue through implementation and migration, and recur when workload, evidence, policy, providers, incidents, or organizational boundaries change.

Check your model

1. Why must the packet distinguish evidence, assumptions, forecasts, and judgments?

They have different confidence and validation paths. Mixing them makes a projected capacity number or subjective risk acceptance look like a measured fact and hides what must be tested or approved.

2. Why compare alternatives only after applying hard gates?

A weighted ranking can allow many minor benefits to offset failure of a non-negotiable safety, security, quality, reliability, governance, or budget boundary. Ineligible options should leave the frontier before optimization.

3. What makes an ADR useful after the original team leaves?

It preserves the decision context, options, reasons, consequences, owner, status, and supersession path, letting future teams understand why the design exists and which changed facts justify replacing it.

Prove the mechanism

Create a review packet for an invented AI product with three credible architecture options. Include context, hard gates, workload, diagrams, ADRs, evidence and risk registers, capacity and cost sensitivity, failure rehearsal, governance, ownership, migration, rollback, decision request, dissent, and five quantified triggers that would supersede the recommendation.

Add a production constraint

Defend the same product under a board-mandated cost reduction, a new high-impact regional use, a tenfold demand forecast, loss of the primary model provider, and departure of the platform owner. Produce the revised option frontier, staged migration, organizational plan, residual-risk decision, and ninety-day evidence roadmap without claiming certainty you do not have.

Artifact: AI architecture review packet

courses/ai-engineering/reference-impl/architect_capstone/architecture_review.py

Download reference implementation

Primary references and next links

References

  1. 1. Well-Architected machine learning design principles

    AWS Well-Architected Machine Learning Lens. Official ML workload principles covering ownership, protection, resilience, reproducibility, modularity, and continuous improvement.

  2. 2. Architectural decision record process

    AWS Prescriptive Guidance. Official guidance for recording an architectural decision, its context and consequences, owner, review lifecycle, immutability, and superseding decision.

  3. 3. The review process

    AWS Well-Architected Framework. Official guidance framing architecture review as an early, continuous, cross-functional, risk-oriented conversation that produces improvement actions rather than a one-time certification.

Continue through the graph

Glossary: architecture decision record · operating envelope · hard constraint · tradeoff frontier · one-way door · residual risk · failure boundary · decision authority · evidence register · architecture trigger