Arc 09 · Staff AI Engineer
Evaluation, Safety, and Reliability
Make quality a release system: task-specific datasets, calibrated judges, human review, security tests, telemetry, and incident learning.
Exit capability: Define a release bar that catches quality, safety, and reliability regressions.
- Mapped lessons
- 7
- Published now
- 0
- Full-arc estimate
- ≈17 hours
- Last edited
- 2026-08-11
Lesson sequence
Live units open into complete labs. Planned units stay visible to show the dependency path, but intentionally have no detail route.
- 01
An Eval Is a Decision System
PlannedStart from a release decision, then choose cases, graders, uncertainty, and escalation.
Design reviewAdvanced105 min estimate
- 02
Build a Golden Dataset That Can Disagree with You
PlannedSample real tasks, hard negatives, long tails, and policy boundaries without teaching to the test.
Build labAdvanced110 min estimate
- 03
Calibrate LLM Judges
PlannedMeasure position bias, verbosity bias, self-preference, variance, and disagreement against human labels.
Failure labAdvanced115 min estimate
- 04
Online Evaluation Without Shipping Blind
PlannedConnect shadowing, canaries, holdbacks, user signals, and rollback thresholds.
Systems labAdvanced105 min estimate
- 05
Observe the Decision Path, Not Just the Model Call
PlannedTrace retrieval, prompts, tools, model versions, policy checks, latency, and cost with controlled data capture.
Systems labAdvanced110 min estimate
- 06
Build a Safety Case for an AI Feature
PlannedTie hazards to controls, evidence, residual risk, ownership, and launch conditions.
Design reviewAdvanced105 min estimate
- 07
AI Incident Response and Correction Loops
PlannedContain harm, preserve evidence, distinguish model and system causes, and turn findings into durable gates.
Failure labAdvanced110 min estimateArtifact: AI incident playbook