Arc 05 · Senior AI Engineer
Transformers and Foundation-Model Internals
Trace the complete path from bytes to tokens, attention, transformer blocks, training objectives, decoding, and inference state.
Exit capability: Explain and capacity-plan the machinery behind modern language models.
- Mapped lessons
- 8
- Published now
- 1
- Full-arc estimate
- ≈19 hours
- Last edited
- 2026-08-11
Lesson sequence
Live units open into complete labs. Planned units stay visible to show the dependency path, but intentionally have no detail route.
- 01
Tokenization as a Compression Contract
PlannedStudy how vocabulary construction changes sequence length, multilingual behavior, cost, and failure modes.
Build labIntermediate105 min estimate
- 02
Derive Attention from Content-Based Routing
PlannedBuild scaled dot-product attention from the need to route information between positions.
Build labIntermediate120 min estimate
- 03
Multi-Head Attention Is Parallel Representation Routing
PlannedExplain head dimension, projections, and what head diversity does and does not guarantee.
Concept labAdvanced95 min estimate
- 04
Position, Context, and Extrapolation
PlannedCompare positional mechanisms by the invariants they encode and how they fail outside training lengths.
Failure labAdvanced105 min estimate
- 05
Assemble and Test a Transformer Block
PlannedCompose attention, MLP, normalization, masking, and residual paths with shape tests.
Build labAdvanced140 min estimateArtifact: Minimal transformer block
- 06
Pretraining Objectives Shape Model Behavior
PlannedConnect next-token prediction, masking, data mixtures, and preference objectives to observable capabilities.
Concept labAdvanced105 min estimate
- 07
Decoding Is a Product Policy
PlannedTreat temperature, top-p, constraints, and stopping as explicit quality and risk decisions.
Failure labIntermediate90 min estimate
- 08
The KV Cache Capacity Plan
LiveDerive per-token cache memory, then connect context length, concurrency, precision, batching, and paging to serving capacity.
Systems labAdvanced120 min estimateArtifact: Tested KV-cache estimator