Arc 05 · Senior AI Engineer
Transformers and Foundation-Model Internals
Trace the complete path from bytes to tokens, attention, transformer blocks, training objectives, decoding, and inference state.
Exit capability: Explain and capacity-plan the machinery behind modern language models.
- Mapped lessons
- 8
- Published now
- 8
- Full-arc estimate
- ≈19 hours
- Last edited
- 2026-08-25
Lesson sequence
Live units open into complete labs. Planned units stay visible to show the dependency path, but intentionally have no detail route. A withdrawn unit is retained only at its previously advertised URL and is not presented as live.
- 01
Tokenization as a Compression Contract
LiveStudy how vocabulary construction changes sequence length, multilingual behavior, cost, and failure modes.
Build labIntermediate105 min estimateArtifact: Tokenizer compression audit
- 02
Derive Attention from Content-Based Routing
LiveBuild scaled dot-product attention from the need to route information between positions.
Build labIntermediate120 min estimateArtifact: Attention routing audit
- 03
Multi-Head Attention Is Parallel Representation Routing
LiveExplain head dimension, projections, and what head diversity does and does not guarantee.
Concept labAdvanced95 min estimateArtifact: Multi-head routing audit
- 04
Position, Context, and Extrapolation
LiveCompare positional mechanisms by the invariants they encode and how they fail outside training lengths.
Failure labAdvanced105 min estimateArtifact: Position extrapolation audit
- 05
Assemble and Test a Transformer Block
LiveCompose attention, MLP, normalization, masking, and residual paths with shape tests.
Build labAdvanced140 min estimateArtifact: Minimal transformer block
- 06
Pretraining Objectives Shape Model Behavior
LiveConnect next-token prediction, masking, data mixtures, and preference objectives to observable capabilities.
Concept labAdvanced105 min estimateArtifact: Training objective audit
- 07
Decoding Is a Product Policy
LiveTreat temperature, top-p, constraints, and stopping as explicit quality and risk decisions.
Failure labIntermediate90 min estimateArtifact: Decoding policy audit
- 08
The KV Cache Capacity Plan
LiveDerive per-token cache memory, then connect context length, concurrency, precision, batching, and paging to serving capacity.
Systems labAdvanced120 min estimateArtifact: Tested KV-cache estimator