InterviewsVector
Capability spine

Arc 05 · Senior AI Engineer

Transformers and Foundation-Model Internals

Trace the complete path from bytes to tokens, attention, transformer blocks, training objectives, decoding, and inference state.

Exit capability: Explain and capacity-plan the machinery behind modern language models.

Mapped lessons
8
Published now
1
Full-arc estimate
19 hours
Last edited
2026-08-11

Lesson sequence

Live units open into complete labs. Planned units stay visible to show the dependency path, but intentionally have no detail route.

  1. 01

    Tokenization as a Compression Contract

    Planned

    Study how vocabulary construction changes sequence length, multilingual behavior, cost, and failure modes.

    Build labIntermediate105 min estimate

  2. 02

    Derive Attention from Content-Based Routing

    Planned

    Build scaled dot-product attention from the need to route information between positions.

    Build labIntermediate120 min estimate

  3. 03

    Multi-Head Attention Is Parallel Representation Routing

    Planned

    Explain head dimension, projections, and what head diversity does and does not guarantee.

    Concept labAdvanced95 min estimate

  4. 04

    Position, Context, and Extrapolation

    Planned

    Compare positional mechanisms by the invariants they encode and how they fail outside training lengths.

    Failure labAdvanced105 min estimate

  5. 05

    Assemble and Test a Transformer Block

    Planned

    Compose attention, MLP, normalization, masking, and residual paths with shape tests.

    Build labAdvanced140 min estimateArtifact: Minimal transformer block

  6. 06

    Pretraining Objectives Shape Model Behavior

    Planned

    Connect next-token prediction, masking, data mixtures, and preference objectives to observable capabilities.

    Concept labAdvanced105 min estimate

  7. 07

    Decoding Is a Product Policy

    Planned

    Treat temperature, top-p, constraints, and stopping as explicit quality and risk decisions.

    Failure labIntermediate90 min estimate

  8. 08

    The KV Cache Capacity Plan

    Live

    Derive per-token cache memory, then connect context length, concurrency, precision, batching, and paging to serving capacity.

    Systems labAdvanced120 min estimateArtifact: Tested KV-cache estimator