Project 01

Building GPT-2 from scratch

A learning project to implement the path from text tokenization to autoregressive training, while keeping the questions, failed assumptions, and verification evidence visible.

Status
Planning
Started
Not started
Planned tools
Python · PyTorch · NumPy · GPT-2 reference implementation

This project has not started yet. For now, this page shows the plan and the structure future entries will follow.

What I want to learn

Understand the full data and gradient path well enough to explain each boundary, reproduce a small model, and diagnose failures without treating the reference implementation as a black box.

Development plan

  1. 01
    In progress

    Foundations and tokenizer

    Text representation, byte-level BPE, vocabulary, and reversible encoding.

  2. 02
    Planned

    Data pipeline

    Sequence windows, batching, targets, and deterministic data loading.

  3. 03
    Planned

    Transformer model

    Embeddings, causal attention, MLP blocks, normalization, and residual paths.

  4. 04
    Planned

    Training and evaluation

    Optimization, checkpoints, sampling, evaluation, and failure analysis.

Entries

As the project develops, I will add the questions, implementation notes, and test results from each stage here.

01

Tokenizer · preparation

Planned

Before writing the tokenizer: what actually becomes a token?

A preview record for the first conceptual boundary in the project: the relationship between Unicode text, UTF-8 bytes, BPE merges, and the token IDs received by the model.

  • GPT-2
  • Tokenizer
  • Byte-level BPE
Read entry