A learning project to implement the path from text tokenization to autoregressive training, while keeping the questions, failed assumptions, and verification evidence visible.
This project has not started yet. For now, this page shows the plan and the structure future entries will follow.
What I want to learn
Understand the full data and gradient path well enough to explain each boundary, reproduce a small model, and diagnose failures without treating the reference implementation as a black box.
Development plan
01
In progress
Foundations and tokenizer
Text representation, byte-level BPE, vocabulary, and reversible encoding.
02
Planned
Data pipeline
Sequence windows, batching, targets, and deterministic data loading.
03
Planned
Transformer model
Embeddings, causal attention, MLP blocks, normalization, and residual paths.
04
Planned
Training and evaluation
Optimization, checkpoints, sampling, evaluation, and failure analysis.
Entries
As the project develops, I will add the questions, implementation notes, and test results from each stage here.
01
Tokenizer · preparation
Planned
Before writing the tokenizer: what actually becomes a token?
A preview record for the first conceptual boundary in the project: the relationship between Unicode text, UTF-8 bytes, BPE merges, and the token IDs received by the model.