01 Tokenizer · preparation

Before writing the tokenizer: what actually becomes a token?

A preview record for the first conceptual boundary in the project: the relationship between Unicode text, UTF-8 bytes, BPE merges, and the token IDs received by the model.

Planning note — not yet tested in code

What I ran into

A language model receives integer IDs, while the input begins as human-readable Unicode text. I want to make every transformation between those two representations explicit before writing the implementation.

The first risk is to copy an existing tokenizer line by line without understanding which layer handles characters, bytes, vocabulary entries, and merge ranks.

The question

Why does GPT-2 begin from UTF-8 bytes and byte-level BPE instead of treating Unicode characters as the base vocabulary, and what invariants must encode and decode preserve?

What I think so far

Starting from bytes gives the tokenizer a finite base alphabet that can represent any UTF-8 input, so unknown characters do not require a separate fallback token.

BPE then learns reusable sequences over that byte representation. This is only a working explanation until it agrees with the original GPT-2 encoder and round-trip tests.

How I will test it

  1. 01

    Round trip

    Check decode(encode(text)) === text across ASCII, Chinese, Japanese, emoji, whitespace, and mixed input.

  2. 02

    Reference

    Compare token IDs and decoded text with the original GPT-2 vocabulary and merge rules.

  3. 03

    Failure cases

    Record malformed bytes, repeated whitespace, combining marks, and boundary cases instead of hiding them behind a fallback.

What I know now

No conclusion yet. This section will change only after the tokenizer exists and the verification plan has been run.

Next step

Read the original encoder implementation, draw the complete encode/decode pipeline, and write the first failing round-trip tests before implementing BPE merges.

Back to project