What I ran into
A language model receives integer IDs, while the input begins as human-readable Unicode text. I want to make every transformation between those two representations explicit before writing the implementation.
The first risk is to copy an existing tokenizer line by line without understanding which layer handles characters, bytes, vocabulary entries, and merge ranks.
The question
Why does GPT-2 begin from UTF-8 bytes and byte-level BPE instead of treating Unicode characters as the base vocabulary, and what invariants must encode and decode preserve?
What I think so far
Starting from bytes gives the tokenizer a finite base alphabet that can represent any UTF-8 input, so unknown characters do not require a separate fallback token.
BPE then learns reusable sequences over that byte representation. This is only a working explanation until it agrees with the original GPT-2 encoder and round-trip tests.
How I will test it
- 01
Round trip
Check decode(encode(text)) === text across ASCII, Chinese, Japanese, emoji, whitespace, and mixed input.
- 02
Reference
Compare token IDs and decoded text with the original GPT-2 vocabulary and merge rules.
- 03
Failure cases
Record malformed bytes, repeated whitespace, combining marks, and boundary cases instead of hiding them behind a fallback.
What I know now
No conclusion yet. This section will change only after the tokenizer exists and the verification plan has been run.