Glossary · Data & representations
Byte Pair Encoding (BPE)
A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
Why it matters
It balances vocabulary size with the ability to represent rare or unseen words as smaller units.
In practice
Train the tokenizer only on approved corpus splits, version its merge rules with the model, and inspect how it segments code, multilingual text, and whitespace.
Common confusion
BPE is one tokenizer family, not a universal description of how every model creates tokens.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.