Glossary · Data & representations

Byte Pair Encoding (BPE)

A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.

Why it matters

It balances vocabulary size with the ability to represent rare or unseen words as smaller units.

In practice

Train the tokenizer only on approved corpus splits, version its merge rules with the model, and inspect how it segments code, multilingual text, and whitespace.

Common confusion

BPE is one tokenizer family, not a universal description of how every model creates tokens.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.