Glossary · Data & representations

Tokenization

Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.

Why it matters

Tokenization determines sequence length, vocabulary boundaries, cost accounting, truncation behavior, and how text or code is represented before embedding.

In practice

Use the exact tokenizer for the target model, version it with artifacts, and test multilingual text, code, whitespace, and special tokens.

Common confusion

Tokenization is not always word splitting, and two models can assign different token counts and IDs to the same input.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.