Glossary · Data & representations
Tokenization
Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
Why it matters
Tokenization determines sequence length, vocabulary boundaries, cost accounting, truncation behavior, and how text or code is represented before embedding.
In practice
Use the exact tokenizer for the target model, version it with artifacts, and test multilingual text, code, whitespace, and special tokens.
Common confusion
Tokenization is not always word splitting, and two models can assign different token counts and IDs to the same input.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.