Glossary · Data & representations
Dataset Split
A documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
Why it matters
Separation prevents the evidence used to choose a system from also serving as independent proof that the chosen system generalizes.
In practice
Split by the real deployment unit, such as user, repository, organization, or time, rather than randomly dividing correlated rows.
Common confusion
A random split is not automatically independent. Near duplicates, future observations, or records from the same entity can cross the boundary.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.