Glossary · Data & representations
Data Deduplication
Detecting and removing exact and near-duplicate examples within or across datasets.
Why it matters
Repetition can distort the training distribution, increase memorization, leak test material, and make evaluation appear stronger than it is.
In practice
Normalize content, use exact hashes and similarity methods, review borderline clusters, and record which version and rule removed each example.
Common confusion
Deduplication is not ordinary data cleaning. Two distinct records can legitimately share text, and two paraphrases can still carry the same leaked information.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.