Glossary · Data & representations

Data Deduplication

Detecting and removing exact and near-duplicate examples within or across datasets.

Why it matters

Repetition can distort the training distribution, increase memorization, leak test material, and make evaluation appear stronger than it is.

In practice

Normalize content, use exact hashes and similarity methods, review borderline clusters, and record which version and rule removed each example.

Common confusion

Deduplication is not ordinary data cleaning. Two distinct records can legitimately share text, and two paraphrases can still carry the same leaked information.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.