How to Deduplicate Systematic Review Search Results
Combine records from multiple databases, remove duplicate citations carefully, and keep the counts needed for a transparent review.
Why duplicates appear
The same report can be indexed by PubMed, Scopus, Web of Science, OpenAlex, Europe PMC, and other sources. Formatting differences make duplicates harder to detect: a DOI can use a URL prefix, punctuation can vary, author names can be abbreviated, and online-first dates can differ from issue dates.
A safe matching order
1. Normalized DOI match
Remove prefixes such as https://doi.org/, normalize case, and compare the identifier. Conflicting valid DOIs should outweigh a similar title.
2. Normalized title similarity
Normalize punctuation, spacing, and case before comparing titles. Review uncertain matches, especially for short or generic titles.
3. Author, year, and title checks
Use the first author, publication year, and title prefix as supporting evidence when a DOI is missing.
Deduplication workflow
- STEP 1
Export complete metadata
Include DOI, title, abstract, authors, year, journal, and source identifiers whenever the database offers them.
- STEP 2
Keep source files separate
Name each export by database and date. Separate inputs make troubleshooting and source counts easier.
- STEP 3
Combine and match
Run exact identifier matching first, followed by normalized title and metadata rules.
- STEP 4
Verify and record counts
Check uncertain pairs and retain the number removed before title and abstract screening begins.
Common deduplication mistakes
- ×Matching only exact titles and missing punctuation or subtitle differences.
- ×Merging records with different valid DOIs because their titles look similar.
- ×Treating multiple reports from one study as duplicate citations without checking whether they report different outcomes or follow-up periods.
- ×Editing source files before preserving the original database exports.
Automatic citation deduplication in Lumina
Lumina checks normalized DOI values, title similarity, and author-year-title combinations when records enter a project. The current title-similarity threshold is 85%. If two records contain different valid DOIs, that identifier conflict prevents a fuzzy-title match from merging them.
Import feedback reports how many records were added and how many were skipped as duplicates. Always review unexpected counts before screening.