Citation Management Guide

How to Deduplicate Systematic Review Search Results

Combine records from multiple databases, remove duplicate citations carefully, and keep the counts needed for a transparent review.

By Lumina Editorial Team Published Updated Methodology guide
Combined academic database search results after DOI and title deduplication
Deduplicate after combining exports so overlap between databases is detected before screening.

Why duplicates appear

The same report can be indexed by PubMed, Scopus, Web of Science, OpenAlex, Europe PMC, and other sources. Formatting differences make duplicates harder to detect: a DOI can use a URL prefix, punctuation can vary, author names can be abbreviated, and online-first dates can differ from issue dates.

A safe matching order

1. Normalized DOI match

Remove prefixes such as https://doi.org/, normalize case, and compare the identifier. Conflicting valid DOIs should outweigh a similar title.

2. Normalized title similarity

Normalize punctuation, spacing, and case before comparing titles. Review uncertain matches, especially for short or generic titles.

3. Author, year, and title checks

Use the first author, publication year, and title prefix as supporting evidence when a DOI is missing.

Deduplication workflow

  1. STEP 1

    Export complete metadata

    Include DOI, title, abstract, authors, year, journal, and source identifiers whenever the database offers them.

  2. STEP 2

    Keep source files separate

    Name each export by database and date. Separate inputs make troubleshooting and source counts easier.

  3. STEP 3

    Combine and match

    Run exact identifier matching first, followed by normalized title and metadata rules.

  4. STEP 4

    Verify and record counts

    Check uncertain pairs and retain the number removed before title and abstract screening begins.

Common deduplication mistakes

  • ×Matching only exact titles and missing punctuation or subtitle differences.
  • ×Merging records with different valid DOIs because their titles look similar.
  • ×Treating multiple reports from one study as duplicate citations without checking whether they report different outcomes or follow-up periods.
  • ×Editing source files before preserving the original database exports.

Automatic citation deduplication in Lumina

Lumina checks normalized DOI values, title similarity, and author-year-title combinations when records enter a project. The current title-similarity threshold is 85%. If two records contain different valid DOIs, that identifier conflict prevents a fuzzy-title match from merging them.

Import feedback reports how many records were added and how many were skipped as duplicates. Always review unexpected counts before screening.