Data-Driven Guide

How Many Papers Should I Screen?

A practical way to estimate screening effort from your own search yield, reviewer workflow, and pilot speed.

By Lumina Editorial Team Published Updated Methodology guide
Deduplicated scientific database search results ready to import for screening
Your screening estimate should start with the actual number of unique records after search and deduplication.

The Short Answer

There is no defensible universal number. Start with the records remaining after deduplication, multiply by the number of independent reviewers, and use a timed pilot to estimate minutes per record. If you plan to stop before every record is assessed, define and validate that rule separately rather than assuming a fixed AI workload reduction.

N

Records after deduplication

Use your actual search yield

R

Independent reviewers

Include verification and conflicts

T

Pilot seconds per record

Measure, do not guess

Build a Review-Specific Estimate

01

Run and deduplicate the search

Count unique records after combining every planned database and supplementary source.

02

Time a calibration sample

Have the actual reviewers screen the same pilot set. Include reading, notes, and decision time.

03

Add workflow overhead

Budget separately for duplicate review, conflicts, full-text retrieval, and protocol-driven quality checks.

How Long Will Screening Take?

The table below is arithmetic, not a benchmark. Replace the example speeds with the median from your pilot. It shows one reviewer screening every record once.

Papers 30 sec/record 60 sec/record 120 sec/record
1,000 8.3 hours 16.7 hours 33.3 hours
3,000 25 hours 50 hours 100 hours
5,000 41.7 hours 83.3 hours 166.7 hours
10,000 83.3 hours 166.7 hours 333.3 hours

For dual independent screening, multiply screening time by two and add time for conflict resolution.

Factors That Affect Screening Volume

1

Search Sensitivity vs. Specificity

Broader searches find more relevant papers but also more noise. A sensitive search for a Cochrane review might return 10,000+ results vs. 2,000 for a focused search.

2

Number of Databases

Each additional database may add unique records as well as duplicates. Estimate this from the combined, deduplicated export rather than a fixed ratio.

3

Topic Popularity

Hot topics (COVID-19, AI, mental health) generate far more results than niche topics.

4

Inclusion Criteria Specificity

Very specific eligibility criteria lead to lower inclusion rates, meaning more papers to screen per included study.

How to Reduce Screening Workload

๐Ÿค– Use AI-assisted screening

Active learning can move likely-relevant records earlier in the queue. Any workload reduction requires a documented and validated stopping approach.

๐Ÿงช Pilot and refine criteria

Clear criteria reduce ambiguity and make reviewer calibration more meaningful.

๐Ÿ“Š Apply stopping strategies

Evidence-based stopping rules tell provide decision support when yield becomes low; they do not prove that no relevant records remain.

๐Ÿ” Refine your search strategy

Better search terms and filters reduce noise without losing relevant papers

Put the Most Promising Records First

Upload your records and let Lumina prioritize the queue while your reviewers retain control over every final decision.