How Many Papers Should I Screen?
A practical way to estimate screening effort from your own search yield, reviewer workflow, and pilot speed.
The Short Answer
There is no defensible universal number. Start with the records remaining after deduplication, multiply by the number of independent reviewers, and use a timed pilot to estimate minutes per record. If you plan to stop before every record is assessed, define and validate that rule separately rather than assuming a fixed AI workload reduction.
Records after deduplication
Use your actual search yield
Independent reviewers
Include verification and conflicts
Pilot seconds per record
Measure, do not guess
Build a Review-Specific Estimate
Run and deduplicate the search
Count unique records after combining every planned database and supplementary source.
Time a calibration sample
Have the actual reviewers screen the same pilot set. Include reading, notes, and decision time.
Add workflow overhead
Budget separately for duplicate review, conflicts, full-text retrieval, and protocol-driven quality checks.
How Long Will Screening Take?
The table below is arithmetic, not a benchmark. Replace the example speeds with the median from your pilot. It shows one reviewer screening every record once.
| Papers | 30 sec/record | 60 sec/record | 120 sec/record |
|---|---|---|---|
| 1,000 | 8.3 hours | 16.7 hours | 33.3 hours |
| 3,000 | 25 hours | 50 hours | 100 hours |
| 5,000 | 41.7 hours | 83.3 hours | 166.7 hours |
| 10,000 | 83.3 hours | 166.7 hours | 333.3 hours |
For dual independent screening, multiply screening time by two and add time for conflict resolution.
Factors That Affect Screening Volume
Search Sensitivity vs. Specificity
Broader searches find more relevant papers but also more noise. A sensitive search for a Cochrane review might return 10,000+ results vs. 2,000 for a focused search.
Number of Databases
Each additional database may add unique records as well as duplicates. Estimate this from the combined, deduplicated export rather than a fixed ratio.
Topic Popularity
Hot topics (COVID-19, AI, mental health) generate far more results than niche topics.
Inclusion Criteria Specificity
Very specific eligibility criteria lead to lower inclusion rates, meaning more papers to screen per included study.
How to Reduce Screening Workload
Active learning can move likely-relevant records earlier in the queue. Any workload reduction requires a documented and validated stopping approach.
Clear criteria reduce ambiguity and make reviewer calibration more meaningful.
Evidence-based stopping rules tell provide decision support when yield becomes low; they do not prove that no relevant records remain.
Better search terms and filters reduce noise without losing relevant papers
Continue Learning
Put the Most Promising Records First
Upload your records and let Lumina prioritize the queue while your reviewers retain control over every final decision.