Screening methodology
When to Stop Screening in a Systematic Review
Stopping rules can identify diminishing yield in an AI-prioritized queue. They cannot prove that every relevant record has been found, so the rule, validation, and acceptable risk belong in the protocol.
Start with the review objective
Prioritization changes the order in which records are seen. Early stopping changes which records are assessed at all. That second decision needs stronger justification because a relevant record may still appear late in the queue.
Screening every record remains the clearest option when the protocol requires exhaustive assessment, when missing a study could materially change a conclusion, or when the ranking has not been validated on comparable data.
Stopping signals and the SAFE workflow
Consecutive irrelevant records
A long sequence of exclusions may indicate declining yield in a ranked queue.
Limit: the result depends on ranking quality and the chosen streak length. A streak does not estimate how many relevant records remain unseen.
Knee detection
The method looks for a bend in the cumulative discovery curve where new inclusions become less frequent.
Limit: curves can be noisy, especially when relevant records are rare. The detected knee is a mathematical feature, not a recall estimate.
Slope or yield threshold
This signal monitors the inclusion rate over a recent window and alerts when observed yield drops.
Limit: window size and threshold can materially change the alert. A low recent yield can still be followed by a late relevant record.
Four-phase workflow
The SAFE stopping procedure
SAFE combines calibration, active learning, model switching and quality review. It is a practical heuristic proposed by Boetje and van de Schoot (2024); it does not guarantee a particular recall.
S — Screen a random sample
Draw a random calibration sample to obtain training labels and a rough relevance estimate. Keep known relevant key records as validation checks.
A — Apply active learning
Combine key-record retrieval, minimum screening volume and an exclusion streak before progressing.
F — Find more relevant records
Use a different model and representation to rank the remaining records, then continue screening.
E — Evaluate quality
Independently reassess exclusions and check inclusion quality. Consider citation searching and document the final decision.
How Lumina implements SAFE
Select SAFE Procedure (Guided) in your project's Early Stopping settings. Enter known relevant DOIs or record IDs and an exclusion-streak length (default 50). Record IDs appear in Results.
Lumina draws 1% of the imported records, with a minimum of two. Complete that sample; extend it randomly if it contains only one label class. Phase A requires all key records to be finally included, at least 10% of records screened, at least twice the sample-based estimate of relevant records screened, and the specified exclusion streak. Those figures are configurable protocol choices where offered, not universal safety thresholds.
Phase F uses a separate character TF-IDF + logistic-regression model and counts a fresh exclusion streak. Phase E presents up to one streak-length of excluded records, ranked highest by the second model, for an independent reviewer who did not previously screen each article. This ranked audit is a Lumina adaptation, not a statistical random-sample recall test. Teams using dual screening may need an additional quality reviewer.
Unresolved decisions block advancement. A flagged exclusion blocks completion. Changes to the record set or decisions invalidate completed checks; restart calibration after reassessment. The owner records inclusion-quality and citation-search notes before completing the workflow. Screening pauses at phase checkpoints; switching to Manual allows continued full screening.
The random sample, settings, phase actions, reviewer identities and quality notes are retained with the project. The review team remains responsible for reporting deviations and unassessed records.
Later evidence supports careful validation: Verdonschot et al. (2026) found that effective model and label combinations differed across four medication-review datasets. A 2026 evaluation of a modified SAFE stop rule fell below 95% recall for three of ten broad-search datasets; it did not test the complete four-phase procedure. Repke et al. (2026) provide a wider comparison of stopping methods, including statistical alternatives.
A defensible decision process
-
01
Pre-specify the rule
Name the algorithm, parameters, minimum training data, and action the team will take when the signal appears.
-
02
Validate on relevant evidence
Use retrospective data, a known benchmark, or a protocol-defined validation sample that resembles the current review.
-
03
Check ranking behavior
Do not stop if inclusions remain frequent, ranking quality appears unstable, or reviewer criteria changed during screening.
-
04
Report the full decision
Report screened and unscreened counts, the stopping rule, parameters, validation, deviations, and any sensitivity analysis.
What to write in the methods section
Adapt this text to what your team actually did. Do not report an estimated recall unless you calculated it with a defensible reference standard.
Sources and further reading
See prioritized screening before configuring a rule
Explore how records are ranked and how human decisions remain visible in the workflow.