MAPEVA-ScR: a four-phase framework integrating transparent rule-based text mining, evidence gap mapping and a mandatory human validation gate for scoping reviews — framework development and empirical validation in occupational health

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Automation is increasingly used to reduce the screening workload of evidence syntheses, but its error is real, variable and rarely reported. The 2025 joint position statement of Cochrane, Campbell, JBI and the Collaboration for Environmental Evidence mandates human oversight of artificial intelligence in evidence synthesis, yet — as of today — no international guideline requires quantitative measurement and reporting of automated-screening error against a human gold standard (1).

Methods

We developed MAPEVA-ScR (MAPeo de EVidencia con VAlidación humana), a four-phase framework for scoping reviews: (A) datalake construction with deterministic, traceable deduplication; (B) transparent, record-level auditable rule-based text mining with confidence tiers; (C) evidence gap mapping that feeds back human screening prioritisation; and (D) a mandatory human validation gate verifying 100% of the candidate corpus and issuing a quantitative screening error report (per-tier precision, confusion matrix, false-positive families, documented changes to conclusions). The framework was empirically validated on a scoping review of epilepsy and occupational fitness (2015–2025; PubMed, Scopus, Web of Science): a funnel of 2,008 → 1,435 → 1,114 → 230 records was reduced to 127 human-confirmed inclusions (present study).

Results

The rule-based filter achieved a global precision of 55.2% (Tier A 75.5%, Tier B 37.5%, Tier C 45.8%); human validation identified 103 false positives classifiable into seven reproducible error families, and 107 of 127 included records required human correction of automated thematic labels. Without the gate, the review’s central thematic conclusion would have been published reversed: automated labelling ranked Aptitude/employability third (12.6% of records), whereas after validation it ranked first (73.2%) (present study).

Conclusions

A mandatory human validation gate with quantitative error metrics closes a reporting gap that current guidelines leave open: in the present study it prevented publication of an inverted synthesis. We propose the screening error report as a minimum reporting standard for any evidence synthesis using automated screening.

Registration and protocol

Not applicable (methods development and validation study reporting original empirical data; no health intervention outcomes).

AI use disclosure

Generative AI assisted with language editing under full human verification (2); the MAPEVA-ScR pipeline uses deterministic rule-based text mining, with no generative AI in screening decisions.

Article activity feed