AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification—requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners.
Objective
To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it.
Methods
In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss’ κ, Krippendorff’s α, Gwet’s AC1, and ICC(2,1)/ICC(2,k).
Results
Human item agreement was moderate (κ 0.409–0.467), as was single-rater reliability (ICC(2,1) 0.539–0.694), against good panel reliability (ICC(2,k) 0.854–0.919). Recorded pass/fail agreement was fair (κ 0.297–0.388) and critical-failure agreement near zero for two skills (κ 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), −0.173 (BVM), 0.664 (medical), and it was most lenient on two skills.
Conclusions
Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration—a baseline, not a limit.