AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification—requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners.

Objective

To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it.

Methods

In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss’ κ, Krippendorff’s α, Gwet’s AC1, and ICC(2,1)/ICC(2,k).

Results

Human item agreement was moderate (κ 0.409–0.467), as was single-rater reliability (ICC(2,1) 0.539–0.694), against good panel reliability (ICC(2,k) 0.854–0.919). Recorded pass/fail agreement was fair (κ 0.297–0.388) and critical-failure agreement near zero for two skills (κ 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), −0.173 (BVM), 0.664 (medical), and it was most lenient on two skills.

Conclusions

Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration—a baseline, not a limit.

Article activity feed