Evidence-Based AI Marking for Secondary Assessment: Alignment, Bias, and Confidence Analysis
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
This study investigates the use of a Large Language Model (LLM) as an independent secondary marker for senior secondary school assessments. Unlike many prior artificial intelligence (AI) marking studies that rely on artificially generated prompts or holistic scoring, the system evaluated here enforces evidence-based marking rules, criterion-level justification, and explicit linkage between awarded marks and quoted student evidence. The study analysed 437 student responses collected across 13 senior secondary subject areas. AI-generated total scores were compared with teacher-awarded scores using Normalised Mean Absolute Error (NMAE), Mean Absolute Error (MAE), exact match rate, and Pearson’s correlation coefficient (r), alongside analyses of directional bias, confidence behaviour, qualitative output quality, and operational efficiency. Results showed moderate overall alignment between AI and teacher total scores (NMAE = 13.24%, MAE = 2.96, exact match = 20.8%, r = 0.51), together with a clear tendency toward leniency by the AI under baseline conditions. Higher model-reported confidence was associated more strongly with lenient marking than with improved accuracy, suggesting that confidence scores reflect evidential clarity or decisiveness rather than calibrated reliability. The findings indicate that transparent and auditable AI marking systems can play a valuable role as consistency-oriented moderation tools and research instruments. While the results do not yet justify fully autonomous use in high-stakes assessment, the study contributes concrete empirical evidence relevant to the design, evaluation, and governance of AI-assisted assessment in senior secondary schooling, and illustrates a viable pathway for responsible, teacher-centred deployment.