How Ethiopia's elite distance runners actually train: an AI-assisted dissection of a multidimensional training structure.

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background: Elite East African distance runners have dominated their events for decades while training in ways that diverge sharply from Western sports-science orthodoxy. Why they are so good has never been settled with certainty. Genetics, lifelong altitude residency, and economic motivation have all been offered without a conclusive answer, partly because the day-to-day texture of the training itself has been hard to observe directly. Consumer GPS watches, now routinely used by elite athletes, leave detailed records of training that can be reconstructed at high resolution. This study takes the detailed footprint left in one cohort's watch data and dissects it for the factors behind their performance, using generative AI to make a decomposition of this depth tractable. The pattern that finally made sense of the data came from human expertise, echoing the very way these runners are coached. Methods: We followed 14 elite Ethiopian distance runners across 97 consecutive days in late 2025, drawing on 22,605 GPS segment records from their training watches together with venue and athlete information gathered in the field. The analysis was carried out through a three-phase human-AI workflow. An autonomous AI tool first searched the data across five investigator-seeded questions. A second AI system then developed selected findings into numerical claims, code, figures, and draft text under direct human guidance. A third, independent AI system stress-tested the methods, statistics, figures, and citations. Across all phases, the investigators set the direction, judged what the data could support, applied corrections and supplied the interpretation. Results: The dissection reconstructs how the cohort trains. Rather than working along a single intensity axis, the athletes spread their training across venues that each combine a particular altitude, surface, and terrain, with long aerobic running on the gentler mid-altitude roads, varied trail running at the high-altitude sites where the route changes from session to session, structured repetitions on the track, and easy recovery spread across the rest. Intensity is set by effort rather than by fixed pace or heart-rate targets, so altitude, surface, and intensity move together and cannot be cleanly separated by statistics. Essentially all the key sessions are morning group runs, and within the long ones the athletes settle into fitness-matched pace bands, so the group itself does the individualisation that Western coaching delivers through prescribed individual zones. Measured against the standard half-marathon and marathon pace scale, the intensity distribution is dominated by easy running with only a small moderate share and leans polarised once altitude is accounted for, matching the training logs of world-class runners recorded at altitude. Taken alone, it looks unremarkable, and that single-axis summary is a projection that hides how the work is placed, which only the multidimensional view recovers. All of it condenses into a compact map with five axes, altitude, surface, intensity, session structure, and session duration, drawn as three lattices on a shared altitude-and-surface grid, which the coaches arrive at by experience and the cohort concentrates into a handful of combinations. Conclusions: In this cohort, the GPS-watch data revealed a multidimensional training pattern in which intensity, altitude, surface, terrain, and group execution varied together rather than falling along a single intensity axis. The multidimensional framework developed to describe the data should be understood as a descriptive model of the regularities visible in that record, with no implication that the coaches used a formal scheme. Its value is that it makes part of an experiential coaching practice measurable. The coaches' tacit knowledge shaped the training, and the training left a structured footprint in the data. The AI-assisted workflow helped expose and verify that footprint, but the organising interpretation depended on human domain expertise. The findings therefore describe how this cohort trained in this setting, and any application beyond it would require judgement grounded in the athletes, coaches, venues, and constraints of the new context. Trial registration: Not applicable.

Article activity feed