Benchmarking Machine Learning Classification of Hanwoo Intramuscular Fat Grade Using Routine Carcass Records

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Accurate classification of intramuscular fat grade in Hanwoo cattle is economically important because beef quality grade strongly influences carcass value. However, most existing machine learning approaches for this trait rely on genomic markers, ultrasound measurements, or image-derived features that require dedicated acquisition systems unavailable in routine slaughter operations. Moreover, small sample sizes and strong class imbalance in livestock carcass datasets can make conventional single-split evaluation unstable, highlighting the need for repeated and leakage-controlled evaluation protocols. This study evaluated whether routinely collected carcass record variables alone can classify a four-class marbling-score grade using machine learning, thereby defining a practical baseline for routine-record-based classification. A total of 386 Hanwoo carcass records were analyzed using seven predictors: age at slaughter in months, meat color score, fat color score, maturity, quantity grade, aging days, and sex. Eleven baseline classifiers and five oversampling strategies were jointly assessed under repeated stratified cross-validation. Gradient boosting was retained as the main baseline because it achieved the highest mean accuracy of 0.579, tied for the highest weighted F1 score of 0.558, and remained competitive in Cohen's κ at 0.316. Adaptive synthetic sampling combined with extreme gradient boosting produced the highest κ value of 0.323, but it did not improve accuracy or weighted F1 relative to the baseline model. Class-wise comparison showed that the augmented model improved the high and premium classes, whereas the baseline model better preserved performance in the middle class and maintained stronger aggregate performance. Cross-validated permutation importance identified sex, age at slaughter in months, meat color score, maturity, and aging days as the largest contributors, with quantity grade showing a smaller positive contribution. These findings show that routine carcass records contain moderate predictive capacity for marbling-score grade, but performance remains constrained for minority high-grade classes. In addition, repeated cross-validation with fold-confined oversampling offers a reproducible evaluation setting for small, imbalanced livestock datasets. The resulting benchmark provides a reference point for future models that incorporate richer pre-slaughter, genomic, imaging, or multimodal information.

Article activity feed