1Animal Genetics and Breeding Division, ICAR-National Dairy Research Institute, Karnal-132 001, Haryana, India.
*Corresponding Author: Sabyasachi Mukherjee, Animal Genetics and Breeding Division, ICAR-National Dairy Research Institute, Karnal-132 001, Haryana, India. Email: sabayasachimukherje@gmail.com
This study evaluates machine learning (ML) algorithms as a robust alternative to traditional linear models, which often overlook complex biological patterns for predicting lactation performance in Karan Fries (KF) cattle. Using phenotypic records from 473 lactating animals, the analysis focuses on predicting 305-day milk yield (305 DMY) and total milk yield (TMY). The population exhibited an average peak yield of 13.28±0.25 kg and a persistent average fat concentration of 4.14%.
This study was done by integrating Wood’s incomplete gamma function to model lactation curves, compressing longitudinal test-day data into mathematically significant parameters. These parameters, alongside fixed environmental factors, served as predictors for three models: Multiple linear regression (MLR), random forest (RF) and extreme gradient boosting (XGBoost).
Results revealed that the RF model significantly outperformed both MLR and XGBoost. For 305DMY, RF achieved a high coefficient of determination (R2 = 0.88) compared to the poor performance of the MLR baseline (R2 = 0.21). A similar hierarchy persisted for TMY, where RF attained an R2 of 0.87, effectively reducing the root mean square error (RMSE) by over 50% compared to MLR. The study concludes that integrating Wood’s parameters with ensemble-based ML models offers a transformative tool for dairy breeding by facilitating early selection and optimized herd management.
Karan fries cattle, Lactation curve, Multiple linear regression, Random forest, Wood’s model, XGBoost