Agricultural Reviews
  • Year: 2026
  • Volume: 47
  • Issue: 4

Systematic Evaluation of ChatGPT-3.5 and ChatGPT-4.0 for Accuracy and Reliability in Veterinary and Animal Science Education

  • Author:
  • Priya Dhattarwal1,*, Vivek Sahu2, Yashwant Singh1
  • Total Page Count: 8
  • Page Number: 618 to 625

1Department of Livestock Production Management, College of Veterinary Science, Rampura Phul, Guru Angad Dev Veterinary and Animal Sciences University, Bathinda-151 103, Punjab, India

2Department of Livestock Products Technology, Guru Angad Dev Veterinary and Animal Science University, Ludhiana-141 004, Punjab, India

*Corresponding Author: Priya Dhattarwal, Department of Livestock Production Management, College of Veterinary Science, Rampura Phul, Guru Angad Dev Veterinary and Animal Sciences University, Bathinda-151 103, Punjab, India, Email: dhattarwalpriya@gmail.com

Abstract

Large Language Models (LLMs) such as ChatGPT are increasingly used in veterinary and animal sciences education, yet systematic performance evaluations across academic domains remain limited. This study quantitatively compared ChatGPT-3.5 and ChatGPT-4.0 in terms of answer accuracy across Animal Sciences, Paraclinical Sciences and Clinical Sciences. A dataset comprising domain-specific questions was presented to both models and responses were scored on a 10-point accuracy scale by subject matter experts. Statistical analyses including mean scores, standard deviations, paired t-tests and one-way and two-way ANOVAs were applied to assess performance differences. Results showed that ChatGPT-4.0 consistently outperformed ChatGPT-3.5 across all domains, with mean differences of +2.0 in Animal Sciences, +2.0 in Paraclinical Sciences and +1.9 in Clinical Sciences, all statistically significant (p<0.001). Moreover, ChatGPT-4.0 exhibited lower variability, indicating more stable accuracy. The overall two-way ANOVA revealed significant main effects for both model version (p<0.001) and academic domain (p = 0.002), with no significant interaction effect. These findings demonstrate that ChatGPT-4.0 provides superior and more consistent performance in veterinary education question-answering compared to ChatGPT-3.5, supporting its adoption as a supplementary learning tool in domain-specific curricula.

Keywords

Ai, Animal science, ChatGPT-3.5, ChatGPT-4.0, Clinical science, Paraclinical science, Veterinary