Large language models show a mean sensitivity of 0.95 and an overall accuracy of 0.78 in answering pharmacology questions, though their mean specificity remains at 0.43, according to a quantitative cross-sectional study conducted in Iran in 2026.
Evaluating AI Across Fifty Persian Multiple-Choice Questions
The study administered a 50-item Persian multiple-choice questionnaire assessing pharmacology knowledge to three distinct AI platforms: ChatGPT-5.5, Gemini 3.1, and Claude 4.6. Researchers analyzed the generated responses using confusion matrix methods, Cochran’s Q, McNemar, and point-biserial tests to measure baseline capabilities in medical education.
Performance metrics varied significantly across the evaluated models. ChatGPT-5.5 achieved a sensitivity value of 1.00, a specificity of 0.43, and an overall accuracy of 0.82. Gemini 3.1 registered a sensitivity of 0.94, a specificity of 0.50, and an accuracy of 0.82. Claude 4.6 recorded a sensitivity of 0.90, a specificity of 0.35, and an accuracy of 0.72.
Performance Declines When Tackling Quantitative Pharmacology
Beyond baseline metrics, the statistical analysis uncovered a distinct drop in capability when the systems faced computational challenges. Both ChatGPT-5.5 and Gemini 3.1 experienced a statistically significant decline in performance when addressing quantitative questions, indicated by a phi coefficient of minus 0.286 at a p-value of less than .05.
Future Directions in Medical Education and Knowledge Support
The findings indicate that these large language models hold tangible potential to support professionals and students within educational and knowledge-based contexts of pharmacology. However, deploying these tools safely requires addressing critical ethical considerations and guarding against hallucinated or inaccurate outputs, particularly given the low specificity scores observed across all tested systems.
Researchers concluded that further investigation is essential to fully understand the performance boundaries and systemic limitations of these technologies before integrating them into high-stakes clinical or educational workflows.