Code Sharing in Healthcare Prediction Model Research: A Scoping Review

A scoping review reveals that a significant proportion of healthcare research papers citing the TRIPOD or TRIPOD+AI reporting standards fail to provide accessible analytical code, raising urgent questions about computational reproducibility in clinical prediction models.

As multivariable prediction models increasingly shape diagnostic and prognostic patient care, independent verification of underlying algorithms remains a critical bottleneck. A comprehensive scoping review utilizing advanced automated pipelines examined PubMed-indexed literature citing the TRIPOD and TRIPOD+AI reporting guidelines up to 11 August 2025 to measure true analytical code availability and structural quality.

  • Multivariable Prediction Models: Mathematical equations combining two or more clinical predictors to estimate an individual’s probability of a current health condition or future medical outcome.
  • Analytical Code Transparency: The practice of publishing the underlying statistical or machine-learning code alongside a research paper, allowing independent researchers to audit and replicate the findings.
  • Automated Repository Auditing: The deployment of specialized large-language-model pipelines to scan for executable source files, license agreements, and software dependency files.

Automating Scoping Reviews for Code Transparency

To characterize code-sharing practices across modern medical literature, researchers designed a robust evaluation pipeline targeting all PubMed-indexed primary research articles citing the TRIPOD or TRIPOD+AI guidelines as of 11 August 2025. Restricting inclusion to articles retrievable through the PubMed Central Open Access API ensured full methodological reproducibility without subscription barriers. The investigation registered its protocol under identifier INPLASY202620080 in accordance with the PRISMA extension for scoping reviews.

Data analysis proceeded in two distinct phases: article-level screening and deep repository characterization. An automated pipeline driven by GPT-5.2 (2025-12-11) extracted metadata, determined eligibility, and identified code repositories. The automated system was validated against human-annotated subsets.

Structural Characteristics and Repository Documentation

Once identified, repositories were ingested using a custom retrieval utility that compiled file trees, README documents, and source scripts into structured text files. A secondary LLM assessment evaluated 14 specific features defined by the TRIPOD-Code executive committee, including software dependencies, licensing information, test frameworks, and random seed control for stochastic processes. Human validation across 35 retrieved repositories confirmed a weighted F1 score of 0.83.

Evaluation Metric Target Feature Model Performance (F1 Score)
Repository Audit Overall 14-Feature Evaluation 0.83 Weighted F1 Score

Future Trajectory of Clinical AI Governance

References

  • PRISMA Extension for Scoping Reviews: Reporting guidance for systematic literature reviews.
  • INPLASY Protocol Registry: Registration record INPLASY202620080.
  • TRIPOD and TRIPOD+AI Statement Guidelines: Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis.
Prediction models in healthcare: a playground for researchers | Webinar
Photo of author

Priya Deshmukh - Senior Editor, Health

Priya Deshmukh Senior Editor, Health Deshmukh is a practicing physician and renowned medical journalist, honored for her investigative reporting on public health. She is dedicated to delivering accurate, evidence-based coverage on health, wellness, and medical innovations.

Sony physical game sales remain high despite digital push