As medical generative AI presents profound risks to trust and fair delivery, experts suggest five practical approaches to better control these dangers.
Medical generative AI carries immense potential for healthcare, but it also introduces profound risks that could damage human trust and undermine efficient, fair care delivery.
Why General AI Safety Standards Fail Medical Domains
Handing medical oversight over to general AI safety scientists without involving established medical technology regulators simply will not work. Foundation models require a mix of horizontal rules and medical-specific regulations that account for the unique demands of healthcare. Regulators must follow the example adopted by the Australian medical device regulator by clearly defining the medical domain, stopping general devices from encroaching upon it, and enforcing rigorous product safety, risk management, and approval processes.
Furthermore, medical technology regulators in the United States and the United Kingdom have recently recognized that GenAI medical tools cannot be regulated in the same way as static, designed-for-purpose medical devices. While certain existing medical laws remain applicable, current regulatory frameworks urgently need adaptation to become fit for purpose.
Managing GenAI as an Integrated Health System
Medical devices powered by AI are not isolated, atomised tools; they operate as complex systems deeply interacting with broader healthcare environments. The UK national commission on Health AI recommends ensuring that implementation happens in partnership across boundaries. Human oversight systems must reach maturity before rollout, remain functional during use, and extend seamlessly between healthcare system deployers and multiple model providers.
If health systems lack mature oversight, or if clinician adaptation and learning curves cannot keep pace with technological changes, deployment must slow down or pause. Healthcare organizations need better resourcing to make these transitions more human-centric and human-sympathetic.
Redefining Human Oversight and Real Transparency
Human oversight is critical, encompassing humans in the loop, on the loop, and above the loop. However, humans are naturally prone to fatigue and miss rare events, while repetitive tasks with limited authority fail to engage true experts. Setting up clinicians as overseers of partially autonomous systems without proper training or technical support sets them up for failure.
To solve this, automated systems—such as large language models acting as judges to monitor risk—are required to support human oversight. Alongside this, real bidirectional transparency is mandatory. Rather than static compliance information, systems must feature built-in open feedback approaches that collect public, interrogatable information on problems and medical mishaps.
Developers must then be forced by regulators, backed by the government, to make changes where changes are needed. Unsafe aspects of tools cannot be allowed to accumulate.
Nature commentary authors
Without strict enforcement of these measures, regulation loses its teeth, undermining public faith and developer respect. As the commentary notes regarding regulatory integrity: If your neighbour does not have to follow the law, why should you? Enforce enforce enforce!