top of page
Screenshot 2025-03-14 091724.png

AI Can Help Detect Disease. But How Do We Know When to Trust AI in Healthcare?

Writer: THEMIS 5.0
THEMIS 5.0
Aug 7
4 min read

Artificial intelligence is rapidly becoming part of everyday clinical practice. Across Europe, AI is now supporting radiologists in detecting lung nodules, identifying breast cancer, assessing stroke patients and prioritising imaging studies. These technologies promise earlier diagnosis, improved efficiency and reduced workload at a time when healthcare systems face increasing demand and workforce shortages. Yet despite remarkable advances in performance, one question continues to dominate discussions among clinicians:


When should we trust an AI recommendation?


Graphic - a radiologist sits at a workstation in a modern hospital, reviewing CT and X-ray images displayed on two computer monitors while other healthcare professionals consult in the background. The image represents the role of trustworthy AI in supporting clinical decision-making, with the clinician remaining at the centre of patient care.
A clinician reviews diagnostic imaging alongside AI-generated insights

This is not simply a technical question. It is a clinical one. For a radiologist reviewing a CT scan or a clinician interpreting an AI-generated recommendation, the decision is rarely based on accuracy alone. Clinical judgement depends on understanding why a conclusion has been reached, how confident the system is, whether it has been validated on similar patients, and what the consequences might be if it is wrong. These are precisely the challenges that our Horizon Europe THEMIS 5.0 (THEMIS) project is addressing.


Accuracy Is Only the Beginning

Artificial intelligence has achieved impressive results in medical imaging. Numerous studies have demonstrated that AI can detect abnormalities with levels of accuracy approaching those of experienced clinicians in specific diagnostic tasks. However, clinical adoption has progressed more slowly than technological development.

Why? Because healthcare professionals are not simply looking for accurate algorithms. They need systems that are safe, transparent, explainable and appropriate for the context in which they will be used.


A recent review in the European Journal of Radiology Artificial Intelligence highlights that explainability has become central to trustworthy AI in radiology, not because clinicians are resistant to innovation, but because they need evidence that AI recommendations are robust, clinically meaningful and free from hidden biases. The review also notes that apparently convincing visual explanations can themselves be misleading if they are not stable or clinically valid. Trust, therefore, is not created by high accuracy alone, it is built through evidence.


The Reality of Clinical Decision-Making

Imagine a radiologist reviewing a CT scan. The AI highlights a suspicious region and recommends further investigation. The clinician immediately begins asking questions.


Why has the AI identified this area?

How confident is the prediction?

Has this model been validated on patients similar to mine?

Could this represent a false positive?

Would I have reached the same conclusion without AI?


These are not hypothetical concerns. A multicentre study published in Radiology found that the way AI explanations are presented significantly influences physician trust and diagnostic performance. Different explanation styles affected how clinicians interpreted AI advice and how much they relied upon it, demonstrating that explainability is not simply a usability feature. it directly influences clinical decision-making.


Equally important is avoiding automation bias. A study article published in Springer's European Radiology showed that when AI intentionally provided incorrect findings, radiologists were more likely to make incorrect follow-up decisions than when interpreting the same images without AI assistance. The research reinforces the need for AI systems that support, not replace, clinical reasoning.


From AI Performance to AI Trustworthiness

This shift, from measuring AI performance to assessing AI trustworthiness, is the foundation of the THEMIS 5.0 healthcare pilot. Rather than asking whether an AI model achieves high diagnostic accuracy, THEMIS asks a broader question:


Is this AI system trustworthy enough to support clinical practice?


To answer that question, the project combines technical assessment with human-centred evaluation. Within the THEMIS platform, AI systems are assessed across multiple trustworthiness characteristics, including accuracy, fairness, robustness, explainability and risk. These assessments are complemented by user-centred methods that recognise different healthcare professionals may require different evidence before deciding whether to rely on an AI system.


A radiologist, for example, may prioritise diagnostic confidence and explainability.

A hospital manager may be more concerned with governance, regulatory compliance and operational reliability. An AI developer may need detailed technical evidence to improve model performance before deployment. THEMIS recognises that trustworthiness is not one-size-fits-all.


Building Trust Through Co-Creation

Perhaps the most distinctive aspect of THEMIS is that these questions were not answered by computer scientists alone. From the outset, the project adopted a co-creation approach, working directly with clinicians, researchers and other stakeholders through participatory workshops and Living Labs to understand what trustworthy AI means in real clinical settings.


Rather than asking healthcare professionals to adapt to technology, the consortium designed its methodologies around the realities of clinical practice. This produced some unexpected insights. For example, during pilot activities, participants reacted negatively to terminology relating to "morals." Although the assessment was never intended to judge whether a person was moral or immoral, the language itself created discomfort. The methodology was refined as a result, demonstrating that trust is influenced not only by algorithms, but also by communication, perception and user experience. These lessons reinforce a simple but powerful idea: people must help shape trustworthy AI if they are expected to trust it.


Keeping Clinicians in Control

One of the guiding principles of THEMIS is that AI should support human expertise, not replace it. The platform does not tell clinicians what decision to make. Instead, it provides evidence to support clinical judgement. It helps users understand where an AI system performs well, where uncertainty exists, what risks remain and whether further improvements may be needed before deployment. This reflects the growing consensus across Europe that trustworthy AI requires more than regulatory compliance. It requires systems that clinicians can understand, question and confidently integrate into their own reasoning.


Trust AI in Healthcare

Artificial intelligence will undoubtedly become an increasingly important part of modern medicine. But successful adoption will depend not only on developing better algorithms, but also on building greater confidence among the professionals who use them.


The THEMIS 5.0 healthcare pilot demonstrates that trust is not achieved through technical performance alone. It is built through transparency, evidence, and co-creation.

And through recognising that every clinical decision ultimately belongs to the healthcare professional, not the AI system. Because the future of healthcare is not about replacing clinicians with artificial intelligence. It is about giving clinicians better evidence to make the best possible decisions for their patients.






trust ai in healthcare

Comments


bottom of page