AI in medical imaging: five conditions for trust
Strong performance on a dataset is not enough to make AI in medical imaging useful. Teams need to check its population, its role, its integration, its oversight and how it is monitored after deployment. Here are five conditions for moving from a technical demo to controlled clinical use.
By Rubens Valcy
Founder of MyTwin
Published on
Contents
- Condition 1 — Define the task and its role in the decision
- Condition 2 — Examine the data and the population
- Condition 3 — Test integration and human factors
- Condition 4 — Organize human oversight
- Condition 5 — Monitor performance after deployment
- A ten-question procurement checklist
- Frequently asked questions
- Sources
Artificial intelligence can help detect, segment, quantify, reconstruct an image or prioritize a worklist. These are different tasks. Their risk, their level of evidence and their impact on the radiologist cannot be assessed with a single number.
High sensitivity can come with many false positives. A good average can hide weaker performance in a subgroup, on another scanner or with a different acquisition protocol. An algorithm never works outside its context: evaluation must start from the intended use and the consequences of an error.
Before choosing
- Condition 1
Define the task and its role in the decision
The exam, the condition, the population, the user, the point in the care pathway and the output produced.
- Condition 2
Examine the data and the population
Validation on external data, with protocols and a population close to your own setting.
- Condition 1
At integration
- Condition 3
Test integration and human factors
With real users, inside the PACS and RIS, across different shifts and exam types.
- Condition 4
Organize human oversight
Who checks, at what point and following which procedure.
- Condition 3
After deployment
- Condition 5
Monitor performance
Versions, errors, incidents, performance by site and device, suspension criteria.
- Condition 5
Condition 1 — Define the task and its role in the decision
The specifications should state the exam, the condition, the population, the user, the point in the care pathway and the output produced. Does the tool draw attention to an area? Does it provide a measurement? Does it sort exams? Does it generate text?
Then describe what the professional does. Do they confirm every result? Can they dismiss the suggestion? How does the interface show uncertainty? Who acts when the AI and the human reading disagree?
This clarification determines the potential regulatory status. In the European Union, software intended for a medical purpose may fall under the Medical Device Regulation. The EU Artificial Intelligence Act adds obligations depending on the category and level of risk.
Condition 2 — Examine the data and the population
Ask where, when and how the development and validation data were obtained. Was performance measured on external data from other centers? Do the scanner vendors, protocols, prevalence levels and demographics resemble your own setting?
A retrospective validation enriched with positive cases can be useful to test a technical capability, but it does not necessarily reflect the real frequency of abnormalities or the interruptions of a clinical department. Predictive values change with prevalence.
Require results with confidence intervals and, where relevant, subgroup analyses. A subgroup that is too small cannot establish equivalence: the absence of a statistical difference is not proof of the absence of bias.
Condition 3 — Test integration and human factors
An accurate tool can still make work harder if it slows down the workstation, adds clicks or displays an ambiguous signal. Test it with real users inside the PACS and RIS, across different shifts and exam types.
Observe:
- the time added or saved;
- how well symbols and scores are understood;
- how often work is interrupted;
- automation bias, meaning the tendency to follow the suggestion without sufficient checking;
- the opposite risk, when repeated alerts end up being ignored;
- how the result is documented in the report.
Training should cover known errors and out-of-scope cases, not just successful examples.
Condition 4 — Organize human oversight
“Human in the loop” only means something if that human has the time, the information and the authority required. Define who checks, at what point and following which procedure. The professional must be able to access the original image and understand the nature of the output.
The World Health Organization (WHO) emphasizes autonomy, transparency, accountability, inclusiveness and sustainability in the use of AI for health. These principles translate into operations: appropriate information, logging, a way to contest results, incident reporting and governance of updates.
That is how MyTwin for clinicians integrates imaging. Alongside the medical record, test results and treatments, imaging is part of the data MyTwin uses to build the patient’s digital twin. Its imaging analyses inform the radiologist, who keeps the reading and the diagnosis. Preoperative 3D models, produced through specialized partner applications, follow the same logic: they help surgeons prepare the procedure without deciding for them.
Dr. Rahal’s story shows what is at stake. As a digestive surgeon, she discovered an unexpected vascular variation in the middle of an operation, which reinforced her conviction that 3D preparation based on imaging matters. The surgical act itself remains hers.
Condition 5 — Monitor performance after deployment
Practices, equipment and populations change. A software update can also alter how the tool behaves. The monitoring plan should include:
- product version and date of changes;
- availability and response time;
- rate of uninterpretable results;
- false positives, false negatives and disagreements reviewed on a sample;
- incidents and near misses;
- performance by site, device and relevant group;
- impact on turnaround times and workload;
- suspension criteria.
The TRIPOD+AI statement and the PROBAST+AI tool provide benchmarks for reporting and assessing prediction models. They do not replace regulatory assessment, but they help identify missing information and risks of bias.
A ten-question procurement checklist
- Exactly which decision or task does the product support?
- Which version was evaluated?
- On which population and which equipment?
- Is there external and prospective validation?
- Which cases are excluded?
- How is uncertainty displayed?
- Who remains responsible for the reading and the action taken?
- How are errors and incidents reported?
- What happens after an update?
- Which clinical or organizational outcomes will be measured locally?
Frequently asked questions
No. It indicates conformity with the applicable EU framework for the claimed use. It does not automatically mean the product outperforms every alternative or that it will improve outcomes in your organization.
No. You need to compare the task, the population, the thresholds, the calibration, the consequences of errors and the integration into the care pathway. Two similar AUCs can correspond to very different uses.
Obligations depend on the context and the system. Beyond the legal minimum, clear information about the tool’s role, its limits and human oversight supports trust. Have the information process validated by the relevant officers.
Activate the planned response: root-cause analysis, checks on data and versions, restriction or suspension if needed, informing stakeholders and reporting in line with applicable obligations.
Sources
- World Health Organization, June 28, 2021, “Ethics and governance of artificial intelligence for health”.
- European Commission, accessed September 16, 2026, “AI Act — Regulatory framework for artificial intelligence”.
- European Commission, accessed September 16, 2026, “Medical Devices — EUDAMED overview”.
- Collins GS et al., 2024, “TRIPOD+AI statement”, BMJ, 385:e078378.
- Moons KGM et al., 2025, “PROBAST+AI”, BMJ, 388:e082505.
This article is provided for information purposes only. It does not replace advice, diagnosis or treatment from a healthcare professional.
