A study by IIIT-H researchers warns doctors against relying entirely on AI for medical diagnosis. After testing four models on chest X-rays, the team found AI heatmaps often miss what radiologists see. Syed Faizan said, “We essentially wanted to examine whether the heatmaps actually correspond to where radiologists say.”
Doctors should exercise caution when using artificial intelligence (AI) tools to interpret medical images, as areas highlighted by AI models may not always align with those identified by radiologists, according to a study by researchers at the International Institute of Information Technology, Hyderabad (IIIT-H).
The study, conducted by the Language Technologies Research Centre (LTRC) at IIIT-H, examined four vision-language models for analysing chest X-rays and found discrepancies between the regions highlighted by the models and those identified by radiologists.
The findings underscore the need for doctors to independently assess AI-generated outputs rather than relying on them entirely for diagnosis.
The research team, led by Parameswari Krishnamurthy, sought to determine whether AI-generated heatmaps accurately represented the regions of an image that radiologists would identify as disease-related.
However, the researchers questioned whether these visual indicators necessarily reflected the models’ ability to identify the actual location of a disease.
“We essentially wanted to examine whether the heatmaps created by vision-language models actually correspond to where radiologists, who look at the image, would say the disease lies,” said Syed Faizan, principal investigator of the study titled “How Well Do Chest X-Ray VLM Attention Overlays Match Radiologist Boxes? A Cross-Model Audit and Radiologist Reader.”
The team evaluated four models, MAIRA-2, MedGemma-4B, LLaVA-Med-1.5, and LLaVA-1.5, against thousands of publicly available chest X-rays. Two radiologists also participated in the study to assess them, allowing the researchers to compare AI-generated highlights with human assessments.
The study, accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2026, is to be presented at the iMIMIC Satellite Event in Strasbourg.
The researchers found that an AI model highlighting an apparently correct region of an X-ray did not necessarily mean that it had identified the disease in the same way as a radiologist.
Heatmaps based on diagnosis
According to Dr. Faizan, a model could arrive at a diagnosis first and then use that conclusion to determine where to place its heatmap, rather than independently identifying the affected region from the image.
To test this possibility, the team removed diagnostic information and examined how the models localised abnormalities. Their performance declined, suggesting that anatomical expectations associated with a diagnosis could influence the regions highlighted by the models.
The study also found a difference in the performance of these different models in the technical audit and radiologists’ assessments.
Dr. Faizan said this difference highlighted the importance of assessing AI tools from a clinical perspective. A model that focuses closely on a suspected abnormality may not provide all the information a radiologist needs, including how far a disease has spread to surrounding areas.
The gap that was addressed
The researchers said their work addressed a gap in the evaluation of medical AI systems, as earlier studies had largely focused on prediction and diagnostic accuracy rather than on whether AI-highlighted regions matched those identified by radiologists.
Study on paraphrase robustness
The LTRC is also investigating another potential weakness in vision-language models — their sensitivity to the wording of medical questions.
Doctors may phrase the same clinical query differently, for example, asking whether a chest X-ray shows pneumonia or whether pneumonia can be ruled out. They may also use technical terminology or more familiar expressions.
The researchers examined whether AI models responded consistently to medical questions worded differently in a study on paraphrase robustness.
The study explored how changes in the phrasing of medical questions affected the responses generated by vision-language models, raising questions about whether the systems genuinely understood the clinical query or depended excessively on the precise wording used.
Ms. Parameswari said the laboratory’s broader objective was to apply Natural Language Processing (NLP) to healthcare and use large language models and vision-language models to assist clinicians with tasks such as documentation, report writing and patient-doctor communication.
The aim was to reduce the time doctors spent on routine administrative work and enable them to devote more attention to clinical decision-making.
The researchers emphasised that the objective of medical AI research should be to support clinicians, not replace them, with human judgement remaining central to healthcare.
