HomeAnalysisMedical AI Diagnosis Faces a Trust Test in IIIT-H X-Ray Study

Medical AI Diagnosis Faces a Trust Test in IIIT-H X-Ray Study

A study by researchers at the International Institute of Information Technology, Hyderabad (IIIT-H), has raised a fundamental question about medical AI diagnosis: when an artificial intelligence model highlights part of a chest X-ray, is it identifying the disease in the same way a radiologist would? The research suggests that the answer cannot be assumed, even when the highlighted region appears to be correct.

The study, conducted by IIIT-H’s Language Technologies Research Centre, examined four vision-language models used to analyse chest X-rays. Its findings point to a gap between diagnostic output and visual explanation. An AI system may produce a plausible diagnosis while directing attention to a region that does not correspond with the area a radiologist identifies as disease-related.

That distinction matters because medical AI is increasingly being developed not only to provide answers but also to show users why it reached them. In clinical settings, a heatmap or attention overlay can appear to offer transparency. Yet the IIIT-H research indicates that a highlighted area may not necessarily demonstrate that the model has independently located the disease. It may instead reflect assumptions associated with the diagnosis the system has already reached.

The study was led by Parameswari Krishnamurthy, with Syed Faizan serving as principal investigator. It evaluated MAIRA-2, MedGemma-4B, LLaVA-Med-1.5 and LLaVA-1.5 against thousands of publicly available chest X-rays. Two radiologists also participated in the assessment, enabling the researchers to compare the areas identified by the models with human assessments.

The research was accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention 2026 and is scheduled to be presented at the iMIMIC Satellite Event in Strasbourg. The supplied study findings, as reported by The Hindu, do not suggest that the models should be discarded. Instead, they indicate that their outputs require independent clinical assessment and should not be treated as a substitute for radiologists.

## The difference between diagnosis and explanation

The central issue is not simply whether an AI model can produce a correct or useful diagnosis. It is whether the visual evidence it presents accurately represents the reasoning behind that diagnosis. These are related but different questions.

As Syed Faizan explained, the researchers wanted to examine whether heatmaps created by vision-language models corresponded to the areas that radiologists would identify as showing disease. A model might highlight an apparently relevant region, but that visual alignment alone does not establish that the model recognised the disease location through the same evidence a clinician would use.

The team tested this possibility by removing diagnostic information and examining how the models localised abnormalities. Their performance declined. The researchers said this suggested that anatomical expectations associated with a diagnosis could influence the regions highlighted by the models.

In practical terms, this creates a risk of mistaking an explanation for proof. If a system first arrives at a diagnosis and then places a heatmap in an expected anatomical location, the overlay may look convincing without fully revealing how the model reached its conclusion. The visual output can therefore appear more clinically meaningful than the underlying process has established.

This is particularly important in chest X-ray interpretation, where the relevant abnormality may not be confined to a single sharply defined area. Dr. Faizan said a model that focuses closely on a suspected abnormality may not provide all the information a radiologist needs, including the extent to which a disease has spread to surrounding areas.

## Why technical accuracy is not enough

The study also found differences between the performance of the models in the technical audit and the assessments made by radiologists. The researchers said this demonstrated the importance of evaluating medical AI from a clinical perspective rather than relying only on technical measures.

The distinction exposes a broader problem in healthcare technology. A model can perform well on a defined technical task and still provide an output that is incomplete or difficult to use safely in a clinical workflow. The question is not only whether the system detects a pattern, but whether its output gives a doctor the information needed to assess the patient.

Earlier studies, according to the researchers, had largely focused on prediction and diagnostic accuracy. The IIIT-H work addresses another part of the evaluation: whether the regions highlighted by AI match those identified by radiologists. This shifts the focus from the result alone to the relationship between the result, the evidence shown to the clinician and the limits of the explanation.

For hospitals and doctors, that distinction has institutional consequences. If an AI tool is used in diagnosis, responsibility cannot be transferred to the software merely because it produces a confidence-building visual overlay. The findings support a workflow in which clinicians independently assess AI-generated outputs and treat them as assistance rather than final judgement.

## The language problem in medical AI

The IIIT-H laboratory is also investigating another weakness in vision-language models: their sensitivity to the way a medical question is phrased. Doctors may ask whether a chest X-ray shows pneumonia or whether pneumonia can be ruled out. They may use technical terminology or more familiar expressions to describe the same clinical concern.

The researchers examined whether changes in wording affected the responses generated by these models in a study on paraphrase robustness. The underlying question is whether a system genuinely understands the clinical query or depends too heavily on the precise language used to frame it.

This issue complements the findings on heatmaps. In both cases, the concern is about whether an AI system is responding to the clinically relevant evidence or relying on indirect signals. A model that changes its response when a question is rephrased, or that highlights a region without reliably reflecting the disease location, may require more careful supervision than its apparent fluency suggests.

For clinical users, this means that the interface cannot be separated from the model’s reliability. The wording of the prompt, the format of the image and the way the result is displayed may all influence how the output is understood. The study material does not establish how these systems would perform across hospitals or patient populations, but it does identify areas that require scrutiny before their outputs are treated as dependable clinical evidence.

## Assistance, not replacement

The laboratory’s broader objective is to apply natural language processing and large language models to healthcare tasks such as documentation, report writing and patient-doctor communication. These applications are intended to reduce the time doctors spend on routine administrative work and allow more attention for clinical decision-making.

That distinction is important because the study does not present medical AI as either an automatic solution or an inherently unusable technology. It draws a boundary between assistance and replacement. AI may support clinicians with documentation, communication and image analysis, but the findings show why human judgement remains central when the system’s explanation does not reliably correspond to clinical interpretation.

The governance challenge is therefore not limited to selecting the most accurate model. Institutions also need to understand what a model highlights, how stable its responses are when questions are rephrased and whether its outputs provide the full information clinicians require. The study’s emphasis on radiologist comparison offers one way of testing that relationship.

The findings also show why public discussion of healthcare AI must move beyond headline accuracy claims. A model’s usefulness depends on the quality of the evidence it presents, the consistency of its responses and the ability of clinicians to recognise when its output is incomplete. Those questions are directly connected to patient safety, professional accountability and the design of hospital workflows.

The IIIT-H study confirms that medical AI diagnosis needs evaluation at more than one level. Prediction, localisation, explanation and clinical usability are not interchangeable measures. The research finds that AI-generated heatmaps may not align with radiologists’ assessments and that model performance can decline when diagnostic information is removed. It also raises questions about the effect of wording on medical AI responses. The next stage is the study’s presentation at the iMIMIC Satellite Event in Strasbourg, while the broader issue remains how healthcare institutions will keep human assessment central as AI tools move closer to clinical practice.


RELATED ARTICLES

Most Popular

Latest News