Consumer AI chatbots missed skin diagnoses in up to 98% of photo cases
A new audit ran nine consumer AI assistants through 13,797 skin-disease questions. The photograph was the problem.
Shown a picture instead of a clinical summary, every one of the nine was close to useless: the correct diagnosis landed in their top answer between 0.4% and 7.0% of the time. The audit, published in Diagnostics by researchers at the Plastic Surgery Hospital of the Chinese Academy of Medical Sciences, took 511 published cases of cutaneous vascular tumours, malformations and related syndromes and put them to ChatGPT 4o, Claude 3.5 Sonnet, Gemini 2.0, Copilot, Wenxiaoyan, Kimi, Doubao, Tongyi and Baixiaoying, through three workflows — text only, text plus image, and image only. The split is clean. Given the clinical summary as text, the same systems named the right condition 32.9% to 53.0% of the time. Given only the image, the correct answer usually did not appear even in the differential list, with omission running from 71.0% to 97.8%.
The second finding is the one worth pausing on. Adding an image to the clinical summary did not reliably help — it lowered top-1 accuracy in seven of the nine systems and raised omission in eight, though the drop was not statistically significant for ChatGPT 4o or Claude 3.5 Sonnet. Whatever these models are doing with a photograph, it is not adding diagnostic signal in the direction people assume.
That matters because this is the consumer use case in its most literal form: a rash, a lesion, a phone camera, a chatbot. Several of the audited conditions are life-threatening, including kaposiform hemangioendothelioma with Kasabach-Merritt coagulopathy and Sturge-Weber syndrome. The expert raters scored text workflows consistently higher than image workflows for appropriateness, by 0.9 to 2.2 points, and for safety, by 0.5 to 0.8 points. The authors land on "preservation of clinical context and specialist oversight."
The caveats are real and the authors state them plainly. The design cannot estimate real-world diagnostic accuracy or isolate what the image itself contributes — the cases came from published reports, so memorised parameters or retrieval augmentation could not be ruled out. It is a single-centre audit in an open-access general-medicine journal, and the systems were queried through their web interfaces in early 2025, which makes the scores a snapshot rather than a standing verdict on the models shipping today.
Our take: the finding is not that these tools are dangerous in the abstract. It is that they degrade in a specific direction — one that flatters them. Text carries the clinical context a diagnosis needs; the image is where a model gets to look confident without being told anything. Anyone holding up a phone to a mole is using the one mode where this generation performs worst.
What to watch: whether the next crop of multimodal medical models shows the image actually improving on text. That is the test the current generation fails.
Have you ever sent a photo to a chatbot to ask what something was? Tell us in the comments.
Sources: Diagnostics (MDPI) · 生物通 bio360 · Vasc511 dataset (GitHub)