Bristol researchers: medicine already handles AI's black boxes

Share
Bristol researchers: medicine already handles AI's black boxes

Two Bristol philosophers argue the medical AI field is optimizing the wrong thing about black boxes. Pharma ran into the same wall decades ago and built a process that works without ever explaining the mechanism.

Emanuele Ratti and Lena Zuchowski, both at the University of Bristol, argue in a paper accepted by Studies in the History and Philosophy of Science that machine learning should stop treating explainability as the gate to trust and instead certify models the way medicine certifies drugs. Their starting point is a 2019 argument by Andrew London: ML shares three epistemic traits with pharmacology — associationism (it tracks correlations, not causes), atheoreticity (it learns from data with no coded theory) and opacity (no one can read the mechanism). London's point was that medicine accepts all three and still ships reliable treatments, because clinical translation substitutes process for explanation. Ratti and Zuchowski say that suggestion was cited for seven years and never actually built into a guide — so they build one.

Their move is to treat the medicine–AI parallel as what philosopher Mary Hesse called a generative analogy: not a loose comparison but a structure that transfers working machinery from one field to the other. What transfers is clinical translation itself, which London and Jonathan Kimmelman describe as assembling an "intervention ensemble" — dose, schedule, population, endpoints, contraindications — until you know the exact conditions inside which a molecule is therapeutically useful. A drug without its ensemble is not a therapy; it is a molecule.

The AI translation is what the paper calls a learning ensemble, with three dimensions. Boundaries of reliability: who is qualified to operate the system, on what hardware and software, trained on data from where, following which methodological steps. Performance: which metrics were chosen and why those. Functional: the intended use, evidence it is achievable, and how that use sits against domain knowledge and norms. The components come from reporting standards the medical field already runs — SPIRIT-AI, CONSORT-AI (both 2020) and TRIPOD+AI (2024) — which the authors repurpose as reliability evidence rather than paperwork.

That reframes the failure mode the field keeps hitting. The authors lean on the 2021 Nature Machine Intelligence finding that COVID-19 X-ray classifiers learned to read portable-scanner markers instead of lung pathology: the model scored well on paper and was useless in a real hospital, because nothing in the accuracy table described the context the model was built for. The same gap explains distribution shift and adversarial fragility. And the paper's sharpest claim is that explainability work is neither sufficient nor necessary here — transparency epistemologies try to solve opacity, while reliability accounting only has to manage it. A saliency map does not tell you whether your model will keep working in the next hospital.

Two honest limits. The framework is normative and illustrative — the authors give examples of components, not an exhaustive list, and they do not say who signs the certificate. It also separates ensemble-reliability (works in its original context) from use-reliability (works in a new one), which is the distinction that kills most clinical AI deployment claims. If you have ever had to tell a vendor benchmark from a real eval, the logic here is the same discipline, and our How to — tell a real benchmark from a marketing one walks the audit the other way round.

What to watch: whether medical AI reporting standards grow a functional-use dimension that regulators can point at — that would be the first place this framework turns into procurement language.

Should a hospital buy a model its own radiologists cannot explain, if the reliability record is documented and audited? Tell us in the comments.

Sources: arXiv 2608.18186 — What Can Artificial Intelligence Learn from Medicine? · The Decoder · DeGrave et al., Nature Machine Intelligence