Illustration credit to Suzan Mozak
Discourse around AI tends to oscillate between two extremes. Either dystopia is already here: jobs evaporating, students outsourcing thinking, and “clinical expertise” reduced to supervising machines. Or, AI is utopia just around the corner: personalised tutors, near-perfect diagnostic accuracy, and doctors so freed from paperwork that the balance between work and life tips in our favour. These polarised narratives are enticing because they offer certainty. They are also deeply misleading.
In my experience, much of the discomfort around AI reflects a very human impulse: the desire to resolve ambiguity too quickly. We often treat AI as a monolith, glossing over important distinctions between “narrow” systems (task-specific) and “general” tools; such as the over-anthropomorphised chatbots and their infamous hallucinations. AlphaFold, for example, represents a remarkable advance in protein structure prediction, but it is bounded and specialised. Conflating such systems with conversational AI amplifies both hope and fear without clarifying how these tools function in practice.
Examine AI in medicine more closely and we see messiness seeping through. In dermatology, it can outperform clinicians in melanoma diagnosis under test conditions, yet performance drops significantly in people of colour due to under-representation in training datasets. Radiology tells a similar story. A 2025 Lancet review found that AI tools can improve sensitivity and may streamline patient lists, but sometimes at the cost of reduced specificity, increasing downstream testing with system-wide consequences. This is neither magic, nor maleficence, but the predictable gap between what technologies promise and how they perform in complex clinical environments.
Two risks stand out. The first is bias: AI systems are trained on historical data and inherit its limitations and assumptions. The second is over-trust. Automation bias – the tendency to defer to algorithmic outputs, even when they’re wrong – is well-documented. A recent study from the Oxford Internet Institute warned that AI chatbots offering medical advice to the public can appear authoritative while failing to identify red flags or provide appropriate safety-netting, risking misplaced reassurance. The danger in our field is perhaps not that AI will replace clinicians, but outputs may not be interrogated properly, with machine-generated recommendations slowly displacing clinical judgement.
This is without even touching on environmental cost. Healthcare already accounts for a significant proportion of global greenhouse gas emissions, and the energy demands of training and running large AI models are substantial. Yet meaningful scrutiny is difficult: companies increasingly limit transparency about training data and model architecture, often citing proprietary intellectual property or national security concerns. Perhaps that is true. It is also convenient. It is also worth remembering that online search results are shaped by actors with their own institutional and commercial interests.
Looking to the year ahead, I suspect we will see less hyperbole and more engagement with the messy work of effective implementation. Recent legal action in the United States concerning the alleged covert use of AI transcription tools in clinical encounters offers a glimpse of the issues that emerge once technology moves from theory into practice. I am increasingly convinced that our response to AI should resist the urge for neat answers. As Roger Kneebone argued; framing the practise of medicine as purely scientific has always been “a dangerous oversimplification”. Clinical work depends on a shifting mix of evidence, judgement, and compassion, applied to individuals in imperfect circumstances. AI does not eliminate that uncertainty; if anything, it makes it more visible.
If clinical reasoning can eventually be done better by an AI, the more important questions are not whether AI belongs in medicine, but what does it mean to be a doctor? What kinds of knowledge should medical schools prioritise? And which parts of medicine should remain in our hands, not because machines cannot do them, but because we don’t want to give them up?
Author’s note: The words in this piece were – honestly – written by me. I’ve no shame to say I used ChatGPT to help me bounce around some ideas, as well as give some suggestions for refining of my phrasing, but on the latter I tended to disagree with it more often than not.
Illustration credit to Suzan Mozak
Discourse around AI tends to oscillate between two extremes. Either dystopia is already here: jobs evaporating, students outsourcing thinking, and “clinical expertise” reduced to supervising machines. Or, AI is utopia just around the corner: personalised tutors, near-perfect diagnostic accuracy, and doctors so freed from paperwork that the balance between work and life tips in our favour. These polarised narratives are enticing because they offer certainty. They are also deeply misleading.
In my experience, much of the discomfort around AI reflects a very human impulse: the desire to resolve ambiguity too quickly. We often treat AI as a monolith, glossing over important distinctions between “narrow” systems (task-specific) and “general” tools; such as the over-anthropomorphised chatbots and their infamous hallucinations. AlphaFold, for example, represents a remarkable advance in protein structure prediction, but it is bounded and specialised. Conflating such systems with conversational AI amplifies both hope and fear without clarifying how these tools function in practice.
Examine AI in medicine more closely and we see messiness seeping through. In dermatology, it can outperform clinicians in melanoma diagnosis under test conditions, yet performance drops significantly in people of colour due to under-representation in training datasets. Radiology tells a similar story. A 2025 Lancet review found that AI tools can improve sensitivity and may streamline patient lists, but sometimes at the cost of reduced specificity, increasing downstream testing with system-wide consequences. This is neither magic, nor maleficence, but the predictable gap between what technologies promise and how they perform in complex clinical environments.
Two risks stand out. The first is bias: AI systems are trained on historical data and inherit its limitations and assumptions. The second is over-trust. Automation bias – the tendency to defer to algorithmic outputs, even when they’re wrong – is well-documented. A recent study from the Oxford Internet Institute warned that AI chatbots offering medical advice to the public can appear authoritative while failing to identify red flags or provide appropriate safety-netting, risking misplaced reassurance. The danger in our field is perhaps not that AI will replace clinicians, but outputs may not be interrogated properly, with machine-generated recommendations slowly displacing clinical judgement.
This is without even touching on environmental cost. Healthcare already accounts for a significant proportion of global greenhouse gas emissions, and the energy demands of training and running large AI models are substantial. Yet meaningful scrutiny is difficult: companies increasingly limit transparency about training data and model architecture, often citing proprietary intellectual property or national security concerns. Perhaps that is true. It is also convenient. It is also worth remembering that online search results are shaped by actors with their own institutional and commercial interests.
Looking to the year ahead, I suspect we will see less hyperbole and more engagement with the messy work of effective implementation. Recent legal action in the United States concerning the alleged covert use of AI transcription tools in clinical encounters offers a glimpse of the issues that emerge once technology moves from theory into practice. I am increasingly convinced that our response to AI should resist the urge for neat answers. As Roger Kneebone argued; framing the practise of medicine as purely scientific has always been “a dangerous oversimplification”. Clinical work depends on a shifting mix of evidence, judgement, and compassion, applied to individuals in imperfect circumstances. AI does not eliminate that uncertainty; if anything, it makes it more visible.
If clinical reasoning can eventually be done better by an AI, the more important questions are not whether AI belongs in medicine, but what does it mean to be a doctor? What kinds of knowledge should medical schools prioritise? And which parts of medicine should remain in our hands, not because machines cannot do them, but because we don’t want to give them up?
Author’s note: The words in this piece were – honestly – written by me. I’ve no shame to say I used ChatGPT to help me bounce around some ideas, as well as give some suggestions for refining of my phrasing, but on the latter I tended to disagree with it more often than not.