Surgical artificial intelligence
There are more than a thousand authorised devices with artificial intelligence and almost no randomised trial showing they improve anyone's life. Being deployed and being validated are different things, and the gap between them is the subject of this decade.
Artificial intelligence entered neurosurgery through the most natural door: images. Detecting, segmenting, classifying. From there it advanced into more delicate territory — predicting the risk of an operation, anticipating complications, recognising structures in surgical video — and into one that is entirely new, that of language models.
This is probably the field where the distinction between what is announced and what is demonstrated is most needed, because the noise is at a maximum and the evidence is scarce and uneven. It is worth going through application by application.
Where it is already in the clinical pathway
The most advanced case is image triage in stroke: software that analyses the scan as soon as it is acquired, detects a large vessel occlusion and alerts the team that can treat it, without waiting for the usual chain of phone calls.
An observational study of more than four hundred and fifty thousand stroke admissions in the English public health system found that in hospitals implementing that tool the thrombectomy rate doubled1. It is a striking figure and it must be read with two caveats the study itself supports. First: it is a before-and-after design, not randomised, and it coincided in time with the deliberate expansion of the British thrombectomy network; separating the software's effect from the health policy's is not possible. Second, and more important: three of the signatories work for the manufacturer of the software evaluated, and the corresponding author holds a dual affiliation with that company1. The figure is real. The causal attribution is not demonstrated. And an article citing that study without saying who signs it is not informing: it is advertising.
In the operating room and on the pathology bench
Here the results are cleaner and more circumscribed. Stimulated Raman histology combined with neural networks allows a brain tumour diagnosis during surgery in under three minutes, against the twenty or thirty of the conventional technique. In a prospective multicentre trial, accuracy was equivalent to the pathologist's interpretation2. The key word is "equivalent", not "superior": the advantage is time and independence from the laboratory, not accuracy. And here too there are declared conflicts: the senior author is an adviser and shareholder in the equipment manufacturer, and several co-authors are employees of that company2.
The field's most solid result is probably the intraoperative molecular classifier using nanopore sequencing, with no authors from the manufacturer and a design that abstains rather than errs3. In twenty-five real operations it was correct in eighteen and in seven did not reach its confidence threshold. That is a useful tool. It is not a digital pathologist.
There is also automated reading of surgical video: systems able to recognise the phases of a procedure, the surgeon's gestures and even the quality with which they are executed4. It is validated in robotic urological surgery, at three hospitals on two continents. Extrapolating it to neurosurgery is, for now, extrapolation.
Where there is still nothing
Risk and complication prediction is the ground where the gap between expectation and evidence is widest. A systematic review of deep learning models published in neurosurgery found that almost all develop their own model and almost none validate someone else's; that code is available in fewer than one in ten papers; and that a considerable share of the information needed to reproduce them is simply not reported. Its conclusion is literal and admits no interpretation: no study described its model as ready for clinical use5.
Across medicine as a whole the picture is no better. Up to 2021 only forty-one randomised trials of machine-learning interventions had been published in the entire discipline, half of them single-centre, more than a third concentrated in one field, and none fully complied with the specific reporting guideline6.
The failure worth knowing
There is a documented case worth keeping at hand. A commercial sepsis prediction model, integrated into the electronic record system of hundreds of US hospitals, was submitted for the first time to independent external validation across almost forty thousand hospitalisations. Its discrimination proved poor: it missed two of every three patients who actually developed sepsis, and generated alerts in nearly one in five admissions7.
That is the example to cite when someone says a model "is already in use in hundreds of centres". Being deployed and being validated are different things.
To that must be added a problem discussed too little because it is uncomfortable: image classifiers trained on large databases selectively underdiagnose women, Black patients and patients of low socioeconomic status, with the worst effect on those belonging to more than one of those groups8. Underdiagnosis labels a sick person healthy and delays their treatment: it is the clinically most dangerous failure mode and the least measured.
Language models
The serious evidence is scarce and, read in full, disconcerting. In one randomised trial, giving physicians access to a language model did not improve their diagnostic reasoning; yet the model working alone scored above the physicians using it9. In another trial by the same group, on a different task, assistance did improve performance, though at the cost of more time per case — and again the model alone matched the assisted physician9.
Whatever the reading, both trials were run on clinical vignettes, not patients. There is still no randomised trial of a language model with a real clinical outcome. And a review of more than five hundred published studies in JAMA supplies the figure that orders everything else: barely five per cent used real patient data, accuracy was the main metric in the great majority, and questions of bias and equity were assessed in a minority of the work10.
What to take away
The US agency declares more than a thousand authorised devices with artificial intelligence, almost all in radiology, and warns on its own page that the list is not exhaustive and was assembled by searching for AI-related terms in authorisation summaries11. Authorising is not validating: the commonest regulatory pathway demonstrates substantial equivalence to a previous device, not that anyone lives longer or better.
Specific guidelines for reporting trials, protocols, early evaluations and prediction models involving artificial intelligence have existed since 202012. They exist, and they are poorly followed.
None of this is an argument against artificial intelligence in medicine. It is an argument about the order of the questions. The operative question of this decade will not be "artificial intelligence, yes or no?", but validated how, supervised by whom, and evaluated against what outcome. And one more, specific to anyone writing about this: who signs the study being cited. In this field, the manufacturer appearing among the authors is not the exception; it is the norm. Saying so out loud is what separates rigorous writing from a brochure.
