The future
Field 08 of 8

Surgical artificial intelligence

There are more than a thousand authorised devices with artificial intelligence and almost no randomised trial showing they improve anyone's life. Being deployed and being validated are different things, and the gap between them is the subject of this decade.

Dr. Mariano PirozzoAugust 20267 min read

Artificial intelligence entered neurosurgery through the most natural door: images. Detecting, segmenting, classifying. From there it advanced into more delicate territory — predicting the risk of an operation, anticipating complications, recognising structures in surgical video — and into one that is entirely new, that of language models.

This is probably the field where the distinction between what is announced and what is demonstrated is most needed, because the noise is at a maximum and the evidence is scarce and uneven. It is worth going through application by application.

Where it is already in the clinical pathway

The most advanced case is image triage in stroke: software that analyses the scan as soon as it is acquired, detects a large vessel occlusion and alerts the team that can treat it, without waiting for the usual chain of phone calls.

An observational study of more than four hundred and fifty thousand stroke admissions in the English public health system found that in hospitals implementing that tool the thrombectomy rate doubled1. It is a striking figure and it must be read with two caveats the study itself supports. First: it is a before-and-after design, not randomised, and it coincided in time with the deliberate expansion of the British thrombectomy network; separating the software's effect from the health policy's is not possible. Second, and more important: three of the signatories work for the manufacturer of the software evaluated, and the corresponding author holds a dual affiliation with that company1. The figure is real. The causal attribution is not demonstrated. And an article citing that study without saying who signs it is not informing: it is advertising.

In the operating room and on the pathology bench

Here the results are cleaner and more circumscribed. Stimulated Raman histology combined with neural networks allows a brain tumour diagnosis during surgery in under three minutes, against the twenty or thirty of the conventional technique. In a prospective multicentre trial, accuracy was equivalent to the pathologist's interpretation2. The key word is "equivalent", not "superior": the advantage is time and independence from the laboratory, not accuracy. And here too there are declared conflicts: the senior author is an adviser and shareholder in the equipment manufacturer, and several co-authors are employees of that company2.

The field's most solid result is probably the intraoperative molecular classifier using nanopore sequencing, with no authors from the manufacturer and a design that abstains rather than errs3. In twenty-five real operations it was correct in eighteen and in seven did not reach its confidence threshold. That is a useful tool. It is not a digital pathologist.

There is also automated reading of surgical video: systems able to recognise the phases of a procedure, the surgeon's gestures and even the quality with which they are executed4. It is validated in robotic urological surgery, at three hospitals on two continents. Extrapolating it to neurosurgery is, for now, extrapolation.

Where there is still nothing

Risk and complication prediction is the ground where the gap between expectation and evidence is widest. A systematic review of deep learning models published in neurosurgery found that almost all develop their own model and almost none validate someone else's; that code is available in fewer than one in ten papers; and that a considerable share of the information needed to reproduce them is simply not reported. Its conclusion is literal and admits no interpretation: no study described its model as ready for clinical use5.

Across medicine as a whole the picture is no better. Up to 2021 only forty-one randomised trials of machine-learning interventions had been published in the entire discipline, half of them single-centre, more than a third concentrated in one field, and none fully complied with the specific reporting guideline6.

The failure worth knowing

There is a documented case worth keeping at hand. A commercial sepsis prediction model, integrated into the electronic record system of hundreds of US hospitals, was submitted for the first time to independent external validation across almost forty thousand hospitalisations. Its discrimination proved poor: it missed two of every three patients who actually developed sepsis, and generated alerts in nearly one in five admissions7.

That is the example to cite when someone says a model "is already in use in hundreds of centres". Being deployed and being validated are different things.

To that must be added a problem discussed too little because it is uncomfortable: image classifiers trained on large databases selectively underdiagnose women, Black patients and patients of low socioeconomic status, with the worst effect on those belonging to more than one of those groups8. Underdiagnosis labels a sick person healthy and delays their treatment: it is the clinically most dangerous failure mode and the least measured.

Language models

The serious evidence is scarce and, read in full, disconcerting. In one randomised trial, giving physicians access to a language model did not improve their diagnostic reasoning; yet the model working alone scored above the physicians using it9. In another trial by the same group, on a different task, assistance did improve performance, though at the cost of more time per case — and again the model alone matched the assisted physician9.

Whatever the reading, both trials were run on clinical vignettes, not patients. There is still no randomised trial of a language model with a real clinical outcome. And a review of more than five hundred published studies in JAMA supplies the figure that orders everything else: barely five per cent used real patient data, accuracy was the main metric in the great majority, and questions of bias and equity were assessed in a minority of the work10.

What to take away

The US agency declares more than a thousand authorised devices with artificial intelligence, almost all in radiology, and warns on its own page that the list is not exhaustive and was assembled by searching for AI-related terms in authorisation summaries11. Authorising is not validating: the commonest regulatory pathway demonstrates substantial equivalence to a previous device, not that anyone lives longer or better.

Specific guidelines for reporting trials, protocols, early evaluations and prediction models involving artificial intelligence have existed since 202012. They exist, and they are poorly followed.

None of this is an argument against artificial intelligence in medicine. It is an argument about the order of the questions. The operative question of this decade will not be "artificial intelligence, yes or no?", but validated how, supervised by whom, and evaluated against what outcome. And one more, specific to anyone writing about this: who signs the study being cited. In this field, the manufacturer appearing among the authors is not the exception; it is the norm. Saying so out loud is what separates rigorous writing from a brochure.

References

Every claim in this article points to its source. The links go to the original work.

  1. Nagaratnam K, et al. Artificial intelligence imaging decision support for acute stroke treatment in England: a prospective observational study. The Lancet Digital Health. 2025;7(12):100927. doi.org/10.1016/j.landig.2025.100927
  2. Hollon TC, et al. Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks. Nature Medicine. 2020;26(1):52-58. doi.org/10.1038/s41591-019-0715-9
  3. Vermeulen C, et al. Ultra-fast deep-learned CNS tumour classification during surgery. Nature. 2023;622(7984):842-849. doi.org/10.1038/s41586-023-06615-2
  4. Kiyasseh D, et al. A vision transformer for decoding surgeon activity from surgical videos. Nature Biomedical Engineering. 2023;7(6):780-796. doi.org/10.1038/s41551-023-01010-8
  5. Huang J, et al. Deep learning for outcome prediction in neurosurgery: a systematic review of design, reporting, and reproducibility. Neurosurgery. 2022;90(1):16-38. doi.org/10.1227/NEU.0000000000001736
  6. Plana D, et al. Randomized clinical trials of machine learning interventions in health care: a systematic review. JAMA Network Open. 2022;5(9):e2233946. doi.org/10.1001/jamanetworkopen.2022.33946
  7. Wong A, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine. 2021;181(8):1065-1070. doi.org/10.1001/jamainternmed.2021.2626
  8. Seyyed-Kalantari L, et al. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature Medicine. 2021;27(12):2176-2182. doi.org/10.1038/s41591-021-01595-0
  9. Goh E, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Network Open. 2024;7(10):e2440969. doi.org/10.1001/jamanetworkopen.2024.40969
  10. Bedi S, et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA. 2025;333(4):319. doi.org/10.1001/jama.2024.21700
  11. U.S. Food and Drug Administration. Artificial intelligence-enabled medical devices. www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices
  12. Liu X, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nature Medicine. 2020;26(9):1364-1374. doi.org/10.1038/s41591-020-1034-x
← PreviousNew biomarkers
← Back to the eight fields