Technology
Tool 06 of 6

Artificial intelligence

Link by link, this is what artificial intelligence actually does today along a neurosurgical patient's path. The pattern repeats: where it measures well, it measures speed. And the two most honest systems here are the ones that abstain.

Dr. Mariano PirozzoAugust 20269 min read

Talking about "artificial intelligence in neurosurgery" in general leads nowhere, because the term covers things at mutually incompatible stages of maturity. It is better to walk the actual chain — from the moment the patient arrives at the emergency department to the moment postoperative management is decided — and ask, at each link, what it does, what it demonstrated and where it fails.

A companion article in this publication takes up the general question: how a medical algorithm is validated and who supervises it. This one takes up the other: what exists today, in this chain, working.

First link: the reading queue

The best-documented case is not neurosurgical but at the front door: the urgent head CT. A system trained on 37,084 studies and tested on 9,499 unseen ones achieved an area under the curve of 0.846, with 73 per cent sensitivity and 80 per cent specificity4. Mediocre figures for a diagnostic problem.

And yet, when it was deployed prospectively for three months, the median time to diagnosis of intracranial haemorrhage fell from 512 minutes to 19. A 96 per cent reduction.

That result explains the pattern that will repeat throughout what follows. The system did not diagnose better than the radiologist: it reordered the queue. It moved suspicious studies to the front of the list. The benefit came not from accuracy but from logistics, and it is an enormous benefit, because in haemorrhage time is tissue.

What happens when the algorithm changes hospital

Here is the field's structural problem, and there are three clean measurements of it.

An independent external validation of a commercial haemorrhage algorithm across 3,605 consecutive CTs found 92.3 per cent sensitivity and 97.7 per cent specificity5. Good figures. What is interesting is the failure analysis: it confused meningiomas and cortical laminar necrosis with blood, and performance dropped in patients with previous neurosurgery. The authors tried to characterise quantitatively what distinguished the studies where the model was right from those where it failed, and could not. Their sentence is the most honest in the whole deployment literature: we were unable to identify the source of this discrepancy, raising concerns about the generalisability of these tools with indeterminate failure modes.

The second measurement is more uncomfortable still. A real-world evaluation of an authorised commercial model, across 101,944 CTs from 74,142 patients in 17 centres, found 82.2 per cent sensitivity6. The dossier submitted to the regulator reported 96.15 per cent, measured on 220 cases. Fourteen points of difference between the file and reality. And the drop concentrates where it matters most: 74.8 per cent in small haemorrhages, 45.5 in subacute ones, 72.2 in outpatients.

The third is directly neurosurgical. A glioma segmentation model for intraoperative ultrasound, trained on data from six institutions, reached a Dice coefficient of 0.90 on its internal test set10. On an external validation database it fell to 0.65. Twenty-five points.

Second link: drawing the tumour

Automatic segmentation is the most mature application, and its story begins with a humbling observation: the human reference standard does not agree with itself. The benchmark that organised the field compared 20 algorithms across 65 studies annotated manually by up to four observers, and found human-to-human disagreement of between 74 and 85 per cent overlap depending on the tumour subregion1. No single algorithm won across all subregions at once, and fusing several consistently beat each one alone.

The clinically most relevant result came later. A network trained at one centre and tested on 2,034 MRIs from 532 patients across 34 institutions in a European trial achieved overlaps of 0.91 for enhancing tumour and 0.93 for non-enhancing abnormality2. And it measured something more valuable than accuracy: reliability. Automated assessment agreed with the central reference reading in 87 per cent of cases; local reading by radiologists agreed in 51 per cent. A 36-point margin.

That is the correct argument for automatic segmentation, and it is not "it measures better than the human". It is that it measures the same way every time. In a disease where the decision to change treatment depends on comparing two MRIs months apart, reproducibility is worth more than accuracy.

The counterpoint comes from the only regulator-cleared product in this block to be evaluated in stratified fashion. An algorithm approved for detecting and contouring brain metastases showed 89.30 per cent overall sensitivity across 435 lesions3. But broken down by size: 99.07 per cent for lesions of 10 millimetres or more, and falling from there. The headline figure rests on the large lesions, which are exactly the ones the human eye does not miss. Three of the authors hold a patent and shares in the company whose product is being evaluated, and they declare it.

Third link: diagnosis during surgery

This is the link where artificial intelligence did something no other technology could, and where it also delivered the field's best lesson in method.

A system combining stimulated Raman histology with a neural network was prospectively evaluated at three institutions and achieved 94.6 per cent accuracy against the pathologist's 93.9 with conventional technique, in under 150 seconds against the usual 20 to 30 minutes7. Non-inferiority in accuracy with an enormous gain in time: the same pattern again.

And there is a detail in the patient flow worth more than the headline figure. Of 291 patients imaged intraoperatively, 13 were excluded because the system did not reach its confidence threshold. That 4.5 per cent is not a failure: it is the system recognising that the case is out of its distribution and declining to answer.

The most radical version of that idea is a molecular classifier based on nanopore sequencing, evaluated in 25 operations with a deliberately high confidence threshold8. In 18 cases it returned a correct classification in under 90 minutes. In 7 — 28 per cent — it did not reach the threshold and abstained. The causes are declared: tumour classes absent from the reference dataset, low tumour purity, exotic cases. And the authors write the sentence that belongs there: for the difficult-to-diagnose cases, the system generally performed less well. It fails precisely where it is most needed, and it says so.

A later prospective multicentre validation, across 301 cases, reached 94.6 per cent concordance with the conventional integrated diagnosis and returned classification within a 30-minute intraoperative window in 18 samples sequenced during surgery, against the several weeks of the usual workflow9. Its competing interests statement is the most extensive I have read in this literature: six authors are co-founders and shareholders of the spin-off company, the first author became a full-time employee during the study, and there is a joint patent application with the sequencer manufacturer. It is all declared, and that is why it can be read.

Fourth link: watching the operation

Automatic recognition of surgical phases and gestures from video is the most talked-about application and the least mature. In neurosurgery there is exactly one validated study: 21 videos of vestibular schwannoma resection, with 81 per cent accuracy in identifying the procedure's three phases11. Its own authors classify it as stage 0 on the IDEAL scale — that is, preclinical.

The real state of the art has to be sought in laparoscopic surgery, where the volume is. A multicentre benchmark with 33 videos and 12 participating teams found phase recognition reaching F1 scores between 23.9 and 67.7 per cent, instrument detection between 38.5 and 63.8, and action recognition — the surgical gesture itself, which is what would matter — between 21.8 and 23.3 per cent12. The best team in the world, in a multicentre setting, does not reach two thirds on the easy task and does not reach a quarter on the hard one.

Single-centre studies report far higher figures. The difference between the two is precisely generalisation.

Fifth link: predicting risk

A systematic review of 47 published predictive models in neurosurgery, assessed against the standard reporting guideline, found median compliance of 82.1 per cent13. Formal reporting, then, is acceptable. But only 15 per cent of those models were externally validated in an independent cohort.

Put another way: 85 per cent of the specialty's predictive models were never tested outside the hospital where they were born.

There is an exception that teaches more than the rule. A model of functional impairment after intracranial tumour surgery was developed in 2,437 patients and externally validated in another 2,427, at centres in seven countries14. The area under the curve was 0.72 in development and 0.72 in validation: identical performance away from home. And 0.72 is modest discrimination. The price of robustness was modesty, and that trade-off is probably the field's most important lesson.

There is a subtler failure mode, and there is a beautiful measurement of it. Two classic prognostic models for head injury were applied to a contemporary multicentre cohort: discrimination held intact — c statistics between 0.83 and 0.92, excellent — while calibration broke systematically in the direction of overpredicting mortality15. The models still rank patients correctly from lower to higher risk, and at the same time hand the clinician the wrong absolute number, because the treatment of head injury has improved since they were trained. A model can age without ceasing to look good.

Who pays for the measurements

There is no quantification of funding bias specific to neurosurgery, so it has to be borrowed. A meta-analysis of 17 studies across seven commercial fracture-detection products found that industry-funded studies reported sensitivity five percentage points higher and specificity four points lower than unfunded ones16. That is trauma, not neurosurgery; the pattern is structurally the same and the extrapolation should be flagged.

In the works cited here the ties are declared and verifiable: patents, shares, direct employment, donated reagents. None of that invalidates any result. What it allows is reading them knowing who signs them, which is the only defence available.

What to take away

Walked link by link, the chain has a recognisable shape. Where artificial intelligence measures well, it measures time: 512 minutes to 19 in triage, 30 minutes to 150 seconds in intraoperative histology, weeks to hours in molecular profiling, ten minutes per MRI in response assessment. Those gains are real, reproducible and clinically important.

Where it almost never wins is in accuracy against the specialist, and where it systematically loses is on changing hospital, scanner or population. Reproducibility — doing the same thing every time — is its underrated virtue; generalisation is its structural failure.

And there is a practical criterion that follows from all of it. The two most trustworthy systems in this walk are the two that abstain: the one that set aside 4.5 per cent of cases and the one that did not answer in 28 per cent. An algorithm that always answers is not more capable: it just does not know when it does not know.

References

Every claim in this article points to its source. The links go to the original work.

  1. Menze BH, et al. The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS). IEEE Transactions on Medical Imaging. 2014;34(10):1993-2024. doi.org/10.1109/TMI.2014.2377694
  2. Kickingereder P, et al. Automated quantitative tumour response assessment of MRI in neuro-oncology with artificial neural networks: a multicentre, retrospective study. The Lancet Oncology. 2019;20(5):728-740. doi.org/10.1016/S1470-2045(19)30098-1
  3. Wang JY, et al. Stratified assessment of an FDA-cleared deep learning algorithm for automated detection and contouring of metastatic brain tumors in stereotactic radiosurgery. Radiation Oncology. 2023;18(1):61. doi.org/10.1186/s13014-023-02246-z
  4. Arbabshirani MR, et al. Advanced machine learning in action: identification of intracranial hemorrhage on computed tomography scans of the head with clinical workflow integration. npj Digital Medicine. 2018;1(1):9. doi.org/10.1038/s41746-017-0015-z
  5. Voter AF, et al. Diagnostic Accuracy and Failure Mode Analysis of a Deep Learning Algorithm for the Detection of Intracranial Hemorrhage. Journal of the American College of Radiology. 2021;18(8):1143-1152. doi.org/10.1016/j.jacr.2021.03.005
  6. Chavoshi M, et al. Real-world performance evaluation of a commercial deep learning model for intracranial hemorrhage detection. npj Digital Medicine. 2025;9(1):66. doi.org/10.1038/s41746-025-02244-3
  7. Hollon TC, et al. Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks. Nature Medicine. 2020;26(1):52-58. doi.org/10.1038/s41591-019-0715-9
  8. Vermeulen C, et al. Ultra-fast deep-learned CNS tumour classification during surgery. Nature. 2023;622(7984):842-849. doi.org/10.1038/s41586-023-06615-2
  9. Patel A, et al. Prospective, multicenter validation of a platform for rapid molecular profiling of central nervous system tumors. Nature Medicine. 2025;31(5):1567-1577. doi.org/10.1038/s41591-025-03562-5
  10. Cepeda S, et al. Deep Learning-Based Glioma Segmentation of 2D Intraoperative Ultrasound Images: A Multicenter Study Using the Brain Tumor Intraoperative Ultrasound Database (BraTioUS). Cancers. 2025;17(2):315. doi.org/10.3390/cancers17020315
  11. Williams SC, et al. Automated Operative Phase and Step Recognition in Vestibular Schwannoma Surgery: Development and Preclinical Evaluation of a Deep Learning Neural Network (IDEAL Stage 0). Neurosurgery. 2025;98(4):799-809. doi.org/10.1227/neu.0000000000003466
  12. Wagner M, et al. Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmark. Medical Image Analysis. 2023;86:102770. doi.org/10.1016/j.media.2023.102770
  13. Warman A, et al. Machine learning predictive models in neurosurgery: an appraisal based on the TRIPOD guidelines. Systematic review. Neurosurgical Focus. 2023;54(6):E8. doi.org/10.3171/2023.3.FOCUS2386
  14. Staartjes VE, et al. Development and external validation of a clinical prediction model for functional impairment after intracranial tumor surgery. Journal of Neurosurgery. 2020;134(6):1743-1750. doi.org/10.3171/2020.4.JNS20643
  15. Yue JK, et al. Performance of the IMPACT and CRASH prognostic models for traumatic brain injury in a contemporary multicenter cohort: a TRACK-TBI study. Journal of Neurosurgery. 2024;141(2):417-429. doi.org/10.3171/2023.11.JNS231425
  16. Husarek J, et al. Artificial intelligence in commercial fracture detection products: a systematic review and meta-analysis of diagnostic test accuracy. Scientific Reports. 2024;14(1):23053. doi.org/10.1038/s41598-024-73058-8
← PreviousNeurophysiological monitoring
← Back to the six tools