Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

Abstract Purpose Machine learning (ML) models have been increasingly applied to predict postoperative facial nerve dysfunction and hearing preservation after vestibular schwannoma (VS) surgery. However, reported performance varies substantially, and the overall diagnostic accuracy and clinical reliability of these models remain uncertain. We conducted a systematic review and diagnostic test accuracy meta-analysis to characterise the current state and methodological readiness of ML-based prediction of these outcomes. Methods PubMed, Embase, and CENTRAL were searched from inception to February 2026. Studies evaluating ML-based prediction of facial nerve function or hearing preservation following VS surgery were included. Diagnostic performance metrics were pooled using random-effects generalised linear mixed models. Sensitivity, specificity, diagnostic odds ratio, and AUC were synthesised, and SROC curves were constructed. The prespecified primary synthesis pooled the single best model per study; small-study effects were assessed with Deeks’ test. Risk of bias (PROBAST) and certainty of evidence (GRADE) were assessed. Results Ten retrospective cohort studies encompassing 1270 patients and 56 ML models met inclusion criteria. In the prespecified primary analysis pooling the single best model per study, the summary AUC was 0.91 for facial nerve dysfunction (sensitivity 0.89, specificity 0.86) and 0.92 for hearing preservation (sensitivity 0.88, specificity 0.96). Pooling all models on held-out test data gave a facial nerve AUC of 0.81; test-set data were too sparse for a stable hearing estimate, for which only training performance could be pooled (AUC 0.79). Tumour size, age, tumour location, and baseline hearing status were the most frequently identified influential predictors. Most studies were at unclear or high risk of bias (PROBAST has no intermediate “moderate” category), and certainty of evidence was moderate for facial nerve dysfunction and low for hearing preservation, the latter reflecting significant small-study effects (Deeks’ p  = 0.004). Conclusion ML-based models demonstrate promising discrimination for predicting postoperative facial nerve and hearing outcomes after VS surgery. However, heterogeneity, limited external validation, and inconsistent reporting of calibration constrain inference regarding transportability and clinical implementation.

More information Original publication

DOI

10.1007/s11060-026-05747-5

Type

Journal article

Publisher

Springer Science and Business Media LLC

Publication Date

2026-09-01T00:00:00+00:00

Volume

179