Abstract / Summary
Machine learning (ML) models have been increasingly applied to predict postoperative facial nerve dysfunction and hearing preservation after vestibular schwannoma (VS) surgery. However, reported performance varies substantially, and the overall diagnostic accuracy and clinical reliability of these models remain uncertain. We conducted a systematic review and diagnostic test accuracy meta-analysis to characterise the current state and methodological readiness of ML-based prediction of these outcomes. PubMed, Embase, and CENTRAL were searched from inception to February 2026. Studies evaluating ML-based prediction of facial nerve function or hearing preservation following VS surgery were included. Diagnostic performance metrics were pooled using random-effects generalised linear mixed models. Sensitivity, specificity, diagnostic odds ratio, and AUC were synthesised, and SROC curves were constructed. The prespecified primary synthesis pooled the single best model per study; small-study effects were assessed with Deeks' test. Risk of bias (PROBAST) and certainty of evidence (GRADE) were assessed. Ten retrospective cohort studies encompassing 1270 patients and 56 ML models met inclusion criteria. In the prespecified primary analysis pooling the single best model per study, the summary AUC was 0.91 for facial nerve dysfunction (sensitivity 0.89, specificity 0.86) and 0.92 for hearing preservation (sensitivity 0.88, specificity 0.96). Pooling all models on held-out test data gave a facial nerve AUC of 0.81; test-set data were too sparse for a stable hearing estimate, for which only training performance could be pooled (AUC 0.79). Tumour size, age, tumour location, and baseline hearing status were the most frequently identified influential predictors. Most studies were at unclear or high risk of bias (PROBAST has no intermediate "moderate" category), and certainty of evidence was moderate for facial nerve dysfunction and low for hearing preservation, the latter reflecting significant small-study effects (Deeks' p = 0.004). ML-based models demonstrate promising discrimination for predicting postoperative facial nerve and hearing outcomes after VS surgery. However, heterogeneity, limited external validation, and inconsistent reporting of calibration constrain inference regarding transportability and clinical implementation.
Topics
Primary Source
Journal of neuro-oncology
Ask Prognia AI
Have questions about this meta-analysis?
Prognia AI can search this source alongside 35M+ PubMed papers and current ESC, AHA, NICE, and ADA guidelines to give you a fully cited clinical answer.