Self-supervised multimodal learning for survival prediction in glioblastoma: a multicenter study from the ReSPOND consortium
Recommended Citation
Yu F, Guo J, Tian Y, Garcia JA, Akbari H, Matsumoto Y, Xing R, Han T, Kazerooni AF, Bakas S, Baid U, Villanueva-Meyer J, Brem S, Lustig R, O’Rourke DM, Bagley S, Calabrese E, Rudie J, Chang S, LaMontagne P, Marcus DS, Balana C, Capellades J, Puig J, Barnholtz-Sloan J, Badve C, Sloan A, Waite K, Colen R, Choi Y. Self-supervised multimodal learning for survival prediction in glioblastoma: a multicenter study from the ReSPOND consortium. Neuro Oncol 2025; 27(Supplement 5):v288.
Document Type
Conference Proceeding
Publication Date
11-11-2025
Publication Title
Neuro Oncol
Keywords
glioblastoma, heterogeneity, adult, foreign medical graduates, o(6)-methylguanine-dna methyltransferase, vision, brain, diagnostic imaging, neoplasms, patient prognosis, stratification, fluid attenuated inversion recovery, multiparametric magnetic resonance imaging, datasets, radiomics, c statistic, convolutional neural networks, autoencoder, multilayer perceptrons
Abstract
PURPOSE: Glioblastoma is the most aggressive adult brain tumor, with a median overall survival of approximately 15 months. It is important to build accurate prognostic models for glioblastoma patients to inform clinical management and trials. This study proposes a self-supervised learning-based approach with multimodal data integration for survival prediction and prognostic stratification of glioblastoma patients on the ReSPOND consortium. METHODS: We curated a multi-parametric MRI dataset (T1, T1CE, T2, FLAIR) of 3,119 glioblastoma patients from 22 institutions across 3 continents. Masked autoencoder (MAE) was adapted to pretrain a Vision Transformer (ViT) encoder by reconstructing the masked image patches. The encoder was utilized for extracting patch embeddings for survival tasks, with cross-attention mechanism to incorporate the molecular and clinical information (age, sex, extent of resection, MGMT) to guide imaging feature aggregation. Imaging and clinical embeddings were fused through a multi-layer perceptron (MLP) for log-risk hazard estimation, optimized using Cox partial likelihood. Model performance and generalizability were assessed via k-fold cross-validation on the ReSPOND consortium and the leave-one-site-out validation was performed on 11 institutions comparing with CoxPH, DeepSurv and DeepHit. Prognostic risk stratification via Kaplan-Meier analysis divided the patients into low-, medium- and high-risk subgroups per site. RESULTS: Multimodal data integration using the proposed framework achieved the highest C-index (0.674 ± 0.017) on the ReSPOND consortium. Integration of clinical information and MGMT consistently boosted the performance of the proposed model across sites (0.615 ± 0.046 vs. 0.662 ± 0.044). The imaging-based approaches, i.e., radiomics and convolutional neural network (CNN) features performed less robustly. The Kaplan-Meier curves and log-rank tests suggested the proposed framework achieved more separable prognostic subgroups. CONCLUSION: The proposed self-supervised multimodal learning framework shows promise for survival prediction and prognostic risk stratification in glioblastoma. It highlights the challenge for clinical model deployment due to the data heterogeneity in multi-institutional cohort.
Volume
27
Issue
Supplement 5
First Page
v288
