The Gynecology Oncology AI Race: A Comparison of Clinical Performance of Four AI Platforms
Recommended Citation
Mekkaoui T, Kheil M, Miller M, Lamah E, Hijaz M. The Gynecology Oncology AI Race: A Comparison of Clinical Performance of Four AI Platforms. Gynecol Oncol 2026; 208:S239.
Document Type
Conference Proceeding
Publication Date
5-1-2026
Publication Title
Gynecol Oncol
Keywords
adjuvant chemotherapy, adult, artificial intelligence, ChatGPT, clinical article, clinical decision making, clinical practice, clinical practice guideline, conference abstract, diagnosis, electrocorticography, endometrium cancer, evidence based practice, female, female genital tract cancer, functional status, genetic screening, gynecologic oncologist, histology, human, middle aged, patient counseling, postoperative monitoring, practice guideline, race, radiotherapy, scoring system, treatment response, uterine cervix cancer
Abstract
Objectives: Artificial intelligence (AI) models may assist clinicians with documentation, patient counseling and clinical decision-making. Existing literature has primarily described the ways in which gynecologic oncologists are using ChatGPT in their clinical practice. In the present study, we sought to compare the performance of four different AI platforms to an academic, multidisciplinary tumor board in the staging of disease and provision of evidence-based treatment recommendations for cases of differing complexity. Methods: A case complexity scoring system that accounts for disease burden, histology, prior treatment response, ECOG status, and extent of cytoreduction achieved with surgery, was developed by the authors. Complexity was categorized as “low” (0–9), “moderate” (10–18), or “high” (19–30). Five gynecologic oncology cases of different complexity levels were selected from an academic institution’s multidisciplinary tumor board (TB) list. Four AI platforms: ChatGPT5, Open Evidence, Claude 4.5, Gemini, were then prompted to assign a stage and generate treatment recommendations for selected cases and results were compared to the TB’s determinations. Results: All platforms demonstrated a high level of concordance in disease staging with no notable discrepancies. For endometrial cancer (EC) with high grade histology and <50% myometrial invasion, all platforms acknowledged the aggressive histology but continued to identify the stage as IA per FIGO 2009 system (Case 2). Adjuvant treatment recommendations for this vignette varied across different platforms, with most suggesting further cytotoxic treatment congruent with 2023 NCCN guidelines. Discrepantly, Open Evidence offered surveillance or radiation therapy (RT) alone as alternative options, and Claude failed to recommend chemotherapy altogether; however, when prompted to edit recommendations per updated guidelines, Claude successfully identified adjuvant chemotherapy as the appropriate treatment. Interestingly, for stage IA cervical cancer where all platforms agreed on post operative surveillance, Claude also suggested RT as an alternative option (Case 1). None of the platforms recommended immunotherapy (IO) for primary treatment of advanced stage EC, but all agreed that it should be included in the recurrent setting. Some platforms generated considerations beyond standard medical recommendations. Open Evidence demonstrated the ability to modify management - including type of RT and dosing of chemotherapy - based on patients’ functional status. Gemini often expanded recommendations to include emerging or experimental approaches. Claude was the only platform to consistently recommend genetic testing and counseling as part of its treatment plan. Conclusions: All AI platforms exhibited accuracy in the staging of different gynecologic cancers. The platforms were relatively consistent in their identification of standard of care treatments but varied in their proposals of alternative or additional therapies. Ultimately, the increased utilization of AI tools in the diagnosis, treatment, and counseling of patients is inevitable, and our findings suggest that different AI platforms may be equally reliable tools to bolster clinical decision-making. [Formula presented]
Volume
208
First Page
S239
