StayCurrentMD · Solving Complex Pediatric Surgical Case Studies: A Comparative Analysis of Copilot, ChatGPT-4, and Experienced Pediatric Surgeons' Performance
Infographic1 min read·Published Jul 2025

Solving Complex Pediatric Surgical Case Studies: A Comparative Analysis of Copilot, ChatGPT-4, and Experienced Pediatric Surgeons' Performance

Infographic comparing AI language models versus pediatric surgeons on complex clinical case assessments

Infographic · Jul 2025 · 1 min read

In brief

In brief

Comparative study evaluating ChatGPT-4 and Microsoft Copilot against experienced pediatric surgeons using 13 complex case vignettes. AI models achieved 47-52% accuracy versus 68.8% for human surgeons, with ChatGPT-4 showing superior differential diagnosis generation but overall limited clinical reliability.

  • ChatGPT-4 scored 52% vs Copilot's 48% on pediatric surgical cases, both significantly below experienced surgeons' 69% (p<0.01).
  • ChatGPT-4 outperformed Copilot in generating differential diagnoses (p<0.05) but showed no advantage for primary diagnosis or diagnostics.
  • Pediatric surgeons rated LLM diagnostic recommendations as only 'average' in completeness and accuracy for complex clinical scenarios.
  • Current AI models demonstrate significant limitations in pediatric surgical decision-making and cannot reliably replace clinical expertise.
  • LLMs show potential as adjunct tools but require substantial improvement before clinical deployment in pediatric surgery.

Written by the GCMD Library team from the infographic.

Three-panel infographic with teal headers. Left panel shows isometric illustrations of data cubes, DNA, and medical imagery alongside icons for 13 clinical cases and a 96-question test. Center panel displays AI chip icon versus three diverse surgeons with purple and teal color blocks. Right panel shows podium-style bar chart with three positions in purple and teal, displaying percentage scores.

Richard Gnatzy, Martin Lacher, Michael Berger, Michael Boettcher, Oliver J Deffaa, Joachim Kübler, Omid Madadi-Sanjani, Illya Martynov, Steffi Mayer, Mikko P Pakarinen, Richard Wagner, Tomas Wester, Augusto Zani, Ophelia Aubert 

The emergence of large language models (LLMs) has led to notable advancements across multiple sectors, including medicine. Yet, their effect in pediatric surgery remains largely unexplored. This study aims to assess the ability of the artificial intelligence (AI) models ChatGPT-4 and Microsoft Copilot to propose diagnostic procedures, primary and differential diagnoses, as well as answer clinical questions using complex clinical case vignettes of classic pediatric surgical diseases.We conducted the study in April 2024. We evaluated the performance of LLMs using 13 complex clinical case vignettes of pediatric surgical diseases and compared responses to a human cohort of experienced pediatric surgeons. Additionally, pediatric surgeons rated the diagnostic recommendations of LLMs for completeness and accuracy. To determine differences in performance, we performed statistical analyses.ChatGPT-4 achieved a higher test score (52.1%) compared to Copilot (47.9%) but less than pediatric surgeons (68.8%). Overall differences in performance between ChatGPT-4, Copilot, and pediatric surgeons were found to be statistically significant (p < 0.01). ChatGPT-4 demonstrated superior performance in generating differential diagnoses compared to Copilot (p < 0.05). No statistically significant differences were found between the AI models regarding suggestions for diagnostics and primary diagnosis. Overall, the recommendations of LLMs were rated as average by pediatric surgeons.This study reveals significant limitations in the performance of AI models in pediatric surgery. Although LLMs exhibit potential across various areas, their reliability and accuracy in handling clinical decision-making tasks is limited. Further research is needed to improve AI capabilities and establish its usefulness in the clinical setting.

The text in the image

Complex cases: AI or pediatric surgeons? | 13 standardized complex clinical cases | Multiple choice test | 96 questions | Multi-National | Survey | 2024 | Comparison | Large language Model | 11 Expert pediatric surgeons | AI | vs. | SURGEONS | Outcome: Correct Answers | Chat GPT-4 | 52% | Surgeons | 69% | Copilot | 48% | 1 | 2 | 3 | EJPS | European Journal of Pediatric Surgery | Richard Gnatzy, et al. apr 2025 | DOI: 10.1055/a-2551-2131 | Autores: José Campos V. | Journal Hive

Try
Intelligent Search· scoped to this infographic · not medical adviceSearch the whole library →