OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
The research introduces OpenTumorBoard, a real-world benchmark featuring 611 patient cases and 19,157 discussion turns extracted from tumor board recordings. The benchmark evaluates models across specialist response generation and full board simulation tasks. Evaluating 14 general-purpose frontier and medical LLMs reveals substantial limitations, with top models scoring 3.43 out of 5 in clinical equivalence and 2.78 out of 5 in alignment with recorded conclusions. Supervised finetuning and reinforcement learning are shown to improve performance on held-out test sets, highlighting opportunities for model adaptation.
OpenTumorBoard includes 611 patient cases and 19,157 discussion turns across ten specialist roles transcribed from 12,534 minutes of recordings.
The best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions.