Skip to main navigation Skip to search Skip to main content

Assessment of bias in scoring of AI-based radiotherapy segmentation and planning studies using modified TRIPOD and PROBAST guidelines as an example

  • Coen Hurkmans (Corresponding author-nrf)
  • , Jean-Emmanuel Bibault
  • , Enrico Clementel
  • , Jennifer Dhont
  • , Wouter van Elmpt
  • , Georgios Kantidakis
  • , Nicolaus Andratschke

Research output: Contribution to journalArticleAcademicpeer-review

42 Downloads (Pure)

Abstract

Background and purpose: Studies investigating the application of Artificial Intelligence (AI) in the field of radiotherapy exhibit substantial variations in terms of quality. The goal of this study was to assess the amount of transparency and bias in scoring articles with a specific focus on AI based segmentation and treatment planning, using modified PROBAST and TRIPOD checklists, in order to provide recommendations for future guideline developers and reviewers. Materials and methods: The TRIPOD and PROBAST checklist items were discussed and modified using a Delphi process. After consensus was reached, 2 groups of 3 co-authors scored 2 articles to evaluate usability and further optimize the adapted checklists. Finally, 10 articles were scored by all co-authors. Fleiss’ kappa was calculated to assess the reliability of agreement between observers. Results: Three of the 37 TRIPOD items and 5 of the 32 PROBAST items were deemed irrelevant. General terminology in the items (e.g., multivariable prediction model, predictors) was modified to align with AI-specific terms. After the first scoring round, further improvements of the items were formulated, e.g., by preventing the use of sub-questions or subjective words and adding clarifications on how to score an item. Using the final consensus list to score the 10 articles, only 2 out of the 61 items resulted in a statistically significant kappa of 0.4 or more demonstrating substantial agreement. For 41 items no statistically significant kappa was obtained indicating that the level of agreement among multiple observers is due to chance alone. Conclusion: Our study showed low reliability scores with the adapted TRIPOD and PROBAST checklists. Although such checklists have shown great value during development and reporting, this raises concerns about the applicability of such checklists to objectively score scientific articles for AI applications. When developing or revising guidelines, it is essential to consider their applicability to score articles without introducing bias.

Original languageEnglish
Article number110196
Number of pages6
JournalRadiotherapy and Oncology
Volume194
DOIs
Publication statusPublished - May 2024

Bibliographical note

Publisher Copyright:
© 2024 The Authors

Keywords

  • Artificial intelligence
  • Bias
  • Checklists
  • Distinctiveness
  • Guidelines
  • Inter-observer variation
  • Oncology
  • Radiation therapy
  • Transparency
  • Neoplasms/radiotherapy
  • Reproducibility of Results
  • Radiotherapy Planning, Computer-Assisted/methods
  • Humans
  • Artificial Intelligence
  • Delphi Technique
  • Checklist
  • Practice Guidelines as Topic

Fingerprint

Dive into the research topics of 'Assessment of bias in scoring of AI-based radiotherapy segmentation and planning studies using modified TRIPOD and PROBAST guidelines as an example'. Together they form a unique fingerprint.

Cite this