258P Cancer prevention in real-world laboratory data: AI-assisted data science methodologies

Logo poskytovatele

Varování

Publikace nespadá pod Filozofickou fakultu, ale pod Lékařskou fakultu. Oficiální stránka publikace je na webu muni.cz.
Autoři

RUSINKO Martin UHER Michal HALÁMKOVÁ Jana

Rok publikování 2025
Druh Stať ve sborníku
Konference ESMO Real World Data and Digital Oncology, November 01, 2025, Volume 10
Fakulta / Pracoviště MU

Lékařská fakulta

Citace
www https://doi.org/10.1016/j.esmorw.2025.100454
Doi https://doi.org/10.1016/j.esmorw.2025.100454
Klíčová slova AI; artificial intelligence; oncology; prevention; data science
Popis Background The rapid growth of real-world data from cancer prevention programs bears a great potential for early cancer detection. Data science work requires statistical and programming skills combined with medical domain expertise, thus relies on tedious collaboration between overloaded data scientists and medical oncologists. Methods We present a clinician-led initiative to analyze retrospective data of 20,000 clients in cancer prevention program in 2006–2022 (identifying over 300 cancer cases) and search for early predictive markers derived from primary data. The underlying outcome is an establishment of a secure analytic framework using generative AI (Artificial Intelligence) coding assistants, enabling clinicians to conduct their own analyses and reduce the initial time investment by data scientists. Results The dataset includes 50,174 laboratory tests from 32,759 complex visits, with 52% repeated visits in a time span of 1–3 years. Over 17,000 unscheduled tests taken beyond the visits are suggested for deeper inspection of results triggering the early retesting. The performance of analytic pipelines varied by modification. The sole use of general chatbots (ChatGPT, Gemini, Copilot) achieved only 30–40% faultless runs in Spyder IDE on real data and often required significant code modification. Stepwise coding with GitHub Copilot assistant within Visual Studio Code (VSC) using complex synthetic data led to 60–70% of successful runs on real data, but required iterative modifications to improve 50% fraction of satisfactory outcomes. Dominant improvement was achieved by preparing multiple sets of characteristic synthetic data for analytic coding in Google Colab, iterative testing and debugging via GitHub Copilot in VSC, and final execution on real data, with 80% of successful runs with similar rate of desired outcomes. Conclusions The results present a viable model for empowering clinical oncologists to accomplish data science analyses on available data and prepare robust datasets for further exploration, and for the training of ML (machine learning) models. This methodology emphasizes the emerging role of generative AI in oncology and offers an opportunity for bridging the expertise gap unavoidable in the flourish of the digital oncology era.
Související projekty:

Používáte starou verzi internetového prohlížeče. Doporučujeme aktualizovat Váš prohlížeč na nejnovější verzi.