Diagnostic Performance of an AI System for Mammography Risk Assessment: Low-Prevalence Retrospective Study
| Autoři | |
|---|---|
| Rok publikování | 2026 |
| Druh | Stať ve sborníku |
| Konference | Proceedings of the 2nd International Conference on AI in Medicine and Healthcare (AiMH' 2026) |
| Fakulta / Pracoviště MU | |
| Citace | |
| Popis | AI tools for mammography are increasingly used to support lesion detection and risk stratification, but clinical utility depends on performance in routine practice where disease prevalence is low and negative examinations predominate. We performed a retrospective diagnostic-accuracy study of consecutive screening mammography examinations acquired in January 2024 at AGEL Hospital Nový Jičín (Czech Republic) on a GE Senographe Pristina system. The reference standard was established by three senior breast radiologists using BI-RADS assessment. Of 338 examinations reviewed, 29 did not achieve consensus and were excluded, leaving 309 examinations for analysis. Consensus labels were grouped as Normal (BI-RADS 1), Benign (BI-RADS 2-3), and Suspicious (BI-RADS 4-5). The index test was the study-level output of an AI mammography system (Carebot AI MMG; Carebot s.r.o., Czech Republic). Two prespecified endpoints were evaluated: (1) any-lesion detection (BI-RADS 2-5 vs BI-RADS 1), with Medium/High Risk considered AI-positive; and (2) suspicion of malignancy (BI-RADS 4.5 vs BI-RADS 1-3), with High Risk considered AI-positive. The analysed cohort included 233 Normal, 68 Benign, and 8 Suspicious examinations (prevalence of BI-RADS 4-5 suspicion: 2.6%). For endpoint 1, sensitivity was 0.895 (95% CI 0.806–0.946) and specificity 0.940 (0.902–0.964). For endpoint 2, sensitivity was 0.875 (0.529–0.978) and specificity 0.857 (0.813–0.892). Predictive values reflected real-world prevalence: for endpoint 1, PPV 0.829 and NPV 0.965; for endpoint 2, PPV 0.140 and NPV 0.996. These findings highlight the importance of reporting predictive values in prevalence-representative cohorts and motivate further multi-centre validation and linkage to outcome-confirmed diagnoses where feasible. |