Background & objective
Commercial AI tools for detecting acute intracranial hemorrhage (AIH) on non-contrast brain CT are proliferating, but independent head-to-head comparisons under real-world conditions are lacking. Procurement often relies on vendor-reported, single-metric numbers. The study set out to evaluate four regulatory-approved commercial AI solutions across multiple performance dimensions — discrimination, confirmatory performance, calibration, and volume measurement.

Methods
Retrospective, single-center study at Eunpyeong St. Mary's Hospital (Catholic University of Korea). Adults who had emergency non-contrast brain CT for suspected AIH (Feb–Mar 2024) were screened; after exclusions, 436 scans were analyzed (AIH prevalence 12%, n=52). Three neuroradiologists established ground truth and manually measured hemorrhage volume (inter-reader agreement ICC 1.00).
The four solutions (blinded as A–D): A = JLK-ICH, B = HyperInsight-ICH (Purple AI), C = Heuron-ICH, D = AVIEW NeuroCAD. All are MFDS-designated innovative devices; A, B, C are also FDA-approved. Metrics: AUROC, AUPRC, and Brier score from probability scores (A, B, C only); sensitivity, specificity, precision, and F1 from binary output (all four); and Bland–Altman volumetric agreement. Sub-analyses covered hemorrhage subtype, etiology, and detection failures.
Results
- Discrimination was uniformly high — AUROC 0.96–0.99 and sensitivity 0.85–0.92, with no significant pairwise differences (P > 0.05).
- Solution B (PurpleAI) led confirmatory performance — highest AUPRC (0.97) and lowest Brier score (0.02), significantly better than the others (P < 0.01).
- Solutions B and D led binary precision-oriented metrics — significantly higher specificity (1.00 and 0.99), precision (0.90–0.98), and F1 (0.87–0.94) than A and C.
- Solution D had the best volume agreement — lowest mean difference (−0.87 mm³) and narrowest limits of agreement (−13.4 to 11.6 mm³).
- Failures clustered in small hemorrhages — false negatives occurred mostly in bleeds <10 mm³, especially SAH; IVH was reliably detected by all. Traumatic vs non-traumatic etiology made no significant difference.
Conclusion
In a real-world emergency setting, all four solutions showed excellent discrimination, but meaningful differences emerged in confirmatory performance (B strongest) and volumetric agreement (D strongest). The tools are not interchangeable — each emphasizes different strengths. The authors argue AI procurement in acute stroke imaging should be guided by a structured, multimetric evaluation aligned to institutional priorities rather than a single aggregate metric.
Limitations
Retrospective, single-center, single CT vendor (Siemens); limited to solutions available on-site; small per-subtype counts (subgroup results are hypothesis-generating); no assessment of topographic accuracy, processing speed, usability, or clinical outcomes.