Dataset related to article "Concordance Between Human and AI-Based clinical follow-up after Total Hip Arthroplasty for prioritizing outpatient services"
Authors/Creators
Description
This record contains raw data related to article "Concordance Between Human and AI-Based clinical follow-up after Total Hip Arthroplasty for prioritizing outpatient services"
Abstract
Purpose
To evaluate concordance between human clinical grading and a multimodal artificial intelligence (AI) virtual follow-up model for prioritising outpatient follow-up after total hip arthroplasty (THA).
Methods
This concordance study compared human reference grading with AI model predictions in 603 THA follow-up cases. The AI system integrated radiographic deep-learning outputs with machine-learning branches based on clinical and comorbidity data. Agreement and classification performance were assessed using cross-tabulation, observed agreement, Cohen’s kappa, sensitivity, specificity, predictive values, likelihood ratios, odds ratio, and AUC with 95% confidence intervals.
Results
Human grading classified 583 cases as normal and 20 as abnormal, whereas the AI model classified 586 as normal and 17 as abnormal. Cross-tabulation showed 582 true negatives, 16 true positives, one false positive, and four false negatives, corresponding to five discordant cases. Observed agreement was 99.2%, and Cohen’s kappa was 0.86 (standard error 0.04; Z = 21.21; p < 0.001). Sensitivity was 80.0% (95% confidence interval [CI] 56.3–94.3), specificity was 99.8% (95% CI 99.0–100.0), positive predictive value was 94.1%, negative predictive value was 99.3%, and ROC area was 0.90.
Conclusion
The AI-supported virtual follow-up model showed excellent concordance with human clinical follow-up and very high specificity in detecting abnormal cases. However, four false-negative cases indicate that the tool should support, rather than replace, clinician-led surveillance.