| Abstract (ENG): |
Artificial intelligence-enabled clinical decision support systems (AI-CDSS) are increasingly embedded in diagnostic and therapeutic workflows, yet evaluation, procurement, and governance often remain anchored in model-centric indicators such as discrimination, calibration, sensitivity, and specificity. These indicators are necessary but insufficient for determining whether a deployed AI-CDSS improves patient-relevant outcomes, clinical workflow, resource use, economic performance, clinician and patient experience, learning, and equity. This article proposes a testable conceptual framework linking the Diagnostic AI Contribution Score (DACS) to outcome-domain prioritization and auditable measurement design. The contribution is theory-building and methodological rather than empirically validating: no new patient-level data were generated or analyzed. We conducted a concept-driven structured narrative synthesis across clinical AI evaluation, clinical decision support, health-services research, value-based care, software measurement, AI governance, and large language model (LLM) deployment. We used the synthesis to derive a DACS-informed framework for selecting outcome domains, specifying exposure–action–outcome linkage, and defining minimum telemetry needed for auditability. The revised framework provides four core outputs: (i) a six-domain taxonomy of outcome metrics for AI-CDSS, separating learning/governance from equity; (ii) a comparison with existing AI evaluation and reporting frameworks to clarify the framework’s incremental contribution; (iii) a DACS-to-domain mapping and explicitly illustrative prioritization heuristic, accompanied by a threshold-sensitivity logic; and (iv) a four-step measurement framework with minimum telemetry requirements and clinical implementation guardrails. LLM-specific extensions address provenance and version tracking, human mediation, documentation-quality audits, latency and compute metering, and workflow-mediated effects. We further specify clinical safety considerations including alert fatigue, override patterns by user group, automation bias and diagnostic anchoring, educational effects, de-implementation criteria, fallback workflows, integration with existing quality-management infrastructure, and patient-facing disclosure where appropriate. The proposed framework should be interpreted as a structured, hypothesis-generating model for proportional outcome measurement and governance. Future work should empirically test inter-rater reliability of DACS scoring, validate DACS-to-domain mappings through expert elicitation and prospective deployments, calibrate thresholds across use cases, and evaluate whether standardized telemetry improves attribution, monitoring, and accountability for AI-CDSS. |
| Citation: |
Kirchhoff, Jan and Berns, Fabian and Schieder, Christian and Schobel, Johannes
(2026)
A testable framework linking diagnostic AI contribution to outcome measurement in clinical decision support.
Frontiers in Digital Health, Section Health Informatics, 8, Paper 1871438. / Hypothesis and theory article.
ISSN 2673-253X
|