Menu
Not yet recruiting NCT07414966

Scalable Clinical Oversight of Large Language Models Via Uncertainty Triangulation

No phase Interventional Coronary Heart Disease (CHD)

For patients and families

In plain language

An automatic summary of structured registry data. It is an orientation aid, not a substitute for the official protocol or a physician assessment.

What is being studied
The protocol lists: SCOUT-Assisted Review Workflow, Standard Manual Review Workflow.
Who it may be relevant to
Registry conditions: Coronary Heart Disease (CHD). Basic parameters: from 18 years · All.
What needs checking
Age, condition and sex are only basic indicators. Prior treatment, laboratory values and other mandatory requirements appear in the eligibility criteria below.
Where it takes place
Center list to be confirmed — check the primary protocol.
Next step
Save the trial, show it to the treating physician, and confirm current recruitment with the study center. Costs, documents and travel →
Official title

Prospective Evaluation of a Model-Agnostic Meta-Verification Framework (SCOUT) for Scalable Clinical Oversight of Large Language Model Outputs in Coronary Heart Disease Diagnosis: A Multi-Reader, Randomized, Crossover Trial

Overview

This prospective, multi-reader, randomized crossover trial evaluates SCOUT (Scalable Clinical Oversight via Uncertainty Triangulation), a model-agnostic meta-verification framework that selectively defers unreliable large language model (LLM) predictions to clinicians by triangulating three orthogonal uncertainty signals: model heterogeneity, stochastic inconsistency, and reasoning critique. The trial assesses whether SCOUT-assisted review can reduce physician review time compared with standard manual review of AI-generated diagnoses while maintaining non-inferior diagnostic accuracy in coronary heart disease (CHD) subtyping.

Detailed description

Background: Large language models are increasingly deployed in clinical workflows, yet requiring clinician review of every AI output negates the efficiency gains that motivate their adoption. SCOUT addresses this efficiency-safety paradox through algorithmic meta-verification.

The SCOUT framework triangulates three orthogonal external signals to determine case-level uncertainty: (1) Model Heterogeneity - whether a structurally different auxiliary LLM agrees with the primary model; (2) Stochastic Inconsistency - whether repeated sampling from the same model yields divergent outputs; (3) Reasoning Critique - whether an external checker model identifies logical flaws in the chain-of-thought reasoning.

In this crossover trial, 7 clinicians of varying seniority (2 junior residents, 3 senior residents, 2 attending physicians) each review all 110 cases under both standard manual review and SCOUT-assisted review workflows. The study evaluates workflow efficiency (primary endpoint) and diagnostic accuracy (secondary endpoint).

Interventions

  • Diagnostic test SCOUT-Assisted Review Workflow
    SCOUT-Assisted Review (Intervention Arm): Physicians review 56 cases processed through the SCOUT framework. For cases classified as low-uncertainty (D(x)=0), the AI prediction is auto-accepted without physician review. For high-uncertainty cases (D(x)=1), the physician reviews the case with access to the main model's chain-of-thought reasoning and the meta-verification audit results. The main model is DeepSeek-V3.1 with chain-of-thought prompting.
  • Diagnostic test Standard Manual Review Workflow
    Physicians perform a full manual review of 54 cases using raw medical records with access to the AI model's predictions and reasoning, but without SCOUT uncertainty stratification or selective deferral.

Primary outcome measures

  • Mean physician review time per case (minutes) [Time frame: Through study completion, an average of 2 hours.]
Secondary outcome measures (2)
  • Diagnostic accuracy (%) [Time frame: Through study completion, an average of 2 hours.]
  • Computational Return on Investment (ROI) [Time frame: Through study completion, an average of 2 hours.]

Eligibility criteria

Inclusion criteria

  • Board-certified or in-training cardiologists at Fuwai Hospital
  • Spanning three experience strata: junior residents, senior residents, attending physicians

Exclusion criteria

  • Clinicians involved in the development or optimization of the SCOUT framework
  • Clinicians involved in the gold-standard adjudication process

Criteria are shown verbatim from the registry (in English). Final eligibility is always assessed by the study center.

Healthy volunteers: No

Study design

Allocation
Randomized
Model
Crossover
Masking
Open label
Primary purpose
Diagnostic

Study locations

Center list to be confirmed — check the primary protocol.

Identifiers

NCT: NCT07414966 · 2025-2702-1

Primary sources (government registries)

View this study on ClinicalTrials.gov ↗