Independent benchmark analysis / 2024 paper; dataset metadata v1.1

MIMIC-IV-Ext Clinical Decision Making

Requesting evidence is harder than receiving the finished case.

MIMIC-CDM reconstructs diagnostic tasks from records for four abdominal conditions. Its two information settings make a useful contrast: an interactive model requests findings, while a full-information model receives a prepared case. We examine the implications of that contrast without treating a four-condition sample as a general emergency-department population. The paper’s per-class accuracy also matters: its denominator is patients with the target disease, not every patient for whom a positive diagnosis might be made. Our independent analysis connects the selection rules, disease balance and result labels to the claims a hospital reader can reasonably make.

01 / What is being tested?

The task, before the score.

input
History of present illness, followed by requested findings; full-information variant supplies a prepared case.
output
Diagnosis and treatment proposal.
unit
One patient case within a selected abdominal pathology.
setting
Retrospective clinical decision simulation; historical open-weight model configurations.

Data origin. 2,400 deidentified MIMIC-IV v2.2 patients selected for appendicitis, cholecystitis, diverticulitis or pancreatitis; all required modalities must be present. [4][5][6]

Patient cases
2,400

Four selected abdominal pathologies, not an all-comer emergency cohort.

Methods [5]
Source database
MIMIC-IV v2.2

Hospital module and clinical notes underpin the derivative.

Background [5]
Diagnostic families
4

Cases with multiple selected target pathologies are excluded.

Methods [5]
Information settings
Interactive / full information

The two settings offer different amounts and timing of evidence.

Results; Methods [4]
Dataset release
v1.1

July 2024 release adds pathology IDs and corrects procedure columns.

Release Notes [5]
  1. 01

    Select the eligible cohort

    Require a target primary diagnosis and the needed clinical modalities; remove diagnosis leakage from selected text.

    [5]
  2. 02

    Reveal information

    Provide findings on request, or use the full-information comparator.

    [4][5]
  3. 03

    Collect the conclusion

    Save the diagnosis, proposed treatment and information-gathering path.

    [4]
  4. 04

    Score by pathology

    Assess the final diagnosis and other behaviors against references, with pathology-specific denominators.

    [4]

Dataset anatomy

The four-condition cohort

Appendicitis

Selected cohort.

957 patients[5]
Cholecystitis

Selected cohort.

648 patients[5]
Diverticulitis

Selected cohort.

257 patients[5]
Pancreatitis

Selected cohort.

538 patients[5]

Published cohort counts; do not interpret as population prevalence. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Per-class diagnostic accuracy

Higher is better within a fixed setting and cohort.

Correct diagnoses divided by the number of evaluated cases with that disease; headline means average across the four pathology classes.

Scoring definition

Accuracy_c = correct diagnoses in class c / evaluated patients in class c; macro = Σ accuracy_c / 4

This restricted cohort cannot supply general emergency-care specificity or positive predictive value. Full-information and interactive results are separate settings. [4]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Historical interactive decision-making results

Hager et al., July 2024; models gather information in the MIMIC-CDM simulation.

Mean of per-pathology diagnostic accuracy · %
050100
Reported
Llama 2 ChatHistorical paper configuration.
45.5%
OASSTHistorical paper configuration.
54.9%
WizardLMHistorical paper configuration.
53.9%

Paper-reported means across pathologies; no current-frontier inference. Do not combine with full-information rows.

Source: Results accompanying Figure 3 [4]

Paper-reported results / selected rows

Historical full-information comparator

Same paper, MIMIC-CDM-FI setting; relevant case information is provided up front.

Mean of per-pathology diagnostic accuracy · %
050100
Reported
Llama 2 ChatFull-information setting.
58.8%
OASSTFull-information setting.
67.8%
WizardLMFull-information setting.
65.1%

These are the paper’s comparison values, not new reruns or an estimate of benefit from a particular clinical intervention.

Source: Results accompanying Figure 3 [4]

04 / Our original analysis

What follows from the design?

01

The missing step is evidence acquisition

Published evidence

All three selected models have lower mean accuracy in the interactive setting. [4]

Our interpretation

An all-information test leaves the ability to choose useful next information largely unmeasured. The observed gap is a setting contrast, not a causal attribution to one component.

02

Cohort balance changes weighting

Published evidence

Appendicitis has 957 cases; diverticulitis has 257. [5][4]

Our interpretation

A pooled case accuracy would weight diseases differently from the paper’s mean of per-class accuracies. Preserve the aggregation definition.

03

A positive-case sample has a boundary

Published evidence

Cases are selected for four known primary pathologies. [4]

Our interpretation

These results cannot supply the false-positive burden for a hospital seeing the complete differential diagnosis.

04

Rich records still have missing dynamics

Published evidence

The data card retains only the first laboratory instance for each test. [5]

Our interpretation

This configuration is poorly suited to claims about longitudinal response to treatment. A larger record volume does not restore omitted temporal behavior.

05 / Scope of the evidence

Where this benchmark stops.

Restricted disease set

The selected four conditions do not represent all abdominal pain or the full hospital diagnostic workload. [4][5]

Retrospective selection

Requiring available imaging, tests and physical examination selects cases with particular documentation and care pathways. [5]

Historical tested models

The 2024 comparison cannot settle performance of later model generations or different agent systems. [4]

Access constraints

Patient-derived data remains credentialed and governed by PhysioNet terms. [5]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Credentialed PhysioNet access
License
PhysioNet Credentialed Health Data License 1.5.0; code MIT.
Conditions
Credentialing, required research training and a signed data use agreement are required for the files.
[5][6]

Evidence trail

Read the originals.

  1. Evaluation and mitigation of the limitations of large language models in clinical decision-making ↗

    Hager, Jungmann, Holland et al.; Nature Medicine. Original MIMIC-CDM study; historical models and restricted disease cohort.

  2. MIMIC-IV-Ext Clinical Decision Making v1.1 ↗

    Hager, Jungmann and Rueckert; PhysioNet. Official data card; 2,400 patients, disease counts, access and release notes.

  3. MIMIC Clinical Decision Making Framework ↗

    Paul Hager and collaborators. MIT code license; source data remains credentialed.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗