Requesting evidence is harder than receiving the finished case.
MIMIC-CDM reconstructs diagnostic tasks from records for four abdominal conditions. Its two information settings make a useful contrast: an interactive model requests findings, while a full-information model receives a prepared case. We examine the implications of that contrast without treating a four-condition sample as a general emergency-department population. The paper’s per-class accuracy also matters: its denominator is patients with the target disease, not every patient for whom a positive diagnosis might be made. Our independent analysis connects the selection rules, disease balance and result labels to the claims a hospital reader can reasonably make.
History of present illness, followed by requested findings; full-information variant supplies a prepared case.
output
Diagnosis and treatment proposal.
unit
One patient case within a selected abdominal pathology.
setting
Retrospective clinical decision simulation; historical open-weight model configurations.
Data origin. 2,400 deidentified MIMIC-IV v2.2 patients selected for appendicitis, cholecystitis, diverticulitis or pancreatitis; all required modalities must be present. [4][5][6]
Patient cases
2,400
Four selected abdominal pathologies, not an all-comer emergency cohort.
Published cohort counts; do not interpret as population prevalence. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Per-class diagnostic accuracy
Higher is better within a fixed setting and cohort.
Correct diagnoses divided by the number of evaluated cases with that disease; headline means average across the four pathology classes.
Scoring definition
Accuracy_c = correct diagnoses in class c / evaluated patients in class c; macro = Σ accuracy_c / 4
This restricted cohort cannot supply general emergency-care specificity or positive predictive value. Full-information and interactive results are separate settings. [4]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Historical interactive decision-making results
Hager et al., July 2024; models gather information in the MIMIC-CDM simulation.
Mean of per-pathology diagnostic accuracy · %
050100
Reported
Llama 2 ChatHistorical paper configuration.
45.5%
OASSTHistorical paper configuration.
54.9%
WizardLMHistorical paper configuration.
53.9%
Paper-reported means across pathologies; no current-frontier inference. Do not combine with full-information rows.
All three selected models have lower mean accuracy in the interactive setting. [4]
Our interpretation
An all-information test leaves the ability to choose useful next information largely unmeasured. The observed gap is a setting contrast, not a causal attribution to one component.
02
Cohort balance changes weighting
Published evidence
Appendicitis has 957 cases; diverticulitis has 257. [5][4]
Our interpretation
A pooled case accuracy would weight diseases differently from the paper’s mean of per-class accuracies. Preserve the aggregation definition.
03
A positive-case sample has a boundary
Published evidence
Cases are selected for four known primary pathologies. [4]
Our interpretation
These results cannot supply the false-positive burden for a hospital seeing the complete differential diagnosis.
04
Rich records still have missing dynamics
Published evidence
The data card retains only the first laboratory instance for each test. [5]
Our interpretation
This configuration is poorly suited to claims about longitudinal response to treatment. A larger record volume does not restore omitted temporal behavior.
05 / Scope of the evidence
Where this benchmark stops.
Restricted disease set
The selected four conditions do not represent all abdominal pain or the full hospital diagnostic workload. [4][5]
Retrospective selection
Requiring available imaging, tests and physical examination selects cases with particular documentation and care pathways. [5]
Historical tested models
The 2024 comparison cannot settle performance of later model generations or different agent systems. [4]
Access constraints
Patient-derived data remains credentialed and governed by PhysioNet terms. [5]
Paul Hager and collaborators. MIT code license; source data remains credentialed.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.