Benchmark analysis / 4 min read

When should a hospital use EHRSHOT or MIMIC-CDM?

Choose between structured-record prediction and diagnostic simulation by matching the benchmark’s inputs, targets and evidence limits.

The short answer

EHRSHOT and MIMIC-CDM both use clinical records, but they test different work. EHRSHOT studies prediction from coded longitudinal history. MIMIC-CDM studies diagnosis under specified information conditions. Our comparison is a task-selection framework, not a contest between dataset sizes. The useful question is which benchmark provides evidence for the system a hospital intends to evaluate, and which important properties remain outside that evidence.

Write the intended output first

A prediction system estimates a task-defined outcome from information available before a specified time. A diagnostic agent gathers or receives evidence and produces a diagnosis, potentially with a treatment proposal. Those outputs can be related clinically while requiring different evaluation designs.

Our proposed first step is to describe the product’s actual output in one sentence. If it predicts a coded future event, inspect EHRSHOT’s task definitions and label timing. If it conducts an information-gathering diagnostic process, inspect MIMIC-CDM’s interactive protocol. A shared hospital context does not make either benchmark a substitute for the other.

Match the information available at the decision time

EHRSHOT uses structured histories up to prediction time. MIMIC-CDM separates case information so it can be revealed on request or supplied together. Each design makes assumptions about what the system is allowed to know before producing its output.

Our deduction is that information availability is part of the task, not a preprocessing detail. A local evaluation should inspect whether records arrive late, whether a code appears only after an event and whether a prepared summary contains information unavailable at the intended moment of use. Matching model architecture is less important than matching the decision’s evidence boundary.

Compare populations before comparing scores

EHRSHOT is drawn from Stanford records with its stated cohort restrictions. MIMIC-CDM is derived from MIMIC-IV and selected for four abdominal pathologies with required modalities. Neither dataset is simply a generic hospital population. Their source institutions and selection processes matter to interpretation.

A practical comparison note records the source setting, eligible patients, exclusions and target events. Then identify differences from the local population that could alter coding, documentation or disease mix. This is a transportability question. It cannot be answered by selecting whichever benchmark has more events, more labels or a more impressive headline result.

Preserve metrics that answer different questions

EHRSHOT emphasizes discrimination metrics for prediction tasks. MIMIC-CDM emphasizes correct diagnosis within each selected pathology. Their percentages and areas under curves should not be placed on a common scale and averaged into a hospital AI grade.

Our original recommendation is an evidence map with distinct endpoints. It can show that a system has evidence for a prediction task and separate evidence for an information-gathering task without inventing a single combined score. Where the intended workflow joins those capabilities, evaluate the handoff directly rather than infer it from success on two disconnected tests.

Plan the smallest evaluation that closes the relevant gap

For a prediction product, the next question might concern local label timing, calibration or performance across the intended cohort. For a diagnostic simulation, it might concern which facts the agent requests and whether the available findings match the local workflow. These are proposed evaluation questions arising from our comparison, not claims that the cited benchmarks already measured them.

Use the official papers for task definitions and the repositories or data cards for implementation and access. Keep a manifest of versions and transformations. A useful benchmark selection ends with a clear boundary: what the published evidence supports, why that matters to the product and exactly which local question remains unanswered.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models ↗Wornow, Thapa, Steinberg, Fries and Shah; Stanford. Versioned EHRSHOT paper with prediction task definitions and patient/label counts.
  2. EHRSHOT official benchmark repository ↗Stanford Shah Lab. Code Apache 2.0; dataset uses separate Stanford Research Use Agreement.
  3. Evaluation and mitigation of the limitations of large language models in clinical decision-making ↗Hager, Jungmann, Holland et al.; Nature Medicine. Original MIMIC-CDM study; historical models and restricted disease cohort.
  4. MIMIC-IV-Ext Clinical Decision Making v1.1 ↗Hager, Jungmann and Rueckert; PhysioNet. Official data card; 2,400 patients, disease counts, access and release notes.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →