Independent benchmark analysis / 2023 paper, arXiv v2

EHRSHOT

The denominator is a prediction time, not just a patient.

EHRSHOT evaluates adaptation of models trained on structured medical records to new prediction tasks with few labels. It offers an unusually useful distinction between unique patients, prediction labels and the size of the broader pretraining population. We analyze the released benchmark’s task inventory and sampling protocol, rather than turn its plots into unsupported numerical rankings. The original contribution here is a denominator map: what one training shot means, which population a test label represents and why event counts cannot be read as independent people. This is structured-record prediction, not a conversational clinical examination.

01 / What is being tested?

The task, before the score.

input
A patient’s structured coded record up to a task-specific prediction time.
output
A binary, multiclass or multilabel prediction.
unit
A labeled prediction event; patients can contribute multiple labels.
setting
Few-shot adaptation with fixed patient splits and held-out task test labels.

Data origin. Deidentified Stanford STARR structured records; released benchmark excludes free-text clinical notes and images. [1][2][3]

Patients
6,739

2,295 train; 2,232 validation; 2,212 test.

Appendix Table 4 [1]
Clinical events
41,661,637

Repository-reported coded event count.

README [2]
Visits
921,499

Repository and paper report this encounter count.

Section 3; README [1][2]
Prediction tasks
15

Nine binary, five multiclass and one multilabel task.

Section 3.2 [1]
Few-shot protocol
k positives + k negatives

Also uses k positive and k negative validation labels for tuning; interpret by task definition.

Section 5 [1]
  1. 01

    Fix patient partitions

    Keep the released train, validation and test patient assignments.

    [1]
  2. 02

    Define prediction time

    Use only the coded history available before the task’s label time.

    [1]
  3. 03

    Sample training and validation labels

    Vary k under the published sampling procedure.

    [1]
  4. 04

    Evaluate held-out labels

    Compare CLMBR-T-base with count-based baselines using AUROC and AUPRC.

    [1]

Dataset anatomy

Four prediction families

Operational outcomes

Length of stay, readmission and ICU transfer.

3 tasks[1]
Lab results

Multiclass laboratory-value prediction.

5 tasks[1]
New diagnoses

Prediction of newly assigned diagnoses.

6 tasks[1]
Chest X-ray findings

Fourteen-label prediction from prior codes; not direct image interpretation.

1 tasks[1]

Task counts sum to 15. Findings prediction uses structured inputs rather than a chest image. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

AUROC and AUPRC

Higher is better within a fixed task, cohort and metric implementation.

Discrimination metrics on held-out task labels; the paper reports task and category-level curves as label availability changes.

Scoring definition

AUROC summarizes ranking across thresholds; AUPRC summarizes precision–recall across thresholds

AUPRC depends on prevalence. Category macro-averages, label-level results and patient counts are not interchangeable. [1]

03 / Measured evidence

Results, with their conditions attached.

This dossier does not reproduce a model ranking. The cited primary sources are the authority for experimental results; the analysis here focuses on task construction, measurement, and interpretation.

04 / Our original analysis

What follows from the design?

01

Labels can outnumber people

Published evidence

The length-of-stay test has 2,195 labels from 1,238 patients. [1]

Our interpretation

A label-level metric is not a trial with 2,195 independent patients. Analyses of uncertainty should respect repeated observations.

02

A shot has a precise sampling meaning

Published evidence

The paper uses k positive and k negative training examples plus matching validation examples. [1]

Our interpretation

Calling this k total labeled examples understates the adaptation evidence. Report the sampling convention as well as k.

03

Rare outcomes change metric context

Published evidence

The ICU-transfer test has 85 positive labels among 2,037; length of stay has 552 among 2,195. [1]

Our interpretation

Similar AUROC values can sit in very different precision–recall settings. Compare AUPRC within its prevalence context.

04

Prediction is not text understanding

Published evidence

The released model and benchmark use coded records and remove clinical notes. [1]

Our interpretation

Strong performance here does not directly establish summarization, dialogue or note-reading quality.

05 / Scope of the evidence

Where this benchmark stops.

Single-institution cohort

Structured records and coding practices come from Stanford; external transportability remains a separate question. [1]

Selected age range

The released cohort excludes ages below 19 and above 88. [1]

No exact plot readings

The paper’s model performance is primarily shown as curves. We do not invent numerical result rows from visual interpolation. [1]

Sample efficiency is conditional

Benefits depend on task, available labels, representation and adaptation procedure. [1]

Lab-class documentation conflict

Paper §3.2 describes five-way lab tasks, while Appendix Table 7 and the official release specify four-way classification. The atlas follows the task table and released label vocabulary. [1][2]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Dataset access under research agreement
License
Code Apache 2.0; dataset EHRSHOT Data Set License 1.0 (Stanford University).
Conditions
The dataset agreement is separate from the repository license and restricts permitted use; consult the current agreement.
[2][3]

Evidence trail

Read the originals.

  1. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models ↗

    Wornow, Thapa, Steinberg, Fries and Shah; Stanford. Versioned EHRSHOT paper with prediction task definitions and patient/label counts.

  2. EHRSHOT official benchmark repository ↗

    Stanford Shah Lab. Code Apache 2.0; dataset uses separate Stanford Research Use Agreement.

  3. EHRSHOT Data Set License 1.0 ↗

    Stanford University. Stanford’s dataset-specific license is separate from the Apache 2.0 software license.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗