Benchmark analysis / 4 min read

How many examples does an EHRSHOT result really use?

Separate EHRSHOT’s patients, prediction labels, pretraining population and positive/negative few-shot sampling.

The short answer

EHRSHOT contains patients, visits, coded events and prediction labels. All are legitimate counts, but they answer different questions. Its few-shot experiments add another distinction between labels used for adaptation and the much larger population used for pretraining. Our analysis follows the denominator from a patient timeline to a scored prediction. This makes it easier to interpret sample efficiency without describing repeated events as independent people or treating a few-shot label as a complete description of the evidence available to a model.

A patient can generate several prediction events

The paper’s task tables separately report unique patients and labels. For example, its length-of-stay test contains more labels than patients. That difference is expected when a person has several eligible events or encounters rather than a single permanent target.

Our deduction is that the unit of prediction and the unit of independence need separate descriptions. A model can be scored on every eligible label while uncertainty analysis accounts for repeated observations from the same person. A result described only by its large label count can make the evidence appear more independent than it actually is.

A shot has a published sampling convention

EHRSHOT’s paper defines its adaptation protocol using positive and negative examples from the training split, with corresponding validation examples for selecting hyperparameters. A stated value of k is therefore not a universal shorthand for k total labeled records. The task and sampling rules supply its meaning.

When comparing a new method, reproduce or explicitly change that convention. Report how rare positives, repeated samples and multiclass or multilabel tasks are handled. Our interpretation is that a sample-efficiency claim requires both a performance result and a complete accounting of which labels were available during fitting and selection.

Pretraining and adaptation are different evidence sources

The released foundation-model baseline was pretrained on a much larger source population before adaptation to the benchmark tasks. The supervised baseline and pretrained representation therefore do not arrive at adaptation with the same prior exposure. That is part of the question the benchmark is designed to study.

A useful report distinguishes pretraining patients, adaptation labels and held-out test patients. It should neither erase the benefit of pretraining nor imply that a model learned everything from a handful of examples. Our original reading is that EHRSHOT helps ask how effectively prior representations can be reused, with the answer conditional on the task and adaptation procedure.

Prevalence belongs beside precision–recall

The test prevalence differs substantially between tasks such as ICU transfer and length of stay. EHRSHOT reports AUROC and AUPRC, which provide complementary views of discrimination. AUPRC should be interpreted with the frequency of positive labels in the relevant cohort.

Our proposed task card places the positive count, total labels and unique patient count beside the metric. This does not replace calibration or clinical utility analysis, but it prevents comparisons that ignore the composition of the test set. A high score on a common label and the same score on a rare label need different practical investigation.

Respect the input boundary and dataset agreement

EHRSHOT’s released benchmark is based on structured codes and excludes free-text notes and images. Its chest-X-ray-finding task predicts labels from prior coded records; it is not an image-reading benchmark. A system that performs well here has not thereby demonstrated note summarization or radiological image interpretation.

The official repository also separates software licensing from the dataset research agreement. For reproduction, identify the permitted data source and the exact preprocessing path. Our site uses the paper’s aggregate counts and task definitions without reproducing patient timelines or inventing exact performance values from plotted curves. The result is a clear description of what evidence exists and what a new experiment must still supply.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models ↗Wornow, Thapa, Steinberg, Fries and Shah; Stanford. Versioned EHRSHOT paper with prediction task definitions and patient/label counts.
  2. EHRSHOT official benchmark repository ↗Stanford Shah Lab. Code Apache 2.0; dataset uses separate Stanford Research Use Agreement.
  3. EHRSHOT Data Set License 1.0 ↗Stanford University. Stanford’s dataset-specific license is separate from the Apache 2.0 software license.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →