Labels can outnumber people
The length-of-stay test has 2,195 labels from 1,238 patients. [1]
A label-level metric is not a trial with 2,195 independent patients. Analyses of uncertainty should respect repeated observations.
Independent benchmark analysis / 2023 paper, arXiv v2
EHRSHOT evaluates adaptation of models trained on structured medical records to new prediction tasks with few labels. It offers an unusually useful distinction between unique patients, prediction labels and the size of the broader pretraining population. We analyze the released benchmark’s task inventory and sampling protocol, rather than turn its plots into unsupported numerical rankings. The original contribution here is a denominator map: what one training shot means, which population a test label represents and why event counts cannot be read as independent people. This is structured-record prediction, not a conversational clinical examination.
01 / What is being tested?
Data origin. Deidentified Stanford STARR structured records; released benchmark excludes free-text clinical notes and images. [1][2][3]
2,295 train; 2,232 validation; 2,212 test.
Appendix Table 4 [1]Repository-reported coded event count.
README [2]Nine binary, five multiclass and one multilabel task.
Section 3.2 [1]Also uses k positive and k negative validation labels for tuning; interpret by task definition.
Section 5 [1]Keep the released train, validation and test patient assignments.
[1]Use only the coded history available before the task’s label time.
[1]Vary k under the published sampling procedure.
[1]Compare CLMBR-T-base with count-based baselines using AUROC and AUPRC.
[1]Dataset anatomy
Length of stay, readmission and ICU transfer.
Multiclass laboratory-value prediction.
Prediction of newly assigned diagnoses.
Fourteen-label prediction from prior codes; not direct image interpretation.
Task counts sum to 15. Findings prediction uses structured inputs rather than a chest image. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better within a fixed task, cohort and metric implementation.
Discrimination metrics on held-out task labels; the paper reports task and category-level curves as label availability changes.
AUROC summarizes ranking across thresholds; AUPRC summarizes precision–recall across thresholds
AUPRC depends on prevalence. Category macro-averages, label-level results and patient counts are not interchangeable. [1]
03 / Measured evidence
04 / Our original analysis
The length-of-stay test has 2,195 labels from 1,238 patients. [1]
A label-level metric is not a trial with 2,195 independent patients. Analyses of uncertainty should respect repeated observations.
The paper uses k positive and k negative training examples plus matching validation examples. [1]
Calling this k total labeled examples understates the adaptation evidence. Report the sampling convention as well as k.
The ICU-transfer test has 85 positive labels among 2,037; length of stay has 552 among 2,195. [1]
Similar AUROC values can sit in very different precision–recall settings. Compare AUPRC within its prevalence context.
The released model and benchmark use coded records and remove clinical notes. [1]
Strong performance here does not directly establish summarization, dialogue or note-reading quality.
05 / Scope of the evidence
Structured records and coding practices come from Stanford; external transportability remains a separate question. [1]
The released cohort excludes ages below 19 and above 88. [1]
The paper’s model performance is primarily shown as curves. We do not invent numerical result rows from visual interpolation. [1]
Benefits depend on task, available labels, representation and adaptation procedure. [1]
Paper §3.2 describes five-way lab tasks, while Appendix Table 7 and the official release specify four-way classification. The atlas follows the task table and released label vocabulary. [1][2]
Evidence trail
Wornow, Thapa, Steinberg, Fries and Shah; Stanford. Versioned EHRSHOT paper with prediction task definitions and patient/label counts.
Stanford Shah Lab. Code Apache 2.0; dataset uses separate Stanford Research Use Agreement.
Stanford University. Stanford’s dataset-specific license is separate from the Apache 2.0 software license.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.