{"publication":"Hospital Bench","url":"https://hospitalbench.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"ehrshot","name":"EHRSHOT","shortName":"EHRSHOT","version":"2023 paper, arXiv v2","creators":"Wornow, Thapa, Steinberg, Fries and Shah; Stanford","paperDate":"2023-11-02","headline":"The denominator is a prediction time, not just a patient.","summary":"EHRSHOT evaluates adaptation of models trained on structured medical records to new prediction tasks with few labels. It offers an unusually useful distinction between unique patients, prediction labels and the size of the broader pretraining population. We analyze the released benchmark’s task inventory and sampling protocol, rather than turn its plots into unsupported numerical rankings. The original contribution here is a denominator map: what one training shot means, which population a test label represents and why event counts cannot be read as independent people. This is structured-record prediction, not a conversational clinical examination.","task":{"input":"A patient’s structured coded record up to a task-specific prediction time.","output":"A binary, multiclass or multilabel prediction.","unit":"A labeled prediction event; patients can contribute multiple labels.","setting":"Few-shot adaptation with fixed patient splits and held-out task test labels."},"dataOrigin":"Deidentified Stanford STARR structured records; released benchmark excludes free-text clinical notes and images.","facts":[{"label":"Patients","value":"6,739","detail":"2,295 train; 2,232 validation; 2,212 test.","sourceIds":["ehr"],"locator":"Appendix Table 4"},{"label":"Clinical events","value":"41,661,637","detail":"Repository-reported coded event count.","sourceIds":["ehr-code"],"locator":"README"},{"label":"Visits","value":"921,499","detail":"Repository and paper report this encounter count.","sourceIds":["ehr","ehr-code"],"locator":"Section 3; README"},{"label":"Prediction tasks","value":"15","detail":"Nine binary, five multiclass and one multilabel task.","sourceIds":["ehr"],"locator":"Section 3.2"},{"label":"Few-shot protocol","value":"k positives + k negatives","detail":"Also uses k positive and k negative validation labels for tuning; interpret by task definition.","sourceIds":["ehr"],"locator":"Section 5"}],"metric":{"name":"AUROC and AUPRC","description":"Discrimination metrics on held-out task labels; the paper reports task and category-level curves as label availability changes.","formula":"AUROC summarizes ranking across thresholds; AUPRC summarizes precision–recall across thresholds","direction":"Higher is better within a fixed task, cohort and metric implementation.","comparability":"AUPRC depends on prevalence. Category macro-averages, label-level results and patient counts are not interchangeable.","sourceIds":["ehr"]},"workflow":[{"label":"Fix patient partitions","detail":"Keep the released train, validation and test patient assignments.","sourceIds":["ehr"]},{"label":"Define prediction time","detail":"Use only the coded history available before the task’s label time.","sourceIds":["ehr"]},{"label":"Sample training and validation labels","detail":"Vary k under the published sampling procedure.","sourceIds":["ehr"]},{"label":"Evaluate held-out labels","detail":"Compare CLMBR-T-base with count-based baselines using AUROC and AUPRC.","sourceIds":["ehr"]}],"slices":[{"label":"Operational outcomes","value":3,"unit":"tasks","detail":"Length of stay, readmission and ICU transfer.","sourceIds":["ehr"]},{"label":"Lab results","value":5,"unit":"tasks","detail":"Multiclass laboratory-value prediction.","sourceIds":["ehr"]},{"label":"New diagnoses","value":6,"unit":"tasks","detail":"Prediction of newly assigned diagnoses.","sourceIds":["ehr"]},{"label":"Chest X-ray findings","value":1,"unit":"tasks","detail":"Fourteen-label prediction from prior codes; not direct image interpretation.","sourceIds":["ehr"]}],"sliceTitle":"Four prediction families","sliceNote":"Task counts sum to 15. Findings prediction uses structured inputs rather than a chest image.","results":[],"analysis":[{"heading":"Labels can outnumber people","evidence":"The length-of-stay test has 2,195 labels from 1,238 patients.","interpretation":"A label-level metric is not a trial with 2,195 independent patients. Analyses of uncertainty should respect repeated observations.","sourceIds":["ehr"]},{"heading":"A shot has a precise sampling meaning","evidence":"The paper uses k positive and k negative training examples plus matching validation examples.","interpretation":"Calling this k total labeled examples understates the adaptation evidence. Report the sampling convention as well as k.","sourceIds":["ehr"]},{"heading":"Rare outcomes change metric context","evidence":"The ICU-transfer test has 85 positive labels among 2,037; length of stay has 552 among 2,195.","interpretation":"Similar AUROC values can sit in very different precision–recall settings. Compare AUPRC within its prevalence context.","sourceIds":["ehr"]},{"heading":"Prediction is not text understanding","evidence":"The released model and benchmark use coded records and remove clinical notes.","interpretation":"Strong performance here does not directly establish summarization, dialogue or note-reading quality.","sourceIds":["ehr"]}],"limitations":[{"title":"Single-institution cohort","detail":"Structured records and coding practices come from Stanford; external transportability remains a separate question.","sourceIds":["ehr"]},{"title":"Selected age range","detail":"The released cohort excludes ages below 19 and above 88.","sourceIds":["ehr"]},{"title":"No exact plot readings","detail":"The paper’s model performance is primarily shown as curves. We do not invent numerical result rows from visual interpolation.","sourceIds":["ehr"]},{"title":"Sample efficiency is conditional","detail":"Benefits depend on task, available labels, representation and adaptation procedure.","sourceIds":["ehr"]},{"title":"Lab-class documentation conflict","detail":"Paper §3.2 describes five-way lab tasks, while Appendix Table 7 and the official release specify four-way classification. The atlas follows the task table and released label vocabulary.","sourceIds":["ehr","ehr-code"]}],"access":{"status":"Dataset access under research agreement","license":"Code Apache 2.0; dataset EHRSHOT Data Set License 1.0 (Stanford University).","restrictions":"The dataset agreement is separate from the repository license and restricts permitted use; consult the current agreement.","url":"https://github.com/som-shahlab/ehrshot-benchmark/blob/main/DUA.md","sourceIds":["ehr-code","ehr-dua"]},"sourceIds":["ehr","ehr-code","ehr-dua"]},{"slug":"mimic-cdm","name":"MIMIC-IV-Ext Clinical Decision Making","shortName":"MIMIC-CDM","version":"2024 paper; dataset metadata v1.1","creators":"Hager, Jungmann, Holland et al.","paperDate":"2024-07-04","headline":"Requesting evidence is harder than receiving the finished case.","summary":"MIMIC-CDM reconstructs diagnostic tasks from records for four abdominal conditions. Its two information settings make a useful contrast: an interactive model requests findings, while a full-information model receives a prepared case. We examine the implications of that contrast without treating a four-condition sample as a general emergency-department population. The paper’s per-class accuracy also matters: its denominator is patients with the target disease, not every patient for whom a positive diagnosis might be made. Our independent analysis connects the selection rules, disease balance and result labels to the claims a hospital reader can reasonably make.","task":{"input":"History of present illness, followed by requested findings; full-information variant supplies a prepared case.","output":"Diagnosis and treatment proposal.","unit":"One patient case within a selected abdominal pathology.","setting":"Retrospective clinical decision simulation; historical open-weight model configurations."},"dataOrigin":"2,400 deidentified MIMIC-IV v2.2 patients selected for appendicitis, cholecystitis, diverticulitis or pancreatitis; all required modalities must be present.","facts":[{"label":"Patient cases","value":"2,400","detail":"Four selected abdominal pathologies, not an all-comer emergency cohort.","sourceIds":["cdm-data"],"locator":"Methods"},{"label":"Source database","value":"MIMIC-IV v2.2","detail":"Hospital module and clinical notes underpin the derivative.","sourceIds":["cdm-data"],"locator":"Background"},{"label":"Diagnostic families","value":"4","detail":"Cases with multiple selected target pathologies are excluded.","sourceIds":["cdm-data"],"locator":"Methods"},{"label":"Information settings","value":"Interactive / full information","detail":"The two settings offer different amounts and timing of evidence.","sourceIds":["cdm"],"locator":"Results; Methods"},{"label":"Dataset release","value":"v1.1","detail":"July 2024 release adds pathology IDs and corrects procedure columns.","sourceIds":["cdm-data"],"locator":"Release Notes"}],"metric":{"name":"Per-class diagnostic accuracy","description":"Correct diagnoses divided by the number of evaluated cases with that disease; headline means average across the four pathology classes.","formula":"Accuracy_c = correct diagnoses in class c / evaluated patients in class c; macro = Σ accuracy_c / 4","direction":"Higher is better within a fixed setting and cohort.","comparability":"This restricted cohort cannot supply general emergency-care specificity or positive predictive value. Full-information and interactive results are separate settings.","sourceIds":["cdm"]},"workflow":[{"label":"Select the eligible cohort","detail":"Require a target primary diagnosis and the needed clinical modalities; remove diagnosis leakage from selected text.","sourceIds":["cdm-data"]},{"label":"Reveal information","detail":"Provide findings on request, or use the full-information comparator.","sourceIds":["cdm","cdm-data"]},{"label":"Collect the conclusion","detail":"Save the diagnosis, proposed treatment and information-gathering path.","sourceIds":["cdm"]},{"label":"Score by pathology","detail":"Assess the final diagnosis and other behaviors against references, with pathology-specific denominators.","sourceIds":["cdm"]}],"slices":[{"label":"Appendicitis","value":957,"unit":"patients","detail":"Selected cohort.","sourceIds":["cdm-data"]},{"label":"Cholecystitis","value":648,"unit":"patients","detail":"Selected cohort.","sourceIds":["cdm-data"]},{"label":"Diverticulitis","value":257,"unit":"patients","detail":"Selected cohort.","sourceIds":["cdm-data"]},{"label":"Pancreatitis","value":538,"unit":"patients","detail":"Selected cohort.","sourceIds":["cdm-data"]}],"sliceTitle":"The four-condition cohort","sliceNote":"Published cohort counts; do not interpret as population prevalence.","results":[{"id":"cdm-interactive","title":"Historical interactive decision-making results","metric":"Mean of per-pathology diagnostic accuracy","unit":"%","lower":0,"upper":100,"scope":"Hager et al., July 2024; models gather information in the MIMIC-CDM simulation.","sourceIds":["cdm"],"locator":"Results accompanying Figure 3","rows":[{"label":"Llama 2 Chat","value":45.5,"display":"45.5%","detail":"Historical paper configuration."},{"label":"OASST","value":54.9,"display":"54.9%","detail":"Historical paper configuration."},{"label":"WizardLM","value":53.9,"display":"53.9%","detail":"Historical paper configuration."}],"note":"Paper-reported means across pathologies; no current-frontier inference. Do not combine with full-information rows."},{"id":"cdm-full","title":"Historical full-information comparator","metric":"Mean of per-pathology diagnostic accuracy","unit":"%","lower":0,"upper":100,"scope":"Same paper, MIMIC-CDM-FI setting; relevant case information is provided up front.","sourceIds":["cdm"],"locator":"Results accompanying Figure 3","rows":[{"label":"Llama 2 Chat","value":58.8,"display":"58.8%","detail":"Full-information setting."},{"label":"OASST","value":67.8,"display":"67.8%","detail":"Full-information setting."},{"label":"WizardLM","value":65.1,"display":"65.1%","detail":"Full-information setting."}],"note":"These are the paper’s comparison values, not new reruns or an estimate of benefit from a particular clinical intervention."}],"analysis":[{"heading":"The missing step is evidence acquisition","evidence":"All three selected models have lower mean accuracy in the interactive setting.","interpretation":"An all-information test leaves the ability to choose useful next information largely unmeasured. The observed gap is a setting contrast, not a causal attribution to one component.","sourceIds":["cdm"]},{"heading":"Cohort balance changes weighting","evidence":"Appendicitis has 957 cases; diverticulitis has 257.","interpretation":"A pooled case accuracy would weight diseases differently from the paper’s mean of per-class accuracies. Preserve the aggregation definition.","sourceIds":["cdm-data","cdm"]},{"heading":"A positive-case sample has a boundary","evidence":"Cases are selected for four known primary pathologies.","interpretation":"These results cannot supply the false-positive burden for a hospital seeing the complete differential diagnosis.","sourceIds":["cdm"]},{"heading":"Rich records still have missing dynamics","evidence":"The data card retains only the first laboratory instance for each test.","interpretation":"This configuration is poorly suited to claims about longitudinal response to treatment. A larger record volume does not restore omitted temporal behavior.","sourceIds":["cdm-data"]}],"limitations":[{"title":"Restricted disease set","detail":"The selected four conditions do not represent all abdominal pain or the full hospital diagnostic workload.","sourceIds":["cdm","cdm-data"]},{"title":"Retrospective selection","detail":"Requiring available imaging, tests and physical examination selects cases with particular documentation and care pathways.","sourceIds":["cdm-data"]},{"title":"Historical tested models","detail":"The 2024 comparison cannot settle performance of later model generations or different agent systems.","sourceIds":["cdm"]},{"title":"Access constraints","detail":"Patient-derived data remains credentialed and governed by PhysioNet terms.","sourceIds":["cdm-data"]}],"access":{"status":"Credentialed PhysioNet access","license":"PhysioNet Credentialed Health Data License 1.5.0; code MIT.","restrictions":"Credentialing, required research training and a signed data use agreement are required for the files.","url":"https://physionet.org/content/mimic-iv-ext-cdm/1.1/","sourceIds":["cdm-data","cdm-code"]},"sourceIds":["cdm","cdm-data","cdm-code"]}],"explorer":{"kind":"coverage","title":"Prediction or diagnostic action?","intro":"Compare EHRSHOT’s prediction targets with MIMIC-CDM’s information settings. Each row states the actual unit being scored.","caution":"Different cohorts, outputs and metrics make a combined hospital-accuracy number inappropriate.","sourceIds":["ehr","cdm"],"parameters":[],"rows":[{"label":"Length of stay","category":"Operational prediction","input":"Coded record up to prediction time","output":"Binary outcome","metric":"AUROC / AUPRC","constraint":"Repeated labels can come from the same patient.","benchmarkSlug":"ehrshot","sourceIds":["ehr"]},{"label":"Thirty-day readmission","category":"Operational prediction","input":"Coded history before discharge","output":"Binary outcome","metric":"AUROC / AUPRC","constraint":"Only events observed in the source system can form labels.","benchmarkSlug":"ehrshot","sourceIds":["ehr"]},{"label":"ICU transfer","category":"Operational prediction","input":"Coded record during admission","output":"Binary outcome","metric":"AUROC / AUPRC","constraint":"Rare positive labels change precision–recall context.","benchmarkSlug":"ehrshot","sourceIds":["ehr"]},{"label":"Future lab result","category":"Clinical prediction","input":"Prior structured record","output":"Four-way class","metric":"AUROC / AUPRC aggregation","constraint":"Appendix Table 7 and the release describe four classes; paper §3.2 prose says five. This atlas follows the task table and release.","benchmarkSlug":"ehrshot","sourceIds":["ehr","ehr-code"]},{"label":"New diagnosis","category":"Clinical prediction","input":"Prior coded timeline","output":"Future diagnosis label","metric":"AUROC / AUPRC","constraint":"Coding and prediction horizon shape the target.","benchmarkSlug":"ehrshot","sourceIds":["ehr"]},{"label":"Chest X-ray findings","category":"Clinical prediction","input":"Prior coded record","output":"Fourteen-label finding prediction","metric":"AUROC / AUPRC","constraint":"The model does not receive the image.","benchmarkSlug":"ehrshot","sourceIds":["ehr"]},{"label":"Interactive abdominal diagnosis","category":"Diagnostic simulation","input":"History, then requested findings","output":"Diagnosis and treatment proposal","metric":"Per-class diagnostic accuracy","constraint":"Four selected diseases do not establish general specificity.","benchmarkSlug":"mimic-cdm","sourceIds":["cdm"]},{"label":"Full-information diagnosis","category":"Diagnostic simulation","input":"Prepared case evidence up front","output":"Diagnosis","metric":"Per-class diagnostic accuracy","constraint":"Removes the need to choose the next evidence request.","benchmarkSlug":"mimic-cdm","sourceIds":["cdm"]}]},"references":[{"id":"ehr","title":"EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models","organization":"Wornow, Thapa, Steinberg, Fries and Shah; Stanford","url":"https://arxiv.org/html/2307.02028v2","note":"Versioned EHRSHOT paper with prediction task definitions and patient/label counts.","locator":"Tables 3–4; Section 5; Appendix C.3","version":"arXiv v2, 2 November 2023"},{"id":"ehr-code","title":"EHRSHOT official benchmark repository","organization":"Stanford Shah Lab","url":"https://github.com/som-shahlab/ehrshot-benchmark","note":"Code Apache 2.0; dataset uses separate Stanford Research Use Agreement.","locator":"README, LICENSE and DUA.md","version":"Accessed 28 September 2026"},{"id":"ehr-dua","title":"EHRSHOT Data Set License 1.0","organization":"Stanford University","url":"https://github.com/som-shahlab/ehrshot-benchmark/blob/main/DUA.md","note":"Stanford’s dataset-specific license is separate from the Apache 2.0 software license.","locator":"Research use, redistribution and clinical-use conditions","version":"Accessed 28 September 2026"},{"id":"cdm","title":"Evaluation and mitigation of the limitations of large language models in clinical decision-making","organization":"Hager, Jungmann, Holland et al.; Nature Medicine","url":"https://www.nature.com/articles/s41591-024-03097-1","note":"Original MIMIC-CDM study; historical models and restricted disease cohort.","locator":"Figure 3; Metrics and statistical analysis; Methods","version":"4 July 2024"},{"id":"cdm-data","title":"MIMIC-IV-Ext Clinical Decision Making v1.1","organization":"Hager, Jungmann and Rueckert; PhysioNet","url":"https://physionet.org/content/mimic-iv-ext-cdm/1.1/","note":"Official data card; 2,400 patients, disease counts, access and release notes.","locator":"Methods; Usage Notes; Access; Release Notes","version":"8 July 2024, v1.1"},{"id":"cdm-code","title":"MIMIC Clinical Decision Making Framework","organization":"Paul Hager and collaborators","url":"https://github.com/paulhager/MIMIC-Clinical-Decision-Making-Framework","note":"MIT code license; source data remains credentialed.","locator":"README and LICENSE","version":"Accessed 28 September 2026"}]}