# Hospital Bench > Independent EHRSHOT and MIMIC-CDM analysis: longitudinal prediction, few-shot denominators, diagnostic simulation and the limits of hospital benchmark transfer. EHRSHOT tests prediction from longitudinal coded records. MIMIC-CDM tests diagnostic decisions from selected retrospective cases. This independent analysis puts their inputs, denominators and information settings side by side. Explore what a prediction label represents, why a selected disease cohort cannot establish general emergency-care performance, and where published evidence stops. We report aggregate metadata and historical study results, with no patient material, new model runs or implied affiliation with the benchmark authors. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [EHRSHOT](https://hospitalbench.ai/benchmarks/ehrshot/): The denominator is a prediction time, not just a patient. Source version: 2023 paper, arXiv v2. - [MIMIC-IV-Ext Clinical Decision Making](https://hospitalbench.ai/benchmarks/mimic-cdm/): Requesting evidence is harder than receiving the finished case. Source version: 2024 paper; dataset metadata v1.1. ## Original analyses - [How many examples does an EHRSHOT result really use?](https://hospitalbench.ai/guides/ehrshot-patients-labels-few-shot/): Separate EHRSHOT’s patients, prediction labels, pretraining population and positive/negative few-shot sampling. - [What does MIMIC-CDM reveal about gathering diagnostic evidence?](https://hospitalbench.ai/guides/mimic-cdm-information-and-diagnostic-accuracy/): Interpret interactive and full-information diagnostic results in the restricted four-condition MIMIC-CDM cohort. - [When should a hospital use EHRSHOT or MIMIC-CDM?](https://hospitalbench.ai/guides/ehrshot-versus-mimic-cdm/): Choose between structured-record prediction and diagnostic simulation by matching the benchmark’s inputs, targets and evidence limits. ## Inspect the evidence - [Evidence JSON](https://hospitalbench.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://hospitalbench.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://hospitalbench.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://hospitalbench.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28