Benchmark dossiers / independent analysis

The benchmark behind the number.

Inspect what each benchmark asks, how its score is produced, and which conclusions its design can support. These are original analyses by Arcophos of the benchmark authors’ work. Any reproduced measurements are labeled with their published source and version.

Dossier01

2023 paper, arXiv v2

EHRSHOT ↗

The denominator is a prediction time, not just a patient.

UnitA labeled prediction event; patients can contribute multiple labels.MeasureAUROC and AUPRC
Dossier02

2024 paper; dataset metadata v1.1

MIMIC-IV-Ext Clinical Decision Making ↗

Requesting evidence is harder than receiving the finished case.

UnitOne patient case within a selected abdominal pathology.MeasurePer-class diagnostic accuracy
What we contribute. Source reconciliation, task and metric interpretation, and tools that make assumptions inspectable. We do not claim to have created these benchmarks or run the reported models.