Building the benchmark is harder than running it: thirteen usable cases for retrospective validation of transcriptomic drug repurposing, and the decisions that would have decided the answer
Description
Transcriptomic drug repurposing rests on a claim that has never been tested at the only endpoint that matters: that reversing a disease expression signature identifies compounds that go on to work in patients. The obvious test is retrospective. Take drugs that were approved for a second indication, query a frozen connectivity pipeline with disease transcriptomics collected independently of the drug, and see where the approved drug ranks.
This report scopes that benchmark rather than running it, and finds that the scoping is where the difficulty lives. Of roughly 72 verified FDA repurposing approvals since 2000, thirteen survive a pre-registered set of inclusion gates, twelve of them cleanly. Getting there required resolving problems that would each have silently changed the result: two of the four strongest candidates are circular, because the transcriptomic data that would test the pipeline is the data that generated the repurposing hypothesis; compound identity by InChIKey skeleton fails in both directions, merging stereoisomers that differ and splitting tautomers that do not; and one pre-registered gate, the requirement that data predate the hypothesis, turns out to be close to structurally unsatisfiable, because public expression repositories postdate most of the repurposing hypotheses that reached approval.
Three properties of the scoring metric are also reported. The self-target control proposed in earlier work does not port to connectivity scoring, because swapping a query's up and down gene sets flips the score's sign exactly by construction, so the naive port is an identity rather than a test. Transcriptional activity across the reference library is not concentrated (Gini 0.274 against 0.333 for a uniform distribution), which removes one way the benchmark could have failed before starting, though 23.2% of the library has no quality-passing signature at all. And the connectivity score returns exactly zero for 44% of compounds, creating a tied block of roughly 11,653 whose tie-breaking convention alone moves a drug's percentile from 28 to 72, a larger swing than any effect the benchmark could detect.
No benchmark result is reported here, and deliberately so. The control that determines whether the metric responds to disease content at all has not been run, and reporting ranks before it would repeat an error this program has already documented once.
Other
This report was written by the author in collaboration with Claude, an AI system made by Anthropic. The analysis it describes was carried out the same way: the author set the questions, the inclusion gates and the pre-registered criteria, and adjudicated every inclusion and exclusion; the model wrote and ran the verification code, and drafted the text. Section 11 records the errors that arose in that process and how each was caught, including several produced by the model and several by the author, because the reliability of the method is part of what the report is about.
Files
L1000_repurposing_benchmark_v2_code_and_data_2026-10-01.zip
Files
(486.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:716b4b2bf6a07c577b8c45ef0ac77868
|
303.7 kB | Preview Download |
|
md5:2a41596e97b7365e6121f9b7ce20f4c5
|
182.4 kB | Preview Download |
Additional details
Related works
- Cites
- Publication: 10.5281/zenodo.23047943 (DOI)
- Is supplemented by
- Software: https://github.com/glenritschel/l1000-repurposing-benchmark (URL)