Topic 09 · Research & knowledge · Deep dive
The evidence gap: a field that publishes faster than it proves
Imaging is one of the most research-intensive fields in medicine — medical-AI papers alone run near 50,000 a year, and radiology is their single biggest home. Yet the knowledge underneath is thinner than the volume suggests. Radiomics studies score around a quarter of the available quality points. Deep-learning tools routinely beat clinicians on paper but have been tested in only a handful of randomised trials. This closing topic is about the distance between what imaging research publishes and what it has actually proven.
By raw volume, imaging research is booming — and AI is the accelerant. The counts below use different search definitions (imaging-specific vs all medical AI), so they aren't directly comparable, but every series points the same way: near-vertical.
A tenfold decade
Across five leading AI conferences, papers rose tenfold from 1,206 (2014) to 12,026 (2024). Medical imaging rode the same wave — radiology consistently ranks as the top specialty for mature AI publications year after year.
Volume ≠ maturity
In one structured review, only 1.8% of 2021's medical-AI papers were judged "mature" (a deployable model or systematic review) — down from 7.7% in 2020 as raw output ballooned. The denominator grew faster than the signal.
The LLM pivot
Since ChatGPT's late-2022 arrival, research shifted from image-based deep learning toward large language models. Imaging remains a top LLM-in-medicine specialty, but text-based work now competes for attention and funding.
Radiomics — mining quantitative features from images that the eye can't see — became a research industry in the 2010s. Its own quality-audit tools tell an uncomfortable story: most studies are retrospective, single-centre, and never externally validated, which is exactly why so little of it has reached the clinic.
The single most important evidence gap in imaging research is between a model's reported accuracy and a trial that shows it helps patients. The landmark BMJ systematic review of deep-learning-vs-clinician studies quantified how wide that gap is — and it has narrowed only slowly since.
The reporting-standards fix
The field's response is a wave of guidelines — CLAIM (imaging AI), TRIPOD+AI, SPIRIT-AI/CONSORT-AI (trials), METRICS (radiomics). They raise reporting quality, but adherence is still partial and voluntary.
Open data changed the game
Public datasets — TCIA, MIMIC-CXR, CheXpert, the RSNA/MICCAI challenges — let thousands train models without collecting their own images. That democratised research, but also concentrated the field on a few, non-representative datasets (Topic 05's equity gap, upstream).
Where the loop closes
This is why Topic 04 found only a fraction of cleared AI is validated to a high level, and why Topic 08's error rate persists: the evidence base was never built to prove clinical benefit, only accuracy. Fixing output quality starts here, in how the field does research.
A field measured by its output, not its proof
Radiology has always been an early technology adopter, and its research output shows it: medical-AI publications rose from under four thousand in 2010 to nearly fifty thousand in 2024, with imaging the single largest specialty inside that total. But volume is a treacherous measure of knowledge. When one structured review applied a maturity filter, the share of medical-AI papers representing a deployable model or a systematic review fell to under 2% in 2021 — not because the science got worse, but because the denominator exploded. The field publishes at a rate that would suggest imaging is among the best-evidenced areas of medicine. On the metrics that matter for patients, it is among the thinnest relative to its output.
Radiomics is the cautionary tale
No sub-field illustrates the gap better than radiomics, which promised to turn every scan into a dense table of quantitative biomarkers. Thousands of studies followed — and when they were audited against the field's own Radiomics Quality Score, the average study earned about a quarter of the available points: 9.4 out of 36. Prospective designs appeared in under 4% of studies, test–retest stability analysis in 6.5%, open data in under 4%, and not one study in the audited oncologic cohort performed a phantom or cost-effectiveness analysis. The deeper irony is that the quality tool itself proved hard to apply reproducibly, with inter-rater agreement as low as 0.30 — the instruments built to police radiomics quality inherited the reproducibility problem they were meant to diagnose.
The distance between accuracy and benefit
The central finding of this topic is that imaging research overwhelmingly measures the wrong thing — diagnostic accuracy on a held-out dataset — and only rarely measures whether using the tool changes what happens to a patient. The landmark BMJ review made this concrete: of more than eighty deep-learning-versus-clinician studies published up to mid-2019, three-quarters claimed performance matching or beating experts, yet only ten registered randomised trials existed, most non-randomised studies were retrospective and at high risk of bias, and code and data were usually unavailable. An accuracy figure and a demonstrated clinical benefit are different claims requiring different evidence, and the field has produced a great deal of the former while barely beginning the latter. This is not unique to imaging, but imaging's sheer publication volume makes the imbalance especially visible.
Where the whole series comes together
This closing topic is, in a sense, the foundation under all the others. Topic 04 found that only a fraction of cleared imaging AI carries high-level evidence — because, as this topic shows, that evidence was rarely generated. Topic 08's stubborn error rate persists partly because research optimised for benchmark accuracy rather than real-world reading conditions. Topic 05's equity gap is reproduced in a research base built on a handful of non-representative public datasets. The encouraging turn is that the field has diagnosed its own problem: reporting standards like CLAIM, TRIPOD+AI and CONSORT-AI, the METRICS successor to RQS, and a slowly rising share of externally validated and prospective work are all pushing toward proof rather than mere publication. The knowledge base of imaging is enormous, growing, and — for now — still catching up to the confidence with which its findings are quoted. Reading it honestly, caveats intact, is the whole purpose of a project like this one.
On the data. Publication counts depend heavily on search terms and databases; the 3,891→49,739 series (2010→2024) is one consistent search, but the intermediate 2020–2021 reference points come from differently-defined searches and are approximate, shown for shape only. The radiomics RQS figures are from specific study cohorts (chiefly Park et al.'s oncologic sample) and vary across sub-fields and reviews. The Nagendran BMJ review covers studies only to June 2019; the RCT count has risen since, though it remains low relative to publication volume. The "translation funnel" chart is explicitly a conceptual illustration of a well-documented attrition pattern, not a measured cascade — no single source provides clean end-to-end proportions, so none is claimed. Figures span 2018–2026 vintages as labelled.