Topic 09 · Research & knowledge · Deep dive

The evidence gap: a field that publishes faster than it proves

Imaging is one of the most research-intensive fields in medicine — medical-AI papers alone run near 50,000 a year, and radiology is their single biggest home. Yet the knowledge underneath is thinner than the volume suggests. Radiomics studies score around a quarter of the available quality points. Deep-learning tools routinely beat clinicians on paper but have been tested in only a handful of randomised trials. This closing topic is about the distance between what imaging research publishes and what it has actually proven.

~50k
medical-AI articles published in 2024 — up from ~3,900 in 2010; imaging is the single largest specialty within them
RSNA Radiol AI / PubMed, 2025
26%
mean Radiomics Quality Score across oncologic radiomics studies — 9.4 of a possible 36 points
Eur Radiol (Park et al.), 2019
10
registered randomised trials of diagnostic deep-learning vs clinicians in imaging — against 81+ non-randomised studies (to 2019)
Nagendran et al., BMJ 2020
75%
of those deep-learning studies claimed performance at least matching experts — but most were retrospective and at high risk of bias
Nagendran et al., BMJ 2020
The publication explosion

Output has grown faster than almost any field in medicine

By raw volume, imaging research is booming — and AI is the accelerant. The counts below use different search definitions (imaging-specific vs all medical AI), so they aren't directly comparable, but every series points the same way: near-vertical.

Medical-AI publications per year

All medical AI (PubMed, broad search). Imaging is the largest slice.
2024~50k articles
49,739
2021maturity review cohort
~22,000
2020pandemic year
~12,000
2010baseline
3,891
Source: Economic Value of AI in Radiology, Radiology: Artificial Intelligence 2025 (search-defined counts: 3,891 in 2010 → 49,739 in 2024). Intermediate years (2020–2021) are approximate reference points from separate "AI in Healthcare year in review" PubMed searches, which use different terms — shown for shape, not precise comparison.

AI-in-radiology papers per year (earlier era)

A narrower, imaging-only count — the take-off decade.
2016–2017per year
700–800
2007–2008per year, a decade earlier
100–150
Source: Pesapane et al., "AI in medical imaging: threat or opportunity?" (Eur Radiol Exp 2018). MRI and CT together account for >50% of these papers; neuroradiology is the single most-studied subspecialty (~one-third), then MSK, cardiovascular, breast, urogenital, thorax and abdomen (6–9% each). A ~5–6× rise in a decade — and that was before the post-2018 deep-learning surge.
A tenfold decade
Across five leading AI conferences, papers rose tenfold from 1,206 (2014) to 12,026 (2024). Medical imaging rode the same wave — radiology consistently ranks as the top specialty for mature AI publications year after year.
Volume ≠ maturity
In one structured review, only 1.8% of 2021's medical-AI papers were judged "mature" (a deployable model or systematic review) — down from 7.7% in 2020 as raw output ballooned. The denominator grew faster than the signal.
The LLM pivot
Since ChatGPT's late-2022 arrival, research shifted from image-based deep learning toward large language models. Imaging remains a top LLM-in-medicine specialty, but text-based work now competes for attention and funding.
The radiomics reproducibility problem

Thousands of studies, a quarter of the quality

Radiomics — mining quantitative features from images that the eye can't see — became a research industry in the 2010s. Its own quality-audit tools tell an uncomfortable story: most studies are retrospective, single-centre, and never externally validated, which is exactly why so little of it has reached the clinic.

Radiomics Quality Score: where the points are lost

Adherence across oncologic radiomics studies (Park et al.), % meeting each item.
Mean overall RQS9.4 of 36 points
26%
Demonstrated clinical utility
19.5%
Test–retest analysisfeature stability
6.5%
Prospective study design
3.9%
Open science / open data
3.9%
Phantom / cost-effectiveness study
0%
Source: Park et al., Eur Radiol 2019 ("Quality of science and reporting of radiomics… RQS and TRIPOD"). Mean TRIPOD reporting adherence was 57.8%. Studies in clinical journals scored higher and used external validation more often than those in imaging journals. RQS itself has been criticised for low inter-rater reproducibility (ICC 0.30–0.55) — the quality tools have a reproducibility problem too.
Publication vs proof

Beats the clinician on paper; rarely tested like a treatment

The single most important evidence gap in imaging research is between a model's reported accuracy and a trial that shows it helps patients. The landmark BMJ systematic review of deep-learning-vs-clinician studies quantified how wide that gap is — and it has narrowed only slowly since.

Deep learning vs clinicians in imaging: the evidence pyramid

Studies to June 2019 (Nagendran et al., BMJ 2020).
Claimed ≥ expert performance61 of 81 studies
75%
Non-randomised studiesmost retrospective, high bias risk
81
Urged further prospective testing31 of 81
38%
Registered randomised trialsongoing or complete
10
Source: Nagendran et al., BMJ 2020 (systematic review, 2010–June 2019). The middle bars use different denominators (a count of studies vs a percentage) — labelled individually. Key finding: data and code availability were lacking in most studies, and human comparator groups were often small. Deep learning went mainstream ~2014, so some lag is expected — but the imbalance is stark.

The translation funnel

From published model to clinical-grade evidence. Illustrative.
Papers publishedthe flood
~50k/yr
Externally validatedtested outside origin data
minority
Prospectively testedreal-world, forward-looking
few
Randomised controlled trialthe gold standard
rare
Conceptual figure — not a measured cascade. The bar widths illustrate the well-documented shape of attrition from publication to trial (Nagendran 2020; Park 2019; radiomics reviews), not counted proportions of a single cohort. Every source in this topic agrees on the direction; none provides a clean end-to-end percentage, so none is asserted here.
The reporting-standards fix
The field's response is a wave of guidelines — CLAIM (imaging AI), TRIPOD+AI, SPIRIT-AI/CONSORT-AI (trials), METRICS (radiomics). They raise reporting quality, but adherence is still partial and voluntary.
Open data changed the game
Public datasets — TCIA, MIMIC-CXR, CheXpert, the RSNA/MICCAI challenges — let thousands train models without collecting their own images. That democratised research, but also concentrated the field on a few, non-representative datasets (Topic 05's equity gap, upstream).
Where the loop closes
This is why Topic 04 found only a fraction of cleared AI is validated to a high level, and why Topic 08's error rate persists: the evidence base was never built to prove clinical benefit, only accuracy. Fixing output quality starts here, in how the field does research.

A field measured by its output, not its proof

Radiology has always been an early technology adopter, and its research output shows it: medical-AI publications rose from under four thousand in 2010 to nearly fifty thousand in 2024, with imaging the single largest specialty inside that total. But volume is a treacherous measure of knowledge. When one structured review applied a maturity filter, the share of medical-AI papers representing a deployable model or a systematic review fell to under 2% in 2021 — not because the science got worse, but because the denominator exploded. The field publishes at a rate that would suggest imaging is among the best-evidenced areas of medicine. On the metrics that matter for patients, it is among the thinnest relative to its output.

Radiomics is the cautionary tale

No sub-field illustrates the gap better than radiomics, which promised to turn every scan into a dense table of quantitative biomarkers. Thousands of studies followed — and when they were audited against the field's own Radiomics Quality Score, the average study earned about a quarter of the available points: 9.4 out of 36. Prospective designs appeared in under 4% of studies, test–retest stability analysis in 6.5%, open data in under 4%, and not one study in the audited oncologic cohort performed a phantom or cost-effectiveness analysis. The deeper irony is that the quality tool itself proved hard to apply reproducibly, with inter-rater agreement as low as 0.30 — the instruments built to police radiomics quality inherited the reproducibility problem they were meant to diagnose.

The distance between accuracy and benefit

The central finding of this topic is that imaging research overwhelmingly measures the wrong thing — diagnostic accuracy on a held-out dataset — and only rarely measures whether using the tool changes what happens to a patient. The landmark BMJ review made this concrete: of more than eighty deep-learning-versus-clinician studies published up to mid-2019, three-quarters claimed performance matching or beating experts, yet only ten registered randomised trials existed, most non-randomised studies were retrospective and at high risk of bias, and code and data were usually unavailable. An accuracy figure and a demonstrated clinical benefit are different claims requiring different evidence, and the field has produced a great deal of the former while barely beginning the latter. This is not unique to imaging, but imaging's sheer publication volume makes the imbalance especially visible.

Where the whole series comes together

This closing topic is, in a sense, the foundation under all the others. Topic 04 found that only a fraction of cleared imaging AI carries high-level evidence — because, as this topic shows, that evidence was rarely generated. Topic 08's stubborn error rate persists partly because research optimised for benchmark accuracy rather than real-world reading conditions. Topic 05's equity gap is reproduced in a research base built on a handful of non-representative public datasets. The encouraging turn is that the field has diagnosed its own problem: reporting standards like CLAIM, TRIPOD+AI and CONSORT-AI, the METRICS successor to RQS, and a slowly rising share of externally validated and prospective work are all pushing toward proof rather than mere publication. The knowledge base of imaging is enormous, growing, and — for now — still catching up to the confidence with which its findings are quoted. Reading it honestly, caveats intact, is the whole purpose of a project like this one.

On the data. Publication counts depend heavily on search terms and databases; the 3,891→49,739 series (2010→2024) is one consistent search, but the intermediate 2020–2021 reference points come from differently-defined searches and are approximate, shown for shape only. The radiomics RQS figures are from specific study cohorts (chiefly Park et al.'s oncologic sample) and vary across sub-fields and reviews. The Nagendran BMJ review covers studies only to June 2019; the RCT count has risen since, though it remains low relative to publication volume. The "translation funnel" chart is explicitly a conceptual illustration of a well-documented attrition pattern, not a measured cascade — no single source provides clean end-to-end proportions, so none is claimed. Figures span 2018–2026 vintages as labelled.

Sources

  1. Economic Value of AI in Radiology: A Systematic Review, Radiology: Artificial Intelligence 2025 (3,891→49,739 medical-AI articles/yr)
  2. Pesapane et al. — AI in medical imaging: threat or opportunity? Eur Radiol Experimental 2018 (100–150 → 700–800 papers/yr; subspecialty split)
  3. AI in Healthcare 2019–2021 review — maturity analysis (1.8% mature in 2021 vs 7.7% in 2020)
  4. AI in Healthcare: 2024 Year in Review — 28,180 PubMed AI/ML records; imaging a top specialty; LLM pivot
  5. AI research towards open and reproducible science (2026) — 1,206→12,026 papers across five AI conferences, 2014–2024
  6. Park et al. — Quality of science & reporting of radiomics (RQS & TRIPOD), Eur Radiol 2019 (mean RQS 26%; 9.4/36)
  7. Reproducibility of the Radiomics Quality Score, Eur Radiol 2023 (inter-rater ICC 0.30–0.55)
  8. Systematic review of RQS applications — EuSoMII Radiomics Auditing Group (RQS structure & limitations)
  9. METRICS & RQS quality appraisal (chondrosarcoma radiomics) — external testing & open-science gaps
  10. Nagendran et al. — AI versus clinicians: systematic review, BMJ 2020 (10 registered RCTs vs 81 non-randomised; 61/81 claim ≥ expert)
  11. AHRQ PSNet summary of Nagendran et al. (10 trials; 81 non-randomised; high bias risk)
  12. Aggarwal et al. — Diagnostic accuracy of deep learning in medical imaging: systematic review & meta-analysis (reporting-standard references)
  13. Effects of AI implementation on efficiency in medical imaging — systematic review & meta-analysis, npj Digital Medicine 2024
  14. A quantitative analysis of global AI medical studies: gaps in randomised controlled trials, npj Digital Medicine 2026
  15. AI in Healthcare: 2023 Year in Review — imaging as leading specialty for mature AI publications