From Training to Doing: What NVIDIA’s GPU Mix Says About the Radiology AI Hype
Radiology definitely does not have a hype shortage. Another month brings another model, another benchmark, another FDA clearance, another foundation model or VLM announcement, another prediction that radiologists are on the way out. Most of these are poor instruments for figuring out where we stand in the adoption curve.
A benchmark score tells you what a model can do in a lab. It does not tell you what the industry is doing with it.
One useful signal comes from somewhere we radiologists may not look: what NVIDIA’s GPUs are actually being used for.
In the 2025 book The NVIDIA Way, Tae Kim traces a distinction that is easy to miss inside the noise of AI coverage.¹ GPUs purchased to train a model are a bet on future capability. GPUs running a model in production, doing inference, are a measure of how much of that capability is already being used.
Training tells you how much effort is going into building AI. Inference tells you how much AI is actually doing work in the world.
Tracked over a few years, that distinction says something specific and checkable about where radiology sits in the cycle.
Something Has Changed
In early 2024, NVIDIA’s own numbers put a rough shape on this. On the company’s fiscal fourth quarter 2024 earnings call, CFO Colette Kress said inference had driven roughly 40 percent of data center revenue over the prior year.² The rest, in broad strokes, was training. A year and a half later, the language shifted from a percentage to a declaration. NVIDIA’s fiscal 2025 annual report describes fiscal 2025 as a break from the past: inference workloads surpassed training, and inference became, in the company’s own words, the dominant workload.³
Independent estimates put a number on how far that shift has gone.
Deloitte’s 2026 technology predictions place inference at roughly two-thirds of all AI compute this year, up from about one-third in 2023 and roughly half in 2025.⁴
Different sources measure this differently, and none of these figures should be read as a precise accounting of GPU-hours. But the direction is not ambiguous.
NVIDIA’s own product language has moved with the numbers. Its fiscal 2026 results lead with the Rubin platform’s inference token cost, not training throughput, and describe compute demand compounding across training and inference simultaneously, with inference now the more urgent constraint.⁵
Why the Crossover Matters
Training answers a narrow question: can we build a model capable of doing X. Inference is what happens every time that model is actually asked to do X again, for a real case, at scale.
A shift toward inference is a shift from building capability toward consuming it. Or, more bluntly, from experiment toward work.
Inference is not synonymous with productive work. A chatbot query is inference. A failed pilot still running in the background is inference. An experimental deployment nobody has validated is inference. But at the scale NVIDIA and Deloitte are describing, across an entire industry, the aggregate direction of travel is still informative, even if any single inference call is not.
Radiology Was Always Going to Be an Early Test Case
If AI really is moving from training toward inference industry-wide, radiology is one of the places that transition should show up first. The inputs are already digital. The workflows are repetitive and measurable. The datasets are large and label-adjacent.
The labor is expensive and chronically short.
The infrastructure, PACS, RIS, EHR, already exists to plug into. And outcomes, unlike a lot of AI use cases, can often be audited retrospectively against ground truth. As of the FDA’s most recent update to its AI-enabled medical device list, radiology still accounts for roughly three-quarters of all AI-enabled device authorizations in the United States, a share that has held steady for years.⁶
Continuing with my common theme, none of that means AI can replace a radiologist. It means radiology is unusually well suited to bounded automation, which is a narrower and more useful claim.
The mammography literature is where that claim gets tested against real patients.
The Mammography Precedent
A trial published in Nature Medicine in early 2026 gives the clearest picture yet of what AI doing real work looks like in practice, rather than in a benchmark table. Researchers at the Córdoba Breast Cancer Screening Unit in Spain ran a prospective, paired trial across 31,301 women undergoing routine mammography between 2022 and 2024.⁷ Every exam was read two ways: the standard double read, and an AI-supported strategy in which exams the AI classified as low risk were called normal without any radiologist reading them at all.
The results are not a clean win, and that is exactly what makes them worth reading closely. Radiologist workload dropped 63.6 percent. Cancer detection rose about 15 percent. But recall also rose, roughly 15 percent relative to baseline, and the trial’s pre-specified non-inferiority margin for recall was not met.
What makes this trial different from most of the AI-and-mammography literature is not the detection number. It is that a defined slice of the clinical workflow ran with no radiologist looking at the image at all, prospectively, inside a real screening program, with a measurable and disclosed tradeoff. That is a different kind of milestone than a lesion marker, a risk score, or a retrospective AUC. Real work moved off a human’s desk, imperfectly, and the imperfection was measured rather than hidden.
Then the Regulatory Line Moved
That same logic just stopped being a research question and became a market fact. On September 3, Vara, a Berlin-based AI company, announced what it describes as the first CE Certification under the EU’s Medical Device Regulation for an autonomous breast triage product, one that classifies certain mammograms as normal with no radiologist reading them at all.8 Every other exam still goes to at least one radiologist. Vara says the certification targets organized population screening programs in Europe, where a shrinking radiologist workforce is straining programs built around reading every mammogram twice.
What makes this worth a radiologist’s attention is not the autonomy claim by itself. Plenty of vendors have implied their tool is trustworthy enough to skip human review. What got Vara a regulator’s signature is the infrastructure sitting underneath the model. The certification covers a system called Atmon, short for Autonomous Triage Monitoring, which sets and continuously monitors each site’s operating point, tracks drift in mammography hardware, system health, and daily performance signals, and can revert a site back to full radiologist reading whenever those signals move outside defined limits.8 Vara’s own framing is that this gives it the continuous human oversight the EU AI Act will soon require of high risk systems, built in rather than bolted on.
Vara did not get this certification on the strength of a single validation study. Its CTO described the achievement as resting on seven years of continuous monitoring of real world performance, case by case, which is what let the company build something it could stand behind without a second reader.⁸ The company also plans to sell Atmon on its own, to customers who never adopt autonomous triage, because the monitoring layer has value independent of the autonomy it enables.
That is the mammography trial’s tradeoff, operationalized. The Córdoba data showed that a bounded slice of work can move to the machine if you are honest about the tradeoff and measure it. Vara’s approval shows a regulator agreeing that this can happen in production, on the condition that the tradeoff is being watched in real time, continuously, with a defined trigger to hand work back to a human the moment performance drifts. The monitoring is not a feature added to the autonomy claim. It is the thing that made the autonomy claim regulator-credible in the first place.
What This Doesn’t Prove
None of this means autonomous radiology has arrived. Vara’s certification is EU-specific, is not available in the United States, and covers one narrow decision, whether a screening mammogram is normal, not diagnosis of any kind. The Córdoba trial ran in one country’s screening program and did not clear its own recall bar. Neither result tells us that today’s models generalize across institutions, that hallucination and reliability problems in more complex imaging have been solved, or that liability and reimbursement questions have caught up to what the technology can already do.
It is worth noticing where both events landed. I have written before that the earliest reading gains would show up in high-volume, relatively homogeneous screening tasks: mammography, lung nodule follow-up, normal chest X-ray triage, bone age, the tasks with well-defined outputs, large training datasets, and measurable ground truth.⁹
Mammography is that kind of task. Neither is evidence that AI has moved past that zone into the more heterogeneous, less standardized reading that fills most of a radiologist’s day. They are evidence that the zone itself just earned a regulator’s signature.
What has changed is narrower, and I think more useful.
The central question is shifting from ‘Can AI do radiology?’ to ‘Which specific pieces of radiology can AI do reliably enough that we should stop using a human to do them the old way?’
The Hype Cycle Is Pointing at the Wrong Layer
For a few years, most of the attention in this field has gone to the model. Whose foundation model scores highest. Whose VLM has the most FDA clearances. Whose benchmark number is biggest this quarter. If inference really is becoming the dominant workload, the bottleneck moves downstream from the model to everything around it: integration, reliability, monitoring, workflow redesign, unit economics, and the regulatory boundary of where in the loop does a human sit, if at all.
Vara’s path to certification is a clean illustration of that shift. The company did not win on model score. It won by building the observability layer that let a regulator trust the model’s output in the absence of a second reader. The winning system in this next phase may not be the one with the highest AUC. It may be the one that can convert 100 units of radiologist work into 70, then 50, then 30, while a monitoring layer proves, continuously and in an auditable fashion, that outcomes remain stable, or improve.
What I’m Watching Next Quarter
I don’t think NVIDIA’s training-to-inference mix is a crystal ball for radiology. I do think it is a useful macro indicator of where the broader AI cycle actually stands, and the events immediately underneath it, a prospective trial willing to publish its own tradeoffs, a regulator willing to certify an autonomy claim tied to continuous monitoring, are the parts worth tracking closely.
Every quarter I am watching NVIDIA’s own commentary on training versus inference, independent estimates of the same split, and, most importantly for healthcare, prospective examples where AI actually changes who or what does the work, with the tradeoff disclosed rather than buried.
In 2024, NVIDIA said roughly 40 percent of its data center revenue was tied to inference. A year later, it said inference had surpassed training outright. Broader estimates now put inference at roughly two-thirds of all AI compute. None of that tells us AI is ready to read every study tomorrow.
It does suggest we are leaving the era where most of the industry’s energy went into teaching models what to do, and entering the one where the harder question is where we can responsibly let them do it, and how we will know if we got that wrong.
Radiology, again, looks like one of the first places we are going to find out.
References
1. Tae Kim, The NVIDIA Way (2025).
2. NVIDIA Q4 fiscal 2024 earnings call, February 21, 2024 (Colette Kress, CFO commentary). https://investor.nvidia.com/news/press-release-details/2024/NVIDIA-Announces-Financial-Results-for-Fourth-Quarter-and-Fiscal-2024/
3. NVIDIA Corporation, Form ARS, Fiscal 2025 Annual Report. https://www.sec.gov/Archives/edgar/data/1045810/000104581025000098/finalforfiling-2025xannual.pdf
4. Deloitte , “Why AI’s Next Phase Will Likely Demand More Computational Power, Not Less,” 2026 TMT Predictions. https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/compute-power-ai.html
5. NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026
6. “Radiology Gets 68 New FDA-Cleared Algorithms,” Radiology Business, June 2026. https://radiologybusiness.com/topics/artificial-intelligence/radiology-gets-68-new-fda-cleared-algorithms
7. Elías-Cabot E, Romero-Martín S, Raya-Povedano JL, Rodríguez-Ruiz A, Álvarez-Benito M. “AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening: a paired, noninferiority trial.” Nat Med. 2026;32(4):1296-1305. https://www.nature.com/articles/s41591-026-04277-x
8. “AI Company Scores World’s 1st Approval for Breast Triage Tool That Skips Radiologist Review,” Radiology Business , September 3, 2026. https://radiologybusiness.com/topics/artificial-intelligence/ai-company-scores-worlds-1st-approval-breast-triage-tool-skips-radiologist-review
9. Ty Vachon, M.D., “What Would Happen If AI Doubled Radiologist Reading Capacity,” LinkedIn. https://www.linkedin.com/pulse/what-would-happen-ai-doubled-radiologist-reading-ty-vachon-m-d–no9hc/
Image credit: The Day the Earth Smiled https://science.nasa.gov/photojournal/the-day-the-earth-smiled/ Easter egg for readers of my first AI articles

