People have been predicting the "imminent" arrival of clinical AI in radiology for a long time. There were pilots a decade ago, proof-of-concept papers going back further, and vendor announcements that promised more than they delivered. So when we talk about pre-reading at scale becoming viable now, it is worth being specific about what changed and why the timing is different from the last wave of optimism.
Two things crossed a threshold at roughly the same time. The model quality got good enough for specific, well-defined tasks. And the backlog problem got bad enough that departments stopped asking whether to automate and started asking how.
What "Good Enough" Actually Means for a Pre-Read
Good enough does not mean a model that can replace a radiologist's judgment across all studies. That bar does not need to be cleared for a pre-read layer to be useful. What needs to be true is narrower: the model has to reliably surface the subset of studies that contain time-sensitive findings, and it has to draft structured language for the studies it marks as routine that a radiologist can verify in significantly less time than writing from scratch.
On the first part, current chest CT and brain MRI triage models have matured to the point where the sensitivity on acute findings like pneumothorax, large vessel occlusion, and midline shift is consistent enough to use in a real worklist-ordering context. Not every modality is there. Not every finding category is there. But the high-acuity, high-consequence finding categories that most departments care most about prioritizing, those are tractable now in a way they were not in 2019 or 2020 when the underlying architectures were earlier and the training datasets were smaller.
On the draft side, the shift from rigid template-filling to generative structured text has closed the gap between what AI outputs and what a radiologist would have dictated for a routine normal study. The jump between generations here was real. Earlier systems produced text that read like it was assembled from lookup tables. The current generation produces something that reads like a dictation, which is the threshold you need before an edit-to-sign workflow is faster than dictate-from-scratch.
The Volume Side of the Equation
Model quality explains why it is possible now. Study volume explains why departments are actually willing to change workflow now.
The numbers are not obscure. Medical imaging order volume has grown consistently faster than the radiologist workforce for most of the past decade. The pandemic compressed what might have been a slow-motion supply-demand gap into something departments could not rationalize away. Departments that had previously managed the backlog through overtime and extended reading hours hit a ceiling. The turnaround time problem became visible to hospital administrators, not just radiology chiefs.
That visibility changed the decision calculus. When the cost of not changing workflow was abstract, the friction of integration felt large relative to the benefit. When the cost of not changing became a documented performance problem, the risk tolerance for workflow change increased. Departments that would not have run a pilot two years ago are running pilots now because the alternative is explaining why urgent studies sit unread for hours.
Why Earlier Attempts Did Not Stick
It is worth being direct about this, because the graveyard of radiology AI pilots that did not scale is instructive. Most of the early tools failed to stick for one of three reasons.
First, they were detection-only. A system that flags a finding but does not integrate into the worklist ordering creates extra work for whoever has to look at the flag. If a radiologist has to leave their PACS environment to check an alert, the interruption cost often exceeds the benefit of the alert. Integration with worklist prioritization, not just detection output, is a prerequisite for workflow impact.
Second, they targeted the wrong studies. High-complexity, high-difficulty studies are not the bottleneck. The bottleneck is the volume of routine studies that are straightforward but require attention. A tool that performs impressively on rare edge cases but does not help with the 70% of studies that are normal or near-normal does not move the number that matters.
Third, they had no path to report drafting. Pre-reading that ends at a triage score is half a workflow. The time savings for a radiologist come from both getting to the urgent study faster and spending less time on the routine one. If the output of the AI is only a priority ranking and not a structured draft, you have addressed the queue order but not the per-study time cost.
We are not saying the earlier tools were poorly made. Some were technically strong. The architecture of how they connected to clinical workflow was the problem.
What the Threshold Crossing Looks Like in Practice
A concrete example from the kind of department we work with: a mid-size regional imaging center in the mid-Atlantic running roughly 200 to 250 CT and MRI studies per day across two reading shifts. Before a pre-read layer, the worklist was ordered by study arrival time, with occasional manual reprioritization when a technologist flagged something to the on-call radiologist. Urgent studies regularly sat 45 to 90 minutes before being read, not because no one cared but because the radiologist working down the queue did not always know the urgent study had arrived until they got to it.
With a pre-read layer that reorders the worklist based on AI triage output, the urgent studies move to the top automatically. The radiologist does not change what they do. They read what is in front of them. The queue order changes, not their reading behavior. That is the simplest version of the value, and it does not require the draft acceptance workflow at all to produce a meaningful clinical outcome.
The draft layer adds a second layer of time savings on top. But the triage-to-worklist integration is the piece that has crossed from useful-in-theory to deployable-in-practice because the models are now reliable enough and the PACS integration path is standardized enough via HL7 order updates and DICOM worklist modification messages that the IT lift is no longer the barrier it was.
The Limits to Acknowledge
Pre-reading is not uniform across modalities. The maturity curve for chest CT is ahead of the maturity curve for musculoskeletal MRI. Plain film triage is tractable. Multi-parametric MRI interpretation is not at the same stage. Departments that want to deploy broadly should have realistic expectations about which modality verticals are ready and which ones need another year or two of model development.
The draft quality is also not uniform across study types. A high-quality AI draft for a normal chest X-ray is not the same as a high-quality AI draft for an abdominal CT with incidental findings. The complexity of the impression section scales with the complexity of the study, and that scaling does not always favor the AI. The value is concentrated in the routine, high-volume study types, which is where most of the workflow time is spent anyway.
The timing argument is not that the technology is finished. It is that the combination of what is ready now with the severity of the problem now makes this the first deployment window where the benefit clearly outweighs the integration cost for a meaningful subset of department workflows. That is a narrower claim than "AI is here," and it is the claim we are actually making.