Blog Diana Chen 7 min read

What AI Can Draft in a Radiology Report Today (and What It Cannot)

Abstract concept of a partially completed document with an AI assistance indicator

Vendors who sell AI report drafting tools have an incentive to describe the capability as broadly as possible. Departments evaluating those tools have an incentive to believe the broadest description, because a tool that drafts everything would meaningfully change their workflow economics. The reality, which any honest account needs to start with, is that current AI drafting capability varies significantly by study type, and the line between what works well and what does not work well yet is worth understanding in specific terms before a department makes a workflow design decision based on it.

Where AI Drafting Works Well Today

The clearest category of reliable AI drafting is structured findings language for normal or near-normal studies in well-supported modalities. A normal chest X-ray, a normal brain MRI, a normal knee MRI with no acute finding, a negative CT of the abdomen that is unremarkable across all reviewed structures. For these studies, the findings section of the report follows a predictable structure: each anatomical region is addressed, the absence of acute finding is documented, and standard language expresses the normal appearance.

This is where current AI drafting systems produce output that a radiologist can verify and accept with light editing. The draft describes the lungs as clear without focal infiltrate. The mediastinum is of normal width. Osseous structures appear intact. The impression states no acute cardiopulmonary finding. A radiologist reading that draft against their own review of a clearly normal study can confirm it, edit any wording that does not match their findings, and sign. The total time is materially shorter than dictating from scratch.

Plain film reading is the highest-volume category where this works. Chest X-rays represent a very large fraction of daily study volume in most departments, and the majority of them are normal or near-normal. A drafting system that handles this category well, with good coverage of the anatomy and standard reporting language that matches the department's style, produces time savings that compound significantly across a shift.

Routine outpatient CTs with limited or no findings are a second strong category. An abdominal CT ordered for a vague complaint that comes back unremarkable, with no significant pathology in the organs reviewed and no significant lymphadenopathy, is well within the current capability of AI drafting for the findings section. The impression for a clear negative is straightforward and the AI generates it reliably.

Where AI Drafting Has Clear Limitations

Complex multi-system findings with interdependent interpretation are not where AI drafting adds value today. An abdominal CT that reveals a liver lesion with characteristics suggesting one differential, an incidental finding in the adrenal gland that needs contextualization, and a renal cyst that requires comparison to a prior study to characterize: the impression for that study requires synthesizing multiple findings, weighing their clinical significance in relation to each other, and producing a prioritized summary that reflects the radiologist's judgment about what the ordering clinician most needs to know. Current AI drafting systems do not produce reliable impressions for this type of study.

Quantitative measurements and structured comparisons to prior studies also remain outside reliable AI drafting. Lung nodule characterization per Fleischner Society guidelines requires measuring the nodule, comparing to prior measurements, and applying the guideline recommendation. This is a structured task that in principle could be automated, and some specialized tools attempt it, but it requires reliable detection, accurate measurement, and correct guideline application, and the error rate in current systems for this specific task is high enough that close verification by the radiologist is always required. The draft, if generated, cannot be lightly reviewed. It must be independently verified, which eliminates the time-saving value of drafting for this finding type.

Studies with significant artifacts are another reliable failure mode. A CT with motion artifact that obscures portions of the anatomy, or an MRI with susceptibility artifact from prior hardware, requires the radiologist to characterize the artifact, note its impact on image quality, and qualify findings accordingly. AI drafting systems trained on clean studies do not generalize well to artifact-heavy studies. A draft that does not acknowledge significant artifact is worse than no draft, because it requires the radiologist to identify the limitation and add language that was not there, which is slower than starting from scratch for this portion of the report.

The Impression Section Is the Hardest Part

Within any given study, the findings section and the impression section have different AI difficulty profiles. The findings section, which documents what is seen in each anatomical region, has a more formulaic structure and is more tractable for AI generation. The impression section, which synthesizes the findings into a clinical conclusion and recommendation, requires judgment that is harder to systematize.

For simple normal studies, the impression is straightforward and AI handles it well. "No acute cardiopulmonary finding" following a normal chest X-ray is not a judgment-intensive conclusion. For any study with findings, the impression requires the radiologist to decide which findings are primary, what the differential diagnosis is, what the recommendation is, and how to express clinical urgency. AI can produce text that looks like an impression, but the clinical judgment embedded in a well-written impression is not yet reliably automated.

A department that designs a workflow where AI drafts the findings and the radiologist generates the impression from scratch may find this a useful division of labor for some study types. The findings section for a normal study requires a radiologist to verify seven or eight normal observations that the AI has already documented. The impression requires a single accurate sentence. The time savings are real even if the impression is dictated rather than accepted from draft.

Voice Recognition vs. AI Drafting: Different Things

It is worth separating AI drafting from voice recognition, because they are sometimes conflated. Voice recognition, as used in radiology dictation systems like PowerScribe or MModal, converts spoken words to text. The radiologist still generates the content by speaking. Voice recognition introduces accuracy errors at the word level, particularly for medical terminology, but the clinical content is entirely the radiologist's.

AI drafting generates content directly from the imaging study. The radiologist receives a pre-populated report and verifies it, rather than speaking the content from scratch. These are categorically different contributions to the report. Voice recognition changes the input modality but not who generates the content. AI drafting changes who generates the content but not the requirement for radiologist verification before signing.

The combination of AI drafting for appropriate study types with voice recognition for the portions the radiologist needs to modify or add is the practical workflow. The AI draft comes in. The radiologist reads the study, reads the draft, makes corrections using voice recognition, and signs. The efficiency gain comes from the proportion of the draft that requires minimal modification versus the proportion requiring significant edit or addition.

Tracking the Line in Your Own Department

The boundary between "AI drafts this well" and "AI does not draft this reliably" shifts as models improve and as departments configure their drafting systems to better reflect their specific reporting style and patient population. The right approach is not to assume the vendor's characterization of capability is accurate for your specific context, but to measure it directly during a pilot period.

Draft acceptance rate by study type and modality is the metric that surfaces the capability boundary. If acceptance rate for chest X-rays is high and acceptance rate for abdominal CTs with findings is low, you have direct evidence of where the drafting adds value and where it does not, calibrated to your department's volume and radiologists' standards. Those numbers are more informative than any general characterization of AI drafting capability.

Being honest with radiologists about what the AI drafts well and what it does not is also important for adoption. A radiologist who expects a reliable draft for every study and receives an unreliable one for a complex case will distrust the system. A radiologist who understands that the AI drafts routine studies well and flags complex ones for full dictation can calibrate their workflow accordingly. The technology is more useful when the expectations are accurate.

Ready to see how Radivault fits your department?

Talk to the team More articles