A radiology AI pilot that runs for 30 days and produces no structured measurement data has wasted both time and political capital. The department has disrupted its workflow, the radiologists have an opinion formed from their individual experience rather than aggregate data, and whoever signed off on the pilot will be making an expansion decision based on anecdotes.
The pilot structure matters as much as the technology being piloted. What follows is how we approach the first 30 days when deploying a pre-reading layer at a new site, and what we've found separates useful measurement from noise.
Start with One Modality, Not the Full Queue
The instinct in many pilots is to turn the pre-reading layer on for everything and measure globally. That approach makes the data much harder to interpret. Different modalities have fundamentally different study types, reading times, and draft acceptance patterns. A blended acceptance rate of 68% for the first month tells you almost nothing about whether the tool is working for CT chest, which might be at 85%, versus MRI brain, which might be at 42%.
Starting with one modality, specifically the one with the highest routine volume and the most templated findings language, gives you a clean signal. For most community imaging centers, this means plain film chest X-ray. For higher-acuity sites with significant CT volume, chest CT without contrast often makes a comparable starting point.
The single-modality constraint also makes the IT footprint of the pilot much smaller. You're configuring one DICOM routing rule, one report template mapping, and training one subset of the radiologist panel. Problems that arise are easier to diagnose. When you expand, you're doing so with a tested baseline rather than starting fresh in a more complex environment.
Establish Baseline Before Day One
The measurement problem with most pilots is that they compare "during pilot" data to "after pilot" data and miss the pre-pilot baseline entirely. You need at least four weeks of pre-pilot data for the same modality you're piloting, from the same radiologist group, collected from the same RIS and PACS event logs you'll use during the pilot.
Three metrics to capture in the pre-pilot baseline period:
First, TAT from study complete (when the technologist marks the acquisition done) to report signed, broken out by day of week. The Monday accumulation pattern mentioned in other posts on operational metrics is real, and if you don't capture it in baseline, your pilot week-one data will look worse than it is simply because you're reading into the Monday queue.
Second, mean time in the reading room from study open to dictation start. This is not the same as overall TAT. It's the cognitive friction time: how long does the radiologist spend orienting to the study before beginning to speak or type? This baseline is almost never tracked before a pilot, but it's highly sensitive to pre-reading assistance because a draft eliminates much of that orientation time for routine studies.
Third, count of addenda per 100 reports. Addenda rate is an indirect proxy for report quality and clarity issues. You're not going to see meaningful change in 30 days, but capturing a pre-pilot baseline gives you the metric for six-month retrospective review.
What to Track During the Pilot Period
During the 30-day pilot, you're tracking six data points daily, all for the pilot modality only:
Coverage rate: what percentage of the day's studies in the pilot modality had a draft available when the radiologist opened them. Anything below 92% needs investigation into the routing configuration.
Draft acceptance rate: the percentage of studies where the radiologist accepted the draft as the report basis versus dictating from scratch. For plain film chest X-ray in a routine outpatient population, a 30-day pilot acceptance rate below 55% usually indicates a language mismatch between the draft style and the site's reporting norms, not a model quality issue.
Time-to-sign delta: for studies with a draft available versus studies where the radiologist opened cold (coverage gaps), what's the difference in time from study-open to report-signed? This is the clearest operational value signal.
Radiologist satisfaction score: a simple weekly 1-5 rating from each participating radiologist. We use four questions covering draft usefulness, workflow friction, confidence in reviewing drafts, and overall preference versus prior workflow. Weekly cadence catches opinion shifts that a single end-of-pilot survey misses.
Urgent flag rate: the number of studies in the pilot modality where an urgent finding flag was triggered, whether by the pre-reading layer or by the radiologist on review. Track this not because you expect it to change but because a sustained deviation from baseline in either direction requires review.
Technical error rate: the number of processing errors, timeout events, or delivery failures in the pre-reading pipeline, per 100 studies. This is an infrastructure health metric, not a clinical metric, but it directly affects coverage rate and radiologist trust.
The Week-Two Trust Inflection
A consistent pattern we see across pilots: radiologist acceptance rate and satisfaction scores dip in week two relative to week one, then recover in weeks three and four. Week one tends to involve heightened scrutiny and careful checking of every draft sentence. Week two is when the fatigue of that scrutiny sets in and some radiologists start expressing skepticism. Week three is when radiologists begin calibrating which study types and sections they can trust quickly versus which require careful review.
This pattern is normal and it's important to communicate it to department leadership before the pilot starts. A pilot report written at the end of week two will look significantly worse than one written at the end of week four, and the difference is behavioral adaptation, not model performance. We've seen pilots almost cancelled at day 17 because someone pulled the data before the adaptation had completed.
We recommend scheduling no formal stakeholder review of pilot data until after day 21.
What a Meaningful Baseline Looks Like Before Expanding
At 30 days, you have a meaningful baseline for expansion decisions if the following conditions hold: coverage rate is above 92% consistently for the last two weeks; draft acceptance rate has stabilized (meaning the week-over-week change is less than 5 percentage points); time-to-sign delta is positive and consistent; and radiologist satisfaction is trending up or holding steady, not declining.
We're not saying you need clinical equivalence evidence from a 30-day pilot to justify expansion. That's not what a departmental pilot is designed to produce. What you need is evidence that the workflow integration is functional, that the radiologists have adapted to the new modality, and that the operational metrics are moving in the right direction. Those four conditions together constitute an adequate basis for expanding to a second modality and increasing the radiologist panel size.
Expansion without those conditions met is premature and typically results in a broader deployment that generates skepticism across a larger group of radiologists, most of whom had no voice in the original decision. That's much harder to recover from than a careful 30-day baseline.
What the 30-Day Data Cannot Tell You
One month is not enough time to evaluate addenda rate changes, rare finding detection, or any patient outcome correlation. It's not a clinical trial. The 30-day pilot answers operational questions: does the infrastructure work, does the workflow integrate without disruption, and do radiologists find it useful enough to adopt? Clinical validation is a separate and longer process.
Confusing these two levels of evaluation leads to pilots that are either over-claimed (citing 30-day data as evidence of clinical benefit) or under-valued (dismissing a successful operational pilot because it didn't produce clinical outcome data that a 30-day window couldn't possibly generate).