AI scribes lost to humans on every quality domain — and lost worst in a noisy room
A head-to-head evaluation published Tuesday in Annals of Internal Medicine ran audio from five standardized primary care visits past 11 commercial ambient AI scribe tools and 18 human clinicians, then had 30 blinded raters score every resulting note on the modified PDQI-9. Human notes won across the board — accuracy, thoroughness, usefulness, organization, comprehensiveness. The gap was widest on an acute low back pain case recorded with background noise: clinicians averaged 43.8 out of 50, the AI tools 20.3. Lead author Ashok Reddy's framing is the practical one: these are draft generators, not note authors. (Annals · UW Medicine summary)
Why it matters to you: the worst-case condition in this study — ambient noise, an unfocused complaint — is a fair description of a 3 a.m. ED bay. And a nocturnist's H&P is the document the day team, the consultants, and eventually the utilization reviewer all read. If your group has an ambient scribe in pilot, the failure modes here are thoroughness and organization, which are exactly the domains a tired reader doesn't audit before signing.
Talking point: worth asking at the next group meeting whether anyone is sampling signed AI-drafted notes against the source encounter, and whether the group has a written expectation that the attending edits before signing rather than attests after.