Video summary
Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
Main summary
Key takeaways
Summary
The video argues that clinical AI note-taking systems—especially ambient “scribes” that transcribe conversations and then generate clinical notes—often produce documentation that looks correct while omitting or altering information in ways that are clinically dangerous. The core issue isn’t obvious hallucinations or transcription glitches. Instead, it’s “quiet failures”: the model outputs text that is technically faithful to what it was given, yet fails to capture what actually matters for patient safety.
Main points and evidence
-
“Dangerous failures look fine.” The speaker shows examples where AI-written notes appear routine on the surface (e.g., headache management) but omit critical context—such as jaw pain on chewing suggesting giant cell arteritis. Missing a single line can change an issue from routine follow-up to an emergency treatment.
-
Severe errors occur at high rates in production. A referenced large real-world study of these kinds of notes found:
- ~1 in 20 notes had errors serious enough to cause significant harm.
- Nearly 1 in 5 involved important omissions (when expanded to all errors).
- >1 in 10 involved hallucinations.
-
Adoption is fast, and tracking is poor. Ambient scribes are already used in about one-third of US practices and growing. The speaker claims adverse event reporting is often missing, meaning errors may not be logged as incidents—creating a “flying blind” situation.
What goes wrong (mechanism)
Ambient scribes have two stages: transcription and then generation.
- Some errors come from misheard words (e.g., sound-alikes).
- However, the speaker emphasizes that most problematic failures remain even with perfect transcription.
The generator can:
- Add information not said,
- Change details that were said,
- Omit important details.
The hardest part for safety systems is determining whether an added/changed/dropped item actually matters in that specific clinical context—because the model lacks “taste/judgment” about clinical importance.
Why existing evaluation approaches fail
- Many teams use an AI “judge” with a rubric, often focused on faithfulness (e.g., whether the note matches the transcript).
- The speaker argues this misses the biggest risk: a judge can be good at spotting literal mismatches, but weak at judging what counts as a clinically serious miss.
- Even when notes pass evaluation, context-dependent omissions can slip through—creating a second silent failure where the verifier effectively “waves it through.”
Proposed fix: continuous, case-calibrated evaluation
The speaker proposes an evaluation loop that updates what “matters” using real expert judgment:
- Discover failure modes from production outputs (build a failure-mode “ontology” from real data, not a guessed list).
- Capture expert corrections and reasoning on representative real outputs.
- Calibrate per output by retrieving the most relevant past expert-reviewed cases and references, then judging the current note against a case-specific standard (not a static rubric).
- Keep the loop running so that if standards or failure patterns shift, the evaluator adapts without requiring retraining everything.
The speaker contrasts this with static rubrics and retrained models, noting that both can “freeze” the wrong notion of quality and go stale as clinical practice and real-world patterns evolve.
Takeaway
Evaluation cannot be a one-time engineering task. The standard of “good and safe” doesn’t exist fully on paper—it must be learned from real outputs and maintained continuously through expert-in-the-loop calibration.
Presenters or contributors
- Sebastian Fox (presenter)
- Composo (company/producer context referenced by the presenter; no additional named contributors in the subtitles)