Video summary
AI and Clinical Reasoning: A New Era of Medical Education and Assessment
Main summary
Key takeaways
Main ideas and lessons
-
Clinical reasoning is fallible and widely affected by diagnostic errors.
- Even with good training and systems, diagnostic errors still occur and significantly impact patients.
- The National Academies of Medicine (NAM) report is referenced as a key decade-old source on the pervasive impact of diagnostic errors.
-
Clinical reasoning (as an educational target) is an iterative, multi-step process.
- It includes:
- gathering information
- deciding tests
- forming a working diagnosis
- discussing with the patient
- creating management plans
- When reasoning goes wrong, it is often multifactorial, but faulty thinking is commonly involved.
- Mark Graber’s work is referenced as a landmark study showing cognitive faults can contribute.
- It includes:
-
What educators knew before generative AI: proven training strategies
- Teach clinical reasoning using approaches that promote:
- Knowledge organization
- problem representations (abstract representations of a case)
- illness scripts (how typical disease presentations unfold)
- management scripts (typical care pathways)
- diagnostic schemas (structured approaches to problems)
- Structural reflection
- metacognition-focused practices that help learners think about their thinking
- diagnostic timeouts
- Team-based thinking
- distributed cognition and diagnostic teams with diverse viewpoints
- leveraging AI to support team processes (not just individual thinking)
- Knowledge organization
- The talk notes that AI will alter how these are implemented and emphasized, but many remain relevant.
- Teach clinical reasoning using approaches that promote:
-
AI’s role in diagnosis: “LLMs outperform humans” is true in constrained settings, but not the end of the story
- In controlled/simulated cases, LLMs often outperform clinicians when information is curated.
- Real clinical environments are more complex; therefore, results don’t automatically mean “replace clinicians.”
- Key point: AI + human does not automatically improve performance unless there is explicit structure.
-
Why “prompting humans to use AI” without structure often fails
- Studies discussed suggest that simply telling clinicians to use an LLM—or adding the LLM into the workflow without a structured reflection strategy—may not improve performance.
- The talk links this to structural reflection literature: reflective prompting alone doesn’t reliably improve outcomes.
-
Promising evidence: structured AI-human collaboration can improve diagnostic performance
- The speaker describes multiple studies (including simulated and real-world clinical settings) showing:
- AI-first vs human-first ordering matters
- “human first / then AI second” can outperform cases where humans adopt AI output prematurely
- AI guidance may be better at improving performance in lower-scoring cases
- Not all cases improve, but meaningful improvements can occur when the interaction is well-designed.
- AI-first vs human-first ordering matters
- The speaker describes multiple studies (including simulated and real-world clinical settings) showing:
-
Real clinical workflow examples
- Urgent care/primary care randomized by clinics (AI + clinician vs standard practice)
- AI flags potential diagnostic discordance.
- Example: a clinician is prompted to reconsider an antibiotic choice when a red-flag discordance appears.
- Independent reviewers found AI + clinician charts scored better across history, testing, diagnosis, and management.
- Over time, clinicians showed reduced unnecessary AI interventions—suggesting learning from feedback.
- Patient-interactive AI conversation + clinician in the loop
- AI gathers patient information via an interactive conversation prior to the visit.
- Clinicians receive transcript-derived output (not necessarily the AI’s full management plan).
- Reviewers found similar diagnostic/management quality; clinicians still integrate context the AI lacks.
- Patients and clinicians reportedly found it helpful (more time for the visit; improved perceived support).
- Urgent care/primary care randomized by clinics (AI + clinician vs standard practice)
-
Curriculum implication: learners will use AI anyway, so training must guide how
- The talk emphasizes a “seat at the table” approach: educators must shape policies, workflows, and assessment.
- Educational tension:
- Upskilling (human-AI synergy improves performance)
- vs deskilling/never skilling/mis-skilling (loss of ability or dependence when tools are introduced incorrectly)
-
Ambient documentation: moving from “should we?” to “how should we do it?”
- AI can reduce documentation burden, affect wellness/burnout, and translate narrative into patient-appropriate literacy.
- Concerns remain about documentation quality (being studied).
- Conclusion: it’s already happening—so education must teach safe, critical use and interpretation.
-
Policies and competency development
- Institutions need:
- guidelines on what tools can be used for what tasks
- privacy and data governance (public tools vs HIPAA/FERPA-compliant institutional tools)
- faculty development so supervision reinforces correct behavior (addressing the “hidden curriculum”)
- Institutions need:
-
Assessment shift: use AI to enhance feedback and measure “diagnostic performance”
- The speaker proposes moving assessment from only “diagnostic reasoning” artifacts to also measuring diagnostic performance (e.g., whether a diagnostic delay occurred).
- Example:
- Disease-based approach starting with VTE (venous thromboembolism) due to high error rates.
- Validation uses Safer Dx (human adjudication tool/work by Hardeep Singh) to classify diagnostic delays.
- A model uses prompts to determine whether residents considered VTE and uses imaging report gold standards to adjudicate outcome/discordance.
- A dashboard provides iterative feedback with standard-setting and trend views.
-
Ambient assessment and feedback (future direction)
- Work discussed uses ambient recordings (not only documentation) to give feedback on communication skills (and potentially clinical skills more broadly).
- Rationale: documentation artifacts may no longer be the only high-quality retrospective signal.
Methodology / instructional framework presented (detailed bullets)
A) Human-AI integration principles (education strategy)
- Teach structured collaboration, not “AI as an unstructured add-on.”
- Encourage learners to commit to their own thinking first (to reduce bias from deferring to AI output).
- Then integrate AI as a second opinion (structured workflow).
- Use structured reasoning tools derived from clinical reasoning literature
- Apply a framework conceptually based on structural reflection:
- generate a differential with supporting vs opposing reasoning
- estimate probabilities / next steps
- Apply a framework conceptually based on structural reflection:
- Require critical appraisal
- After AI output, teach learners to evaluate evidence quality and safety before acting.
B) “AI-first opinion vs second opinion” teaching model (implementation concept)
- Step 1: Human-first differential/plan
- Learners produce their own differential diagnosis and management considerations.
- Step 2: Generate AI-assisted second opinion
- Learners obtain AI output after committing to their own reasoning.
- Prompts scaffold the learner through structured reflection.
- Step 3: Critically appraise and reconcile
- Learners compare AI output to their own reasoning:
- What changed?
- What was newly considered?
- What stayed consistent?
- Learners compare AI output to their own reasoning:
- Step 4: Reflection and commitment
- Learners explicitly decide how they will use AI in future cases.
- They assess whether the AI helped or distracted from correct reasoning.
C) Faculty/supervisor micro-skills framework for ward use (“Depth AI”)
- The talk references a framework (Depth AI) to guide faculty conversations.
- Faculty prompts include:
- What tools did you use?
- How did you use them? (what prompts did you enter)
- What evidence did you verify? (e.g., learner answers “I didn’t” → teach evidence checking)
- Safety / ethics / privacy checks
- How did you assess accuracy/safety?
- Feedback and improvement loop
- How might you change your AI use next time?
- Tailor teaching to gaps
- Where the learner is weak determines which general principles to emphasize.
D) Institutional policy approach (privacy/data governance)
- Do not treat all tools equally
- Public tools (e.g., “OpenEvidence”) vs evidence-vetted tools (e.g., “UpToDate Expert AI”) vs internal HIPAA/FERPA tools (e.g., “UltraViolet AI”) have different constraints and reliability profiles.
- Data handling rules
- Public tools: avoid patient-identifiable or detailed EHR data.
- Institutional compliant tools: may allow appropriate integration of EHR data (where set up by the institution).
- Assessment controls
- For certain assessments, AI usage may be prohibited or restricted.
- Guardrails must be explicit; otherwise “using AI when not disallowed” becomes a policy gap.
E) Assessment methodology using AI (diagnostic performance tooling)
- Ground level: define the human behavior / skill being assessed
- Start by specifying clinical reasoning/performance behaviors (e.g., differential quality, reasoning explanation).
- Use AI to score artifacts
- Example: admission note reasoning elements (e.g., did they include a differential; did they explain reasoning).
- Disease-based performance validation workflow (example VTE)
- Create prompts to detect whether residents considered the condition (e.g., VTE).
- Use imaging report narratives and timing rules as gold standards.
- Use Safer Dx (human-reviewed diagnostic delay adjudication) as the validation benchmark.
- Provide dashboard feedback on:
- cases with/without delays
- trend over time
- contextual factors and balancing measures (e.g., avoid “everyone ordered imaging” behavior dominating interpretations)
- Resident advisory panels and quality improvement
- Residents participate in choosing feedback formats and goals.
- Implement dashboards with data visibility and reflective exercises.
Speakers / sources featured
People / speakers
- Bob Wachter (introducer; Chair of the Department of Medicine at Moffitt)
- Berdychev (primary visiting professor speaker; NYU Grossman School of Medicine; NYU titles described in intro)
- Michelle Guy (Master Clinician inductee)
- Raman Chawla (Master Clinician inductee)
- Vivek Jain (Master Clinician inductee)
- Scott Steiger (Master Clinician inductee)
- Sunny Wang (Master Clinician inductee)
- Gurpreet Dhaliwal (namesake of the master clinician lectureship; mentioned as present)
- Eric Jameson (audience member; UCSF instructor/faculty; asks questions)
- Nina (audience member; internal medicine resident at the program; asks questions)
- Verity (audience member; asks question about hidden curriculum and training residents to supervise students)
- Dr. Okter (appears to refer to Bob Wachter; introduction contains a brief mis-transcription)
- Daniel Sartori (implementation lead mentioned)
- Lauren Hiry (resident advisory panel member mentioned)
- Bijal Rajbhat (resident advisory panel member mentioned)
- Dr. Briscoe (referenced as connected to the Depth AI framework; may join virtually)
- Dr. Abdo (referenced as connected to the Depth AI framework)
- Mark Triola (NYU colleague referenced as creating the prompt-a-thon/platform; appears in curriculum/tools discussion)
- Marc Triola (same person as above, referenced with full name)
- Rodman (referenced as a leader in the field of AI education competencies; name appears in transcript)
- Hardeep Singh (Safer Dx referenced; leader in diagnostic error field)
- Yuliya Umcheva (referenced in ambient assessment work)
- Amy Jessi Burke-Ravelle (referenced in ambient assessment work)
Organizations / reports / tools referenced
- National Academies of Medicine (NAM) report on diagnostic errors
- Safer Dx (Hardeep Singh / diagnostic error adjudication tool)
- UpToDate and UpToDate Expert AI
- OpenEvidence
- UltraViolet AI (described as institutional HIPAA/FERPA-compliant)
- Depth AI framework (for faculty conversations)
- One-minute preceptor / micro-skills (referenced as an analogy for the feedback framework)
- OSCEs (assessment context referenced)
- EHR (Electronic Health Record)
- HIPAA / FERPA (privacy compliance frameworks referenced)