In one sentence
At two US medical schools, 103 students reported how they already use large language models — heavier users felt more confident, 86.4% wanted formal training, and LLMs ranked least helpful for flashcards.
What the researchers did
Herpin and colleagues ran a 30-item cross-sectional email survey at the University of Pennsylvania’s Perelman School of Medicine and the University of Michigan Medical School (IRB-approved; distributed April 2025 with a two-week reminder). Items drew on emerging AI competency frameworks and covered use patterns, self-rated knowledge, privacy confidence, perceived educational value, and hopes for curriculum.
Responses were summarised with descriptive statistics; students were clustered into low, moderate, and high LLM-usage groups (k-means). Group differences used non-parametric tests; ordinal logistic regression examined educator-driven LLM engagement. About 7.7% of Michigan and 5.2% of Penn full-time students responded (n = 103 overall; demographics completed by 83).
What they found
- High vs low users reported greater LLM knowledge (3.46 ± 0.82 vs 2.40 ± 0.91, p = 0.0004) and more confidence judging HIPAA-related compliance (p = 0.016).
- High users agreed more strongly that LLMs improve medical education (4.14 ± 0.79 vs 2.16 ± 1.19, p < 0.0001) and foresaw more clinical applications (p = 0.0001).
- Usage (p = 0.007), not institution (p = 0.714), predicted educator-prompted LLM engagement.
- Students rated LLMs most helpful for fact-finding, literature summarisation, and differential diagnosis — and least helpful for flashcards.
- High users noticed inaccuracies more often (p < 0.0001).
- 86.4% endorsed formal training, especially critical thinking and ethical/legal skills.
What this means for learners and educators
- Silence is not a policy: students are already using tools; structured literacy (limits, privacy, critique) matches what they say they want.
- Treat LLMs as research and reasoning aides, not as a replacement for spaced flashcard systems — respondents themselves ranked cards among the weakest use cases.
- Heavier users may be better at spotting errors, not only more enthusiastic — curricula can lean into verification habits rather than blanket bans.
- Faculty modelling matters: engagement tracked usage, not which of the two schools students attended.
Limitations and what we don't know yet
Response rates were low, so enthusiasts (or critics) may be over-represented. All outcomes are self-reported, not exam scores or clinical performance. Two research-intensive US schools may not match other countries or programme types. Clustering usage into three groups simplifies a messy continuum. The survey cannot prove that training improves safe use — only that students ask for it.