← All resources
Blog9 min read

Medical Affairs Can Measure Everything but the Conversation. Virtual Humans Can Help

Medical Affairs professionals discussing the shift from training completion to readiness.

The Conversation Nobody Sees

Fifteen minutes that go sideways

Picture an MSL we'll call Priya, six weeks out of onboarding, sitting with a cardiologist who has actually read the Phase III paper. He wants to know why the primary endpoint was a composite. He asks whether the benefit holds in patients with reduced kidney function, then compares the whole thing to the drug he already prescribes. Halfway through her answer he mentions a use that isn't in the label, casually, the way a busy clinician thinks out loud.

Priya knows this material cold. She scored well on the certification test and told her manager she felt ready, and none of that helps her much in this moment.

What the organization writes down instead

Her company will record that she completed the curriculum, that she visited the account, and maybe that she filed an insight afterward. The fifteen minutes that mattered are gone. Nobody watched them, nobody scored them, and the only person who can describe what happened was one of the people inside the conversation.

That creates a basic measurement problem for Medical Affairs. Organizations can track whether training happened and whether an engagement occurred much more easily than they can observe how well the scientific exchange actually went.

Medical Affairs Measures What's Easy to Count

A serious investment with a narrow window

Medical Affairs puts real time into readiness. A 2024 benchmarking report based on input from 22 Medical Affairs executives representing 20 companies reported an average of 188 hours of onboarding training and nearly 100 hours of continuous training per year for MSLs [1]. Then look at what commonly counts as evidence that the training worked: a knowledge test, a completion status, a self-rating, and perhaps a manager's impression from a role play.

Those measures can be useful, but they are not the same thing as observing performance in a difficult scientific exchange. A knowledge test can tell you whether someone knows the evidence. It cannot tell you how that person will use the evidence while a skeptical expert interrupts, challenges the study design, raises a competitor, or asks a question that crosses into an unapproved use.

A medical science liaison in conversation with a clinician, illustrating the conversation training metrics miss.

The people in these roles say the same thing

This measurement problem shows up in the research. In a nationwide survey of 179 MSLs in Spain, 56% did not agree that their existing performance metrics reflected their true value. The same study illustrates how broad the MSL role has become: respondents identified off-label information management, KOL relationships, continuing medical education for health care professionals, and involvement in clinical trials among important responsibilities [2].

A much larger global survey of 1,023 Medical Affairs professionals across 63 countries found a similar disconnect. Ninety-two percent reported using activity-based measures such as the number of KOL engagements. Yet only 7% preferred purely quantitative metrics, while 52% preferred qualitative metrics. When respondents were asked what MSL KPIs should measure, 70% selected the quality of KOL/HCP relationships or engagements and 67% selected the quality of actionable insights gathered [3]. Those are closer to properties of an interaction than simple counts of how many interactions occurred.

Measuring quality is harder

The Medical Affairs Professional Society has described the same challenge. Its Field Medical metrics guidance states that measurement needs to capture not only how often activities occur but also the qualitative impact of those activities [4]. MAPS' current Medical Affairs Competency Framework is similarly broad, with 42 competencies across seven domains, spanning scientific and technical knowledge, strategy, evidence generation, customer engagement and scientific communication, leadership, business knowledge, and medical governance and compliance [5].

A clinician reviewing Medical Affairs measurement, with the finding that 56% say metrics do not reflect their value.

The difficult part is not deciding that these capabilities matter. It is finding repeatable ways to observe them when someone has to use several of them at the same time.

The Gap Shows Up Well Beyond Field Medical

Launches move faster than the training can

A launch makes the problem easier to see. Data changes, a competitor publishes, a safety question emerges, a new subgroup analysis becomes important, and the scientific narrative gets updated again. Teams read the updated material and then face the same unknown as before, which is whether anyone can discuss it under pressure.

Rehearsing against a scenario library that gets updated as the evidence changes is a different kind of preparation than another slide review. It also gives a launch leader something to look at before the first meetings happen.

Medical Information and compliance-sensitive moments

The unsolicited off-label question is the cleanest example of a conversation nobody gets to practice. FDA's guidance on communications from firms to health care providers about scientific information on unapproved uses draws careful lines [6], and the people who have to hold those lines usually learn them from a policy document rather than from repetition.

The same is true for inquiry handling in Medical Information, where the skill is clarifying what's actually being asked, knowing where the boundary sits, and escalating cleanly. Those behaviors can be taught and tested. Right now they are mostly neither.

Advisory boards, trial sites, and global teams

Advisory board preparation involves predictable stakeholder profiles and predictable hard questions, which makes it a reasonable candidate for rehearsal. So does site-facing communication, where consent discussions, adverse event conversations, and protocol deviations all depend on how somebody talks under stress.

Global teams add another layer, and here it's worth being careful. A platform that supports forty languages has not thereby validated a scenario for local label, local prescribing context, terminology, or scoring equivalence. A translated English rubric is not automatically a working French or Japanese assessment.

Where AI Virtual Humans Come In

Practice that doesn't wait for a trainer

Live role play helps, but it happens once or twice, it takes trainer time, and the difficulty swings depending on who plays the expert. AI virtual humans change that arithmetic. Instead of picking answers in a branching scenario, someone speaks out loud with a simulated cardiologist, investigator, or advisory board member who asks follow-ups, disagrees, brings up a competitor, and wanders into territory that isn't on label.

The exchange is then scored against a rubric the organization wrote itself, covering scientific accuracy, how the limits of the evidence were framed, whether the person answered the question that was actually asked, how a boundary was handled, and whether a next step was agreed on. Then they do it again, and that part matters more than the avatar on the screen.

Early evidence, read honestly

The research is encouraging and it is young. A 2025 multicenter randomized crossover study in JMIR Formative Research compared an AI virtual patient simulator with actor-based consultation training across 396 medical students. Both improved self-rated communication skills, the actors did somewhat better and were better liked, and the AI cost roughly half as much per learner [7]. A 2026 study in JMIR Medical Education found that students who received AI-generated feedback after virtual patient interactions later scored higher on their OSCE, with a medium-to-large effect size of 0.74 [8].

Those are medical students, not Medical Affairs teams. There is no published evidence yet that this kind of practice improves real-world MSL or Medical Affairs performance, and any claim otherwise is ahead of the data.

What the technology gets wrong, and what has changed since

Researchers building a GPT-4 virtual patient reported that the system sometimes misjudged whether a learner followed a communication framework, occasionally slipped out of character, and misquoted the transcript back to the learner. They fixed much of it with better criteria and clinician review, and still described the feedback as imperfect [9]. Worth saying plainly: that work ran on GPT-4, and the models available now are far better at holding character, following long instructions, and quoting a transcript correctly. The gap between 2023 and today is enormous.

Better is not the same as verified, though. Nobody has published evidence that current models score a regulated scientific exchange the way a qualified human reviewer would, so the answer is the same either way. Check the scoring against people before it counts for anything.

Governance Decides Whether Any of This Gets Used

Adoption is running ahead of the rules

A 2026 cross-sectional survey of 367 Medical Affairs professionals across 48 countries found that only 27% were using AI for work related to KOL engagement, while close to three quarters expected AI to have a significant future impact. The same study identified real gaps in policy, governance, and training [10].

That's the practical barrier. The question isn't whether an AI application can be built. It's whether Medical Affairs, legal, regulatory, privacy, and compliance can approve it.

What approval actually requires

Every product claim and safety statement the virtual expert or the scorer uses should map to an approved source with a version, an owner, and a review date. Scenarios should be fictional and should prohibit identifiable patient information, and transcripts are employee data, so retention, access, and permitted use need deciding up front rather than drifting into place. In applicable EU deployments, Article 50 transparency provisions of the AI Act carry obligations about informing people when they are interacting with an AI system [11].

Practice and certification are not the same word

Unlimited formative practice with feedback is one thing, and a readiness decision with consequences attached is another. Before an AI score influences who goes into the field or who signs off on a launch cohort, somebody has to compare it against trained human raters and find that they agree, especially on the compliance items where a miss is costly. A number produced by a model isn't a certification until that work is done.

Medical professionals discussing a boundary-sensitive question, with a reminder that practice is not certification.

From Completion to Readiness

The question that changes

The real prize is that an organization can finally see where it breaks down. A team might learn that everyone recalls the efficacy number but a large share can't explain the study's external validity limitation when challenged, or that veterans handle the off-label question cleanly while new hires over-answer it every time. That's coaching material built from transcripts instead of from a single score, and it feeds directly into what gets built next.

None of this replaces managers or scientific experts. It gives people somewhere to fail while failing is still cheap, which is why we keep putting pilots in simulators long after they've earned the license. Before somebody walks into the room, the question stops being whether they finished the curriculum, or even whether they know the science, and becomes whether they can have the conversation.

References

  1. Medical Science Liaisons (MSL) Training: Developing First-Class Medical Science Liaisons through Proficient On-Boarding and Continuous Training Programs. Research and Markets, 2024.
  2. The medical science liaison role in Spain: a nationwide survey. PubMed, 2022.
  3. Challenges of Key Performance Indicators and Metrics for Measuring Medical Science Liaison Performance: Insights from a Global Survey. Pharmacy, 2025.
  4. Field Medical Metrics and KPIs Guidance Document. Medical Affairs Professional Society.
  5. Medical Affairs Competency Framework. Medical Affairs Professional Society.
  6. Communications From Firms to Health Care Providers Regarding Scientific Information on Unapproved Uses of Approved or Cleared Medical Products: Questions and Answers. US Food and Drug Administration.
  7. Web-Based AI-Driven Virtual Patient Simulator Versus Actor-Based Simulation for Teaching Consultation Skills: Multicenter Randomized Crossover Study. JMIR Formative Research, 2025.
  8. AI-Generated Feedback Following Virtual Patient Interactions and Medical Student Performance: Nonrandomized Quasi-Experimental Study. JMIR Medical Education, 2026.
  9. Development of a GPT-4-Powered Virtual Simulated Patient and Communication Training Platform for Medical Students. JMIR Formative Research, 2025.
  10. Perceptions, Utilization, and Impact of Artificial Intelligence on Medical Science Liaisons: A Global Cross-Sectional Survey of Medical Affairs Professionals. Cureus, 2026.
  11. Transparency obligations under Article 50 of the AI Act. European Commission.
← Back to all resources