Skip to content

How to Build an HCP Question Set: Voices, Moments, and Why Keyword Lists Fail

Keyword lists fail as AI probes. Build an HCP question set for AI monitoring from voices crossed with moments: six probe families, one guardrail, governance.

The Juncture team8 min read
AI answer monitoringHCP engagementprompt designShare of Answerpharma marketing

Every AI answer monitoring program stands on one input you fully control: the questions you ask. The engines are not yours. The model updates are not yours. The sources the machine reads are mostly not yours. The probe set is. It is the instrument, and the readings are only as good as the instrument.

Most teams build that instrument the way they built SEO keyword lists a decade ago. Brand name plus "dosage". Brand name plus "side effects". Brand name plus "vs alternatives". Run those and you get answers back, and scores, and a dashboard that trends nicely. What you do not get is any resemblance to the answers your actual audience is reading, because nobody talks to an assistant in keywords. A clinician with a decision to make types a sentence with a patient in it. A hospital pharmacist types a sentence with a protocol in it. A payer types a sentence with a budget in it. If your probes do not sound like those sentences, you are carefully monitoring a channel that does not exist.

The fix is not a longer list. It is a different construction method. A durable HCP question set is built as voices crossed with moments: who is asking, multiplied by when they ask. This article covers why phrasing changes the answer, how the voices and moments grid works, the six probe families every pharma set should cover, the one guardrail that keeps the exercise safe, and the governance step most teams skip.

Why phrasing changes the answer

The premise behind keyword-style probes is that a model answers the topic, so the wording is cosmetic. The research says the opposite. In a study of prompt sensitivity, subtle formatting changes alone, with the meaning held constant, swung model task accuracy by up to 76 points. Not the question. The formatting of the question.

It gets sharper in medicine. An MIT team perturbed patient messages with typos, extra whitespace, uncertain phrasing, and removed gender markers, then watched treatment recommendations move. Across nine types of altered messages, models showed a 7 to 9 percent increase in recommending patients self-manage rather than seek care, and the same study found roughly 7 percent more errors for female patients. Nothing clinical changed in those messages. The asker sounded different, so the answer was different.

That is the mechanism your question set has to respect. Assistants adapt register and content to the person they infer they are talking to. Phrase a question like a clinician and you get guideline language, monitoring parameters, and a dosing table. Phrase the same underlying question like a worried patient and you get simplified reassurance with the hedges moved around. Both answers exist. Both carry your brand. A keyword probe retrieves neither, because it sounds like no one.

And the audience is asking in its own registers at scale. An AMA survey of nearly 1,700 physicians found 81 percent now use AI professionally, more than double the 2023 rate, and Doximity's physician cohorts show AI use rising from 47 percent in April 2025 to 63 percent by early 2026, with literature search the most common use. Those clinicians are not typing "brand + dosage". They are describing a patient.

Voices crossed with moments, not keywords

Build the set on a grid with two axes.

Voices are who is asking. For most brands, five cover the field: a treating specialist, a hospital pharmacist, a payer preparing a coverage position, a patient, and a caregiver. Each voice carries its own vocabulary, its own concerns, and its own register, and the engines respond to all three.

Moments are when they ask. Five again: initial diagnosis, considering a switch, a safety review, checking an interaction, and preparing a formulary decision. A moment is a decision context, and the decision context changes what a good answer looks like far more than the topic does.

Cross them and you get a grid of up to twenty-five cells. Not every cell is live: a caregiver rarely prepares a formulary decision. But each live cell yields one to three questions phrased the way that person would actually ask at that moment, in their words, at their reading level, with their stakes attached. A specialist at the switch moment asks about washout and monitoring. A payer at the formulary moment asks about comparative value and the evidence behind it. A caregiver at the interaction moment asks whether the new tablet is safe next to the pills already in the cabinet.

Two properties make the grid durable. It is complete by construction: an empty cell is visible, so you notice the payer questions you never wrote, whereas a keyword list hides its gaps. And it is stable: voices and moments change slowly even as language shifts, so you can rephrase a question inside a cell without losing the thread of what that cell has measured for a year.

The six probe families a pharma set should cover

Once the grid is drafted, audit it against six families. Every pharma question set should touch all six.

Diagnosis questions. Where does the condition get recognized and treatment initiated? These reveal whether you appear at all when the category is framed clinically, before any brand is named.

Safety questions. Adverse events, contraindications, warnings, use in special populations. These are the answers where a paraphrase dropping one clause becomes a compliance exposure, so they deserve the densest probing.

Competitor comparison questions. "Is X better than Y for this patient" is one of the most natural assistant questions there is, and the engines answer it every day. You want to know how that comparison reads before a formulary committee does.

Dosing questions. Initiation, titration, renal and hepatic adjustment, missed doses. Precise, checkable, and the easiest place to catch a model asserting numbers your label does not contain.

Patient-style questions. Lowercase, misspelled, anxious, compound. More than 40 million people ask ChatGPT healthcare questions every day, and almost none of them write like a medical affairs reviewer. If every probe in your set is grammatical, you are not sampling the channel as it exists.

Off-label-trap questions. Questions a curious user would genuinely ask that invite the model toward unapproved territory: the adjacent indication, the pediatric use, the combination nobody studied. You never assert the off-label use in the question. You ask the open question and score whether the engine volunteers it, because that volunteered sentence is the single most expensive drift you can catch.

One moment, three voices: Varigel at the switch

Take a fictional brand, Varigel, and hold one moment fixed: considering a switch from current therapy. Render it in three voices.

The treating specialist asks: "I have a patient inadequately controlled on first-line therapy and I am considering a switch to Varigel. What washout is needed and what should I monitor in the first month?" The engine answers in clinical register: transition guidance, monitoring parameters, maybe a dosing table. The thing to score is whether that protocol matches the label or was assembled from a review article and a forum thread.

The hospital pharmacist asks: "We are getting switch requests to Varigel on our unit. What does the label say about transitioning from other agents in the class, and what interactions matter in patients on multiple therapies?" The answer goes class-level: comparisons across agents, interaction mechanics, often a competitor's data sitting inside your brand's answer. The thing to score is the comparison itself, and whether the safety language survived it.

The patient asks: "my doctor wants to switch me to varigel, is it better than what im taking now? will the side effects be worse". The answer simplifies, reassures, and sometimes rounds a hedged claim up into a promise the label never made. The thing to score is what got dropped on the way down to plain language.

One moment. Three questions. Three different answers, each needing its own score against the same approved label. A keyword probe like "Varigel switching" would have retrieved a fourth answer that none of these three people will ever read.

The public-questions-only guardrail

One rule holds the whole method together: every question in the set must be built from public knowledge, phrased as something any member of your audience could ask. No pipeline data, no unpublished results, no internal strategy, no confidential document pasted in for context. The strategy lives in which questions you choose and how you score the answers, and both of those stay on your side. The test is simple: if the question itself appeared in a screenshot tomorrow, it should read as an ordinary user asking an ordinary thing. Monitoring is an observation discipline. The label you grade against never crosses the boundary, and neither does anything else you would not say out loud.

Review before run

A probe set is published words. A badly phrased probe can itself imply an off-label use or embed an unapproved claim in its premise, which means question drafting needs the same discipline as content drafting, scaled down. The working pattern: drafted questions land in a review queue, a reviewer who knows the label checks each one for phrasing, premise, and the public-questions rule, and nothing probes an engine until it is approved. Version the set when it changes so trend lines stay comparable, and re-review it when the label changes, because a question that was safe under the old indication may be a trap under the new one. This is a lightweight gate, hours not weeks, and it is the difference between a monitoring program and a liability generator.

Where this leaves you

The question set is the one part of AI answer monitoring you fully control, so build it like an instrument: voices crossed with moments, audited against the six families, phrased the way real people ask, public-only, and reviewed before anything runs. From there the method continues as measurement: how to monitor AI answers covers running and scoring, the Share of Answer metric covers the number you trend, and the HCP behavior data covers why the audience is already on the other end of these probes.

This is also exactly how Juncture's Answer Monitor builds its probe sets. Its Prompt Creator treats voices and moments as first-class objects: pick the voices, pick the moments, draft or import the questions each cell suggests, and every draft lands in a review queue where nothing runs against ChatGPT, Gemini, Perplexity, Google AI Overviews, or Claude until a reviewer approves it. Bring one brand and one moment. We will render it in three voices and show you three answers you have never seen.

Sources

  1. Sclar, Choi, Tsvetkov and Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design," arXiv, 2023. arxiv.org
  2. MIT News, "LLMs factor in unrelated information when recommending medical treatments," June 2025. news.mit.edu
  3. American Medical Association, "More than 80% of physicians use AI professionally, AMA survey," 2026. ama-assn.org
  4. Doximity, "2026 State of AI in Medicine Report," 2026. doximity.com
  5. Fierce Healthcare, "40M people use ChatGPT to answer healthcare questions, OpenAI says," 2026. fiercehealthcare.com

People also ask

Questions this raises

What are voices and moments in an AI question set?
Voices are who is asking: a treating specialist, a hospital pharmacist, a payer, a patient, a caregiver. Moments are when they ask: initial diagnosis, considering a switch, a safety review, checking an interaction, preparing a formulary decision. Crossing them produces a grid, and each live cell yields questions phrased the way that person would actually ask at that moment. The grid makes gaps visible and keeps the set stable over time, which a flat keyword list cannot do.
Why do keyword lists fail for AI answer monitoring?
Because nobody talks to an assistant in keywords, and assistants adapt both register and content to the asker. Research has shown model accuracy swinging by up to 76 points from prompt formatting alone, and a medical study found 7 to 9 percent shifts in treatment recommendations when only the style of a patient message changed. A probe like "brand + dosage" retrieves an answer that no real clinician, pharmacist, or patient will ever read, so the scores it produces do not describe the channel.
What question types should a pharma AI monitoring set cover?
Six probe families: diagnosis questions, safety questions, competitor comparison questions, dosing questions, patient-style questions written the informal way patients actually type, and off-label-trap questions. The last family asks open questions a curious user would genuinely ask and scores whether the engine volunteers an unapproved use, because volunteered off-label content is the most expensive drift to miss.
Is it safe to send monitoring questions to AI models?
Yes, if you hold the public-questions-only guardrail: every probe is built from public knowledge and phrased as something any member of your audience could ask. No pipeline data, unpublished results, or internal strategy ever goes into a prompt. The approved label you grade answers against stays on your side, so monitoring remains an observation discipline where the only thing crossing the boundary is a question the audience was going to ask anyway.
How should teams govern an HCP question set before running it?
Route every drafted question into a review queue, and let nothing probe an engine until a reviewer who knows the label approves it. The review checks phrasing, premise, and the public-questions rule, because a badly phrased probe can itself imply an off-label use. Version the set whenever it changes so trends stay comparable, and re-review it whenever the label changes.

See it on your brand

See Juncture run on your brand.

Bring an asset and a brand. We will pre-check the asset against the label and show how the machine answers about the brand today, inside and out.