Skip to content

Reference quality: the silent risk in your pharma claims library

Pharma claims library reference quality: five checks for missing, stale, or non-substantiating references, and why label changes make currency a moving target.

The Juncture team9 min read
reference qualityclaims librarySmPCsubstantiationMLR

Every claim in a pharma claims library rests on a reference. The claim is the visible half of the record: the approved statement, its status, its owner, its lineage through medical, legal, and regulatory review. The reference is the load-bearing half: the evidence that makes the statement sayable at all. Strip the reference away and a claim is just a sentence somebody once approved the sound of. Yet in most libraries the references are the least examined objects in the system. The claims are governed, versioned, and workflowed. The substantiation underneath them is assumed.

That assumption deserves scrutiny, because the most rigorously refereed citation environment we have, peer-reviewed medical literature, fails at a measured rate. A 2017 review in PLOS One recalculated quotation accuracy across the medical literature and found that 14.5 percent of cited assertions in medical research articles are quotation errors, and that among the content errors, 64.8 percent are major, meaning the referenced source fails to substantiate, is unrelated to, or contradicts the assertion citing it. Those are professional authors with named accountability, writing for reviewers whose job includes checking exactly this. A claims library assembled over years, across agencies, markets, and system migrations, with nobody assigned to re-verify substantiation, is not beating that baseline. It is simply not measuring itself.

The risk is silent because every failure mode leaves the library looking intact. The claim still renders. The reference field still contains something. The audit trail still shows an approval. What has quietly gone wrong sits one layer down, and it takes one of five specific forms.

The five ways a reference quietly fails

The reference is missing entirely. The claim was approved in a meeting where everyone knew the substantiation, and the knowing never became a record. Or the reference existed in the previous system and did not survive the migration. Either way the library now asserts something it cannot show evidence for, and no daily workflow ever surfaces that, because assets reference the claim, not the claim's paperwork.

The reference points at a document, not a place in one. "Varigel SmPC" is not a reference. It is a homework assignment. An SmPC runs to dozens of pages across ten sections; a reference without a section and paragraph pointer forces every future verifier to re-derive the original reviewer's reading from scratch. In practice nobody does, so the citation is trusted rather than checked, which means it is not really functioning as a citation at all.

The numbers no longer match. The claim says 42 percent. The source reports 38.7 percent, in a subgroup, at a different timepoint. Number drift creeps in through rounding, through copy-paste across asset generations, through a well-meant harmonization of figures between markets. Each step felt small. The distance between the claim and its evidence is now material, and nothing in the library measures that distance.

The label has moved and the citation has not. The reference says "SmPC, section 5.1" and that was true when it was recorded. The label has since been revised: the population narrowed, a warning added, an efficacy description requalified. The reference is not wrong about which document to consult, it is wrong about which version, and version is the part nobody wrote down. This is the failure mode that ages into every library by default, because labels move and reference records do not.

The hierarchy is inverted. A claim about the indication cites a congress abstract. A safety statement cites the pivotal publication instead of the label that distilled it. In the EU, the SmPC forms an intrinsic and integral part of the marketing authorisation, and its content cannot be changed except with the approval of the originating competent authority. For statements about what the product is approved to do, the label is the agreed position and the publication is commentary. Citing the publication where the label should stand swaps a legally settled source for a merely persuasive one.

Five checks that turn reference quality into a measurement

Each failure mode implies its own test, which is what makes reference quality tractable: it decomposes into five checks that can be run claim by claim.

  1. Reference present. Does every claim carry at least one attached, resolvable reference? No reference, no pass, regardless of how obviously true the claim feels.
  2. Cited section substantiates. Does the specific cited section, at the page and paragraph level, actually support the claim's meaning? A reference to a whole document scores as unverified until a substantiating section is found and pinned.
  3. Numbers match. Does every figure in the claim, the percentage, the timepoint, the population, appear in the cited source with the same value and the same scope?
  4. Label version current. Is the cited label version the version the regulator currently publishes? If the label has been revised since the citation was recorded, the claim needs re-verification against the current text, even if nothing else changed.
  5. Evidence hierarchy respected. Do label-derived statements, indication, dosing, safety, cite the label itself, with publications in supporting roles rather than substituting ones?

Each check is binary at the claim level, which means the results aggregate honestly: a library of 300 claims produces a pass rate per check, a ranked list of failures, and a number that can be tracked quarter over quarter. Reference quality stops being a vague anxiety and becomes a score with addresses attached.

Why label currency is a moving target

The fourth check deserves its own treatment, because it is the one that can never be run once and considered done.

A label is not a fixed document. The EMA is explicit that SmPCs are a key part of the marketing authorisation of all medicines authorised in the European Union, the basis of information for healthcare professionals, and kept updated throughout the lifecycle of a medicine as new efficacy or safety data emerge. On the US side the volume has been quantified: roughly 400 to 500 FDA label changes occur every year, most arising from spontaneous post-marketing reports. Spread that across a portfolio and the conclusion is uncomfortable: for any mature brand, the probability that the label text behind a years-old citation is still the current text approaches zero. A claims library that recorded its references at launch and never re-checked them is stale by construction, silently, with every claim still wearing its original approval.

The same regulatory machinery that creates the problem also makes it checkable. Regulators publish the current approved label as a matter of course: the EMA publishes the product information for centrally authorised medicines, and the National Library of Medicine's DailyMed describes itself as the official provider of FDA drug label information, carrying the most recent labeling submitted to the FDA and currently in use. Currency is therefore not an archaeology project. It is a lookup: compare the version your reference records against the version the regulator publishes today. That this comparison is possible, cheap, and almost never performed is the purest expression of the silent risk.

A worked example: Varigel

Varigel is a fictional specialty brand with a claims library of about 120 claims, built at launch and extended through two label updates. Run the five checks across it and two claims tell the story.

The first claim reads: "Varigel is indicated for adults with moderate to severe disease who have had an inadequate response to first-line therapy." Its reference field says "Varigel SmPC." Check one passes: a reference is present. Check two cannot even begin: there is no section pointer, so there is nothing specific to verify. The checker reads the current SmPC, finds that section 4.1 carries the indication wording, confirms the claim's phrasing tracks it faithfully, and pins the result: SmPC section 4.1, the version identifier of the current label, the date checked. The claim's status does not change. Its evidence does. The next reviewer who touches this claim inherits an address instead of a document, and check four now has a version to watch.

The second claim reads: "Response to Varigel is maintained through 52 weeks." Its reference field says "Varigel SmPC, section 4.2." It looks better groomed than the first claim: it has a pointer. But section 4.2 is posology, dosing and administration, with no durability data at all. The checker searches the rest of the label and finds the 52-week data in section 5.1, where it is reported for the biologic-naive subpopulation, a narrower group than the claim's unqualified phrasing. This is not a typo to silently repair. The cited section does not substantiate the claim, and the section that comes closest supports a narrower statement than the one approved. The finding goes to the medical reviewer with the discrepancy spelled out: cited section, actual content, candidate section, scope gap. The reviewer decides whether the claim gets requalified, re-referenced, or retired. That division of labor is the point. The checker finds; the reviewer judges.

One library, two claims, and both would have sailed through another year of reuse. The first was trusted because the document was right, even though the reference did no work. The second was trusted because it looked precise, even though it pointed at the wrong place. Neither failure announces itself, which is exactly why the checks have to.

Where this leaves you

A claims library is only as strong as its references, and reference strength is not a feeling. It is five checks: present, substantiating, numerically matched, current, and hierarchically sound. All five run against materials you already have: claims you already govern and labels the regulator already publishes. The only missing ingredient in most organizations is that nobody owns the run.

This is work Content Intelligence is built to do continuously: it reads claims and references from the system of record, runs the five checks, watches the published label so currency failures surface when the label moves rather than when an audit does, and routes every finding to the MLR reviewer with the evidence attached. It backs the reviewer's judgment, it never replaces it. The full framework, with the scoring rubric and the failure taxonomy, is published open and citable in our reference quality report at /research/reference-quality-framework. Start with your ten most reused claims. If all ten survive all five checks, your library is stronger than most. If they do not, you have just found the silent risk while it is still silent.

Sources

  1. Mogull, S. A., "Accuracy of cited 'facts' in medical research articles: A review of study methodology and recalculation of quotation error rate," PLOS One, 2017. journals.plos.org
  2. European Commission, "A Guideline on Summary of Product Characteristics (SmPC)," Revision 2, September 2009. health.ec.europa.eu
  3. European Medicines Agency, "How to prepare and review a summary of product characteristics." ema.europa.eu
  4. Kircik, L. et al., "United States Food and Drug Administration Product Label Changes," The Journal of Clinical and Aesthetic Dermatology, 2016. pmc.ncbi.nlm.nih.gov
  5. National Library of Medicine, DailyMed. dailymed.nlm.nih.gov

People also ask

Questions this raises

What is reference quality in a pharma claims library?
Reference quality is the degree to which every claim in a pharma claims library is backed by evidence that actually substantiates it. It decomposes into five checks: a reference is present, the cited section substantiates the claim, the numbers match the source, the cited label version is current, and the evidence hierarchy is respected so label-derived statements cite the label itself. Most libraries govern claims tightly but never re-verify references, which makes reference quality the silent layer of risk under an otherwise well-run library.
What are the most common ways references fail in a claims library?
Five failure modes cover most of it: the reference is missing entirely; it points at a whole document with no section or paragraph pointer; the numbers in the claim no longer match the cited source; the label has been revised since the citation was recorded, so the reference is silently stale; and the evidence hierarchy is inverted, with a publication cited where the approved label should be. All five leave the library looking intact, which is why they persist.
How do you check whether a claim's cited label version is still current?
Compare the version recorded in the reference against the version the regulator currently publishes. The EMA publishes the product information for centrally authorised medicines and keeps SmPCs updated throughout a medicine's lifecycle, and DailyMed, the National Library of Medicine's service, carries the most recent FDA labeling currently in use. Because the current approved label is public, currency is a lookup rather than a judgment call, but it has to be repeated, since labels are revised many times over a product's life.
Why should promotional claims cite the label rather than a publication?
In the EU the SmPC forms an intrinsic and integral part of the marketing authorisation and sets out the agreed position on the product; its content cannot be changed except with the approval of the competent authority. For statements about indication, dosing, or safety, the label is the legally settled source and a publication is commentary on it. Citing a congress abstract or a journal article where the label should stand inverts the evidence hierarchy and substitutes a persuasive source for an authoritative one.
Does automated reference checking replace MLR review?
No. The checks find discrepancies: a cited section that does not substantiate, a number that does not match, a label version that has moved. What happens next, requalifying, re-referencing, or retiring a claim, is a medical and regulatory judgment that stays with the MLR reviewer. Juncture's Content Intelligence runs the five checks continuously and routes findings to the reviewer with the evidence attached; it backs the reviewer, it never replaces them.

See it on your brand

See Juncture run on your brand.

Bring an asset and a brand. We will pre-check the asset against the label and show how the machine answers about the brand today, inside and out.