How Do I Score My Pharma Reps on Call Plan Adherence?
Score call plan adherence with a weighted multi-KPI matrix, not raw call volume. List eight or nine adherence behaviors — target-tier coverage, planned-versus-actual calls, frequency on high-tier HCPs, off-plan call rate, CRM logging accuracy — assign each a weight and a 1-to-5 level, then compute composite = sum of (weight × level). Wire bonus and coaching to that composite.
The end-to-end process from brand plan to rep scorecard
Adherence scoring is a chain, and the chain breaks at whichever link you neglect. It starts upstream, well before any rep gets scored: the brand team builds a target list from prescribing data, segments HCPs into tiers, and assigns a planned call frequency per tier — typically something like tier 1 gets 2 calls a month, tier 2 gets 1, tier 3 gets 1 a quarter. That plan is the denominator for every metric downstream. If the plan is wrong, a perfect adherence score measures perfect execution of a bad strategy, and you will have a field force that is disciplined, motivated, and pointed at the wrong doctors.
The second link is data capture. Rep activity has to land in the CRM against the same HCP master file the brand team targeted from, or planned-versus-actual comparison is apples to oranges. This is where most programs quietly fail — not on scoring math, but on record-matching. A rep calls on a practice, logs it against the practice account instead of the individual prescriber, and the tier-coverage metric shows a gap that does not exist. Before you score anyone, run a reconciliation pass: pull a month of calls, match every logged call to a target-list record, and measure your unmatched rate. Anything above 5% and you fix the data layer before you publish a scorecard, because the first rep who finds a matching error in their own numbers will spend the next quarter arguing about the tool instead of working the plan.

The third link is normalization. Territories are not equal. A rep covering rural Montana with 40 targets spread over 300 miles cannot hit the same raw call frequency as a rep covering three hospital systems in downtown Philadelphia. Score adherence as a percentage of *their own* plan, not as an absolute count, and adjust the plan itself for geography and target density rather than adjusting the score after the fact. Percentage-of-plan is the only formulation that lets you put every rep on one leaderboard honestly.
The fourth link is the composite calculation, which is deliberately simple arithmetic: each KPI gets a weight (weights sum to 100%) and each rep gets a 1-to-5 level per KPI. Multiply, sum, and you have one number between 1.0 and 5.0. Simplicity matters more than sophistication here. A rep has to be able to reconstruct their own score on the back of a napkin, because a score they cannot reproduce is a score they will not trust, and a score they do not trust changes nothing about Monday morning.
The fifth link is exception review, and it is the one most teams skip. Some misses are legitimate — an office closed for renovation, a physician on sabbatical, a hospital that suspended rep access. Build a formal exception queue where a rep flags a target as unreachable with a reason code, the manager approves or rejects within a week, and approved exceptions drop out of the denominator. Without this, reps learn the scorecard punishes things outside their control, and once they believe that, they stop treating it as a fair instrument.

The final link is the feedback loop back to the brand plan. Adherence data is diagnostic in both directions. If 80% of the field misses the same tier-2 segment, that is not a discipline problem — it is a plan problem, and the target list needs revisiting. Reading the scorecard only as a judgment of reps wastes half its value.
Where the scorecard creates or leaks revenue
The revenue case for adherence scoring rests on a single premise: the brand team's target list is a better predictor of prescription lift than a rep's personal routing preference. That premise is usually right, because the target list is built on prescribing data and the rep's route is built on who returns their calls. But the premise is not automatic, and the honest version of this argument acknowledges that a great rep with deep local relationships sometimes knows something the segmentation model does not. The scorecard should have enough room in it — through the exception queue and through a modest weight on rep-nominated targets — to absorb that intelligence rather than steamroll it.
The clearest leak is tier substitution. Reps drift toward accessible HCPs because access is the binding constraint on their day, not motivation. A tier-1 specialist with a gatekeeper and a two-week lead time costs an hour of effort per contact; a tier-3 generalist in the same building costs ten minutes. Under a raw call-count metric, the economically rational move for the rep is obvious and wrong for the brand. The weighted matrix corrects the price signal: if tier-1 coverage carries 25% of the composite and total call count carries 5%, the hour spent on the specialist is now worth more to the rep than six easy calls. You are not appealing to discipline; you are repricing effort.

The second leak is frequency collapse. Pharma promotional response tends to be non-linear — the second and third contact in a period usually carry more incremental influence than the first, because message retention builds. A rep who calls 100% of their tier-1 targets once, rather than 50% of them at planned frequency, will look fine on a coverage metric and produce far less lift. This is why coverage and frequency must be separate lines on the matrix with separate weights. Collapse them into one "adherence" number and you have hidden the exact trade-off that matters.
The third leak sits downstream in the mix between field and non-personal channels. Adherence scoring that only counts in-person calls will push reps away from email, remote detailing, and speaker-program follow-up even when those are the right channel for a given HCP — and for some access-restricted specialists, remote is the *only* channel. A modern matrix scores channel-appropriate reach, not just windshield time. The same logic shows up in adjacent RevOps contexts: a B2B SaaS team that scores only outbound dials pushes reps off LinkedIn and referral motions that convert better. The structural error is identical — scoring the medium instead of the outcome.
The fourth leak is compliance exposure, which is a revenue issue even though it does not look like one. Aggressive adherence targets create pressure to log calls that did not meaningfully happen. A "call" that was a hallway wave recorded as a detail is a data-integrity failure and, in a regulated environment where sample accountability and aggregate-spend reporting depend on the same records, a genuine risk. Build the matrix so that logging accuracy is itself a scored line, and audit a random sample of calls — signature capture, sample transactions, duration — rather than assuming volume pressure produces clean data.

The upside case is more modest than vendor material suggests, and it should be stated carefully: what a well-run adherence program reliably buys you is *predictability*. You know where promotional effort actually landed, so when the brand underperforms you can separate an execution failure from a targeting failure from a market failure. That diagnostic clarity is worth more over a product lifecycle than any single quarter's coverage lift, because it is what lets you stop repeating the same wrong fix.
Concrete numbers and how to set the weights
Start with the number of lines. Eight or nine KPIs is the working range, and the boundaries are real. Below six, the matrix is too coarse to coach from — a rep sees a low score and cannot tell which behavior to change. Above ten, weights get so thin that a full grade level on any one line moves the composite by a rounding error, and reps rationally ignore the tail. Nine lines with weights ranging from 5% to 25% is a shape that holds up.
A defensible starting distribution for a specialty field force looks roughly like this: target-tier coverage 25%, call frequency on tier-1 and tier-2 targets 20%, planned-versus-actual call match 15%, off-plan call rate 10%, CRM logging accuracy and timeliness 10%, sample and signature compliance 10%, follow-up cadence after a first contact 5%, non-personal channel reach 5%. That is eight lines summing to 100%. For a primary-care force with a broad, shallower target list, shift weight from frequency toward coverage — coverage 30%, frequency 12% — because breadth is the strategy there.

Level definitions need to be written down before anyone is scored, in plain percentage-of-plan terms so nobody argues about interpretation. A workable ladder for coverage: level 5 at 90%+ of planned targets touched in the period, level 4 at 80–89%, level 3 at 70–79%, level 2 at 55–69%, level 1 below 55%. Set level 3 at whatever your current field median actually is, not at an aspirational number. If you set the middle of the ladder above where the field lives today, the launch month produces a wall of 1s and 2s, the field concludes the instrument is rigged, and you spend your credibility before the program has done anything.
On period length: score monthly, review quarterly, pay on a rolling three-month composite. Monthly scoring gives coaching a tight enough loop to matter. Paying on a single month punishes a rep for one bad week of weather or a conference that pulled their targets out of the territory. A rolling three-month average smooths that noise without letting a genuinely disengaged rep hide for long.
On the bonus link, be conservative in year one. Putting 50% or 60% of variable comp on a brand-new composite is how you generate a grievance queue instead of a behavior change. A defensible year-one split puts something in the range of 20–30% of variable compensation on the composite, with the remainder on the outcome metrics the field is used to. Raise that share in year two once the data has survived a full cycle of rep scrutiny and you have fixed the matching errors the field will absolutely find for you.

On movement expectations: the honest answer is that the composite drops in month one, because you are now measuring things nobody was optimizing for. That drop is not a failure signal and leadership needs to be told so in advance, in writing, before the first scorecard goes out. Set the month-one expectation as a baseline, not a grade, and set the first improvement checkpoint at 90 days.
On effort: assume the alignment conversation about weights takes three to five working sessions with brand, sales leadership, and RevOps in the room, and that data reconciliation takes longer than anyone budgets. The scoring math is an afternoon. The agreement and the data hygiene are the project.
Pitfalls that quietly kill an adherence program
Scoring what the CRM happens to capture rather than what matters. Metrics get chosen by data availability, so the matrix fills up with clean-but-trivial lines — call duration, logging latency — while the hard-to-measure things that actually drive prescriptions, like message quality and objection handling, get zero weight. If a behavior matters and is hard to measure, put it on the matrix as a manager-assessed line with an explicit rubric rather than dropping it. A subjective line with a clear rubric beats an objective line that measures nothing.

Publishing rankings before publishing the method. If reps see their position on a leaderboard before they see the weights, the level definitions, and the exception process, the program is read as surveillance. Publish the matrix first, run it in shadow mode for a full period with no consequences attached, let the field find the errors — and they will find real ones — then turn on the scoring.
Letting the target list go stale. Adherence to a plan built on year-old prescribing data measures obedience, not effectiveness. Re-tier at least twice a year, and immediately after a competitor launch, a label expansion, or a formulary change. This is the single most common failure: the scoring machinery is maintained meticulously while the plan it enforces slowly decays.
Treating the composite as a verdict instead of a pointer. The composite exists to route attention. The coaching conversation is never "your composite is 3.2"; it is "your composite is 3.2 because tier-1 coverage is at 2.0, here are the specific targets you have not touched in six weeks, and two of them have an access issue we should solve together." A manager who reads the number aloud and stops there has automated a bad review, not improved anything.
Over-fitting to gaming behavior. When reps find a loophole — logging a drive-by as a detail, front-loading calls into the last week of the period — the instinct is to add a rule. Three quarters of that and the matrix has fourteen lines and a rulebook nobody reads. Prefer fixing the incentive shape over adding rules: if end-of-period stuffing is the problem, score frequency distribution across weeks rather than banning late calls.

Ignoring the manager layer. District managers are the transmission mechanism, and if they are not scored on how they use the matrix — coaching notes filed, ride-alongs targeted at low lines, exceptions cleared on time — the scorecard stays a report the reps receive rather than a system the org runs.
Assuming the pharma case is unique. It is not, and it helps to say so internally. Medical device teams score account-plan adherence, field service orgs score preventive-maintenance route compliance, and enterprise sales orgs score named-account coverage — same structure, same failure modes. Borrowing a working comp design from an adjacent function inside your own company is usually faster than building from scratch.
A selection checklist for tooling and rollout
Build the matrix before you buy anything. Every platform in this space works better against a defined methodology, and a tool selected first will quietly impose its own scoring model on you. A spreadsheet is sufficient to prove the matrix and force the weighting argument into the open — and that argument, between the VP who wants volume weighted and the brand director who wants tier coverage weighted, is the actual deliverable of the first month.

Then decide where the teeth need to live, because that determines the category. If your problem is data fidelity and planned-versus-actual capture, the answer is your life-sciences CRM layer — Veeva CRM and IQVIA's Orchestrated Customer Engagement are the established options, and Salesforce Health Cloud is viable if you already run Salesforce and are willing to build the scorecard as custom dashboards. If the problem is that nobody sees their score, the answer is a visibility and gamification layer such as Spinify, Ambition, or Hoopla. If the problem is that the score does not reach the paycheck, a commission platform like QuotaPath makes the composite-to-comp link explicit. If the problem is that the target list itself is wrong, no scoring tool helps — that is an HCP analytics problem, and Komodo Health sits in that category. Pricing across these varies widely and most enterprise life-sciences platforms quote rather than publish, so treat any number you hear secondhand as unverified.
Two requirements are non-negotiable regardless of category. First, you must be able to change the weights yourself, overnight, without a ticket — brand priorities shift on competitor launches and formulary decisions, and a matrix that takes six weeks of IT work to re-weight is a matrix that will be wrong for six weeks. Second, reps must be able to see their own levels and the gap to the next one in near-real time. A score revealed at the quarterly review is a historical document; a score visible on Tuesday is a decision input.
Sequence the rollout: build the matrix in a spreadsheet, shadow-score one district for a full period, fix the matching errors that district finds, publish to the full field with no comp consequence for one more period, then attach the bonus. Two periods of running visibly-but-harmlessly is the cheapest insurance available against a credibility failure you cannot undo. RevOps owns the mechanics of this — the data pipeline, the reconciliation, the weight governance — while brand owns the target list and sales leadership owns the comp link. Naming those three owners explicitly at kickoff prevents the most common organizational failure, which is a scorecard everyone uses and nobody maintains.
Related questions
How often should we re-weight the adherence matrix?
Quarterly as a default rhythm, plus immediately after any brand-plan change — a competitor launch, label expansion, or formulary shift. Re-weighting more often than monthly makes the target unstable and reps stop steering toward it, because the destination moves faster than they can travel.
Should tenured reps be scored on the same matrix?
Same lines, same weights, different level expectations only if territory plans genuinely differ. Tenured reps often have harder territories, which should be reflected in the plan itself rather than in a softer score. Two matrices for two groups destroys comparability and invites accusations of favoritism.
What if a rep's territory has fewer high-tier targets?
Score percentage-of-own-plan, never absolute counts. A rep with 18 tier-1 targets and a rep with 45 are both measured against their own planned coverage and frequency. If the plans themselves are unbalanced, fix the territory alignment — that is a plan problem, not a scoring problem.
Can this approach work for a contract sales organization?
Yes, and adherence scoring matters more there because you are buying execution against a plan you designed. Write the composite into the contract's service levels, require the same reconciliation reporting, and audit logged calls at a higher sample rate than you would internally.
How does adherence scoring interact with sample accountability?
Keep them as separate scored lines. Sample compliance is a regulatory obligation with its own audit trail; adherence is a commercial-execution metric. Blending them lets a strong compliance record mask a coverage failure, which is exactly the substitution the matrix exists to prevent.
FAQ
What is actually wrong with scoring reps on raw call volume?
Raw volume rewards accessibility rather than strategic value. A rep logging fifteen calls on low-decile prescribers outranks a rep logging ten on planned high-tier specialists, even though the second rep executed the brand plan and the first did not. The number is not merely incomplete — it actively points effort in the wrong direction, because reps optimize whatever you display. Adding tier coverage and planned-versus-actual as separately weighted lines restores the correct price signal on rep effort.
How do I stop reps from gaming the composite?
Assume gaming and design for it rather than policing it after the fact. The three common exploits are logging low-effort contacts as details, stuffing calls into the end of a period, and flagging reachable targets as exceptions. Counter each structurally: audit a random call sample against signature and sample records, score frequency distribution across weeks rather than period totals, and require manager approval with reason codes on exceptions. Adding rules invites new loopholes; changing what the score rewards closes them.
What is a realistic timeline from decision to first scored period?
Plan on roughly a quarter. Three to five sessions to agree weights and level definitions, several weeks of data reconciliation to get the unmatched-call rate under control, one shadow period with no consequences, then go live. Teams that compress this to a month almost always skip reconciliation, and the resulting matching errors surface in the first published scorecard, which is the worst possible moment for the field to find them.
Should the composite be visible to the whole field or only to each rep?
Each rep should see their own lines and levels in detail. Whether the full ranking is public is a culture decision with a real trade-off: public leaderboards drive urgency in competitive teams and demoralize the bottom quartile in others. A middle path that works well is publishing anonymized distribution — where each rep sits relative to the median — so everyone has context without a named bottom of the list.
How does this connect to broader RevOps practice?
It is the same discipline applied in a regulated setting: define the plan, instrument the activity, weight the behaviors that lead outcomes, and connect the composite to compensation and coaching. The pharma specifics — tiering, sample accountability, access restrictions, aggregate-spend reporting — change the KPI names, not the architecture. RevOps teams that have built named-account coverage scorecards in B2B will recognize the entire structure.
What single mistake sinks these programs most often?
Letting the target list decay while maintaining the scoring machinery perfectly. An immaculate adherence score against an eighteen-month-old tier list measures compliance with a stale strategy and produces confident, well-documented, misdirected effort. Re-tier on a fixed cadence and after every market event, and treat widespread misses on one segment as evidence the plan is wrong before concluding the field is undisciplined.
Sources
- U.S. Food and Drug Administration — Prescription Drug Advertising and Promotional Labeling
- PhRMA Code on Interactions with Health Care Professionals
- Centers for Medicare & Medicaid Services — Open Payments Program
- Office of Inspector General — Compliance Program Guidance for Pharmaceutical Manufacturers
- Veeva Systems — Life Sciences CRM
- IQVIA — Orchestrated Customer Engagement
- Salesforce Health Cloud
- Harvard Business Review — Motivating Salespeople: What Really Works
- McKinsey & Company — Life Sciences Insights
Related on PULSE
- [How Many Sales Reps Do I Need to Hire for My Logistics Company?](/knowledge/tl0058)
- [How Many Salespeople Do I Need to Hire for My Car Dealership?](/knowledge/tl0052)
- [How Many Producers Do I Need to Hire for My Insurance Agency to Grow My Book?](/knowledge/tl0015)
- [How Do I Figure Out How Many People to Schedule Each Day and at What Times for My Single Store?](/knowledge/tl0002)
- [How Many Attendants Should I Schedule Each Day at My Car Wash?](/knowledge/tl0067)










