Pulse - Value Added
FRACTIONAL CRO · MARYLAND-BASED, NATIONWIDE · $0→$200M

Kory White

RevOps & Revenue Leadership

Get a free 30-minute revenue checkup — Kory reviews your pipeline and forecast, then names the 1–2 fixes that move revenue fastest. 25 yrs scaling teams $0→$200M.

Free 30-min revenue checkup →
Hire a Fractional CROHow We Help?LinkedInRésuméCRO Syndicate
← Library
Knowledge Library · pulse-tools
13/13 Gate✓ IQ Certified10/10?

How Do I Score My Inside and Outside Reps on the Same Scale?

Pulse ToolsHow Do I Score My Inside and Outside Reps on the Same Scale?
📖 3,593 words🗓️ Published Aug 6, 2026
Direct Answer

Build one weighted scorecard both motions share. List eight or nine KPIs every seller influences — pipeline created, deal size, win rate, cycle time, attach, retention, activity — assign each a weight, score every rep 1-to-5 per line, then sum weight × level into a single composite. Inside and outside reps land on the same comparable number.

The job a shared rep scorecard is hired to do

The job is not "measure salespeople." Most organizations already measure salespeople to death. The job is to produce one defensible number that survives a room full of skeptics — a sales leader who thinks the field carries the company, an inside manager who thinks the phone team does three times the work for half the credit, a finance partner who has to fund both, and a rep who wants to know why the person two desks over got promoted.

Separate scorecards fail at exactly that moment. When inside reps are tracked on dials, connects, and monthly closed-won, and outside reps are tracked on territory revenue, key-account penetration, and quarterly bookings, there is no arithmetic that puts those two lists next to each other. Leaders end up making the comparison anyway — by gut, by tenure, by who is loudest in the pipeline meeting. That is the actual failure mode. It is not that comparison is impossible; it is that comparison happens informally and nobody can audit it.

A shared scale is hired to do four concrete jobs:

Make ranking defensible. When you stack-rank for promotion, territory expansion, President's Club, or a reduction in force, you need a number whose construction you can show. "Sarah scored 3.9, Marcus scored 3.4, here is the line-by-line" ends an argument that "Sarah just seems stronger" never will.

Make coaching specific. A composite is useless as a coaching tool, but its components are gold. A rep at composite 3.2 whose weakest line is a level 2 on win rate has a completely different development plan from a rep at 3.2 whose weakest line is a level 2 on activity. The first needs discovery and qualification work; the second needs a calendar intervention. One number for ranking, eight or nine lines for coaching.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 1

Make pay explainable. The moment comp is tied to a published composite, the "inside versus outside is unfair" conversation stops being a grievance and becomes a weighting debate — which is a debate you can actually resolve with data.

Make pivots cheap. This is the underrated one. When a product moves from a field motion to an inside motion — which happens constantly as deals commoditize, as procurement moves online, as a mid-market segment matures — you do not rebuild your performance management system. You change three weights. Both teams re-aim within a week.

The adjacent version of this problem shows up anywhere a company runs two delivery motions against one outcome: a distributor with road reps and counter sales, a manufacturer with territory managers and a desk team, an insurance agency with captive producers and a call center, a services firm with billable consultants and a proposal desk. Whenever two groups are compensated from the same revenue pool with different day-to-day work, the shared-scale problem is the same problem wearing different clothes. Everything below transfers.

Choosing KPIs both motions can actually be held to

The single biggest mistake is picking KPIs that quietly favor one motion. If you weight "meetings held" heavily, the inside team wins by structural advantage — they can book eight in a day. If you weight "average deal size" heavily, the outside team wins the same way. Neither result tells you anything about who is better at their job.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 2

The test for every candidate KPI: can a rep in either seat move this number through effort and skill, rather than through the shape of their territory? If the answer is no, either drop it or normalize it.

A workable eight-to-nine line set for most B2B teams:

Pipeline created (self-sourced). Weight it on self-sourced rather than total, or the rep who inherits an inbound-heavy territory scores high for opening their laptop. Score levels against a rep's own segment: an inside rep sourcing 3× quota coverage in SMB and an outside rep sourcing 3× in enterprise are both level 5, even though the absolute dollars differ by an order of magnitude.

Deal size versus segment median. Not raw deal size — deal size *relative to the median for that segment*. An outside rep closing $180K against a $150K enterprise median is doing the same thing as an inside rep closing $22K against an $18K mid-market median. Both are ~20% above. Both are level 4. This normalization is the single most important design choice in the whole matrix, and it is where most homegrown scorecards break.

Win rate on qualified opportunities. Measure from a consistent stage gate — usually the first stage where a buyer has confirmed a problem and a timeline. Both motions can be held to this. Guard against gaming by tracking the ratio of created-to-qualified alongside it; a rep who qualifies only sure things will show a suspiciously thin funnel above the gate.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 3

Sales cycle time versus segment baseline. Again relative. Enterprise cycles run months, inside cycles run weeks. Score the delta against the baseline for the segment, not the raw day count.

Attach or expansion rate. Multi-product attach, seat expansion, service-plan attach — whatever "sold more than the minimum viable deal" means in your business. Both motions influence this heavily and neither has a structural edge.

Retention or renewal contribution. Even where a CS team owns renewals, the seller's contribution to a clean renewal — accurate scoping, realistic promises, warm handoff — is scoreable. This line is the antidote to a scorecard that rewards closing bad deals fast.

Activity quality, not activity volume. Volume is where the fairness argument dies. Instead score the quality signal: multithreading depth (how many buyer-side contacts are engaged), next-step discipline (percentage of open opps with a scheduled next step), or CRM hygiene as a proxy for forecast reliability. An outside rep with four contacts engaged at an account and an inside rep with four contacts engaged are doing the same job.

Forecast accuracy. Underrated and beautifully motion-neutral. Score the variance between a rep's commit and their actual, over a rolling two or three quarters. A rep who calls their number within 10% is a level 5 whether they sell over Zoom or over lunch.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 4

Playbook adherence or coachability. The one subjective line most teams need. Keep it to a single line, define the levels explicitly, and require manager evidence. If you cannot define what level 3 looks like in a sentence, drop the line.

Two KPIs to deliberately *exclude*: raw call/email counts (structurally biased, trivially gamed) and total territory revenue (measures territory assignment more than rep skill). If leadership insists on including revenue attainment, use percent of quota rather than dollars — quota is where you already encoded territory difficulty.

How the scorecard fits the RevOps stack

The scorecard is not a standalone artifact. It is a computed layer that sits downstream of your CRM and upstream of your ranking, coaching, and comp processes. Getting the data flow right matters more than the tool you pick, because a beautiful matrix fed by unreliable stage data is just a confident wrong answer.

The practical build sequence in most RevOps functions: CRM is the system of record for opportunities and activity; a warehouse or reporting layer normalizes it into per-rep, per-period facts; the scorecard applies segment baselines and weights; the composite feeds three consumers — the stack-rank, the 1:1 coaching agenda, and the comp calculation.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 5

Three implementation notes that save months of rework:

Compute the composite in one place. If the ranking uses one calculation, the comp system uses a slightly different one, and a manager keeps a private spreadsheet with a third, you have recreated the exact problem you were solving. Pick one source of truth — a warehouse view, a dashboard, a purpose-built matrix tool — and make everything else read from it.

Snapshot every period. Store the scored matrix per rep per quarter, immutably. You will need it for comp disputes, for promotion committees, and for the far more valuable exercise of asking whether last quarter's high scorers actually became this quarter's top performers. If your composite has no predictive relationship with future results, your weights are wrong and only the historical snapshots will tell you.

Keep the levels ordinal and documented. A 1-to-5 level should map to a written definition ("level 4 = 15-30% above segment median"), not a manager's impression. Publish the bands with the weights. The scale earns trust from the fact that two managers scoring the same rep would land in the same place.

Where teams reach for tooling: a CRM like Salesforce can host the whole thing in custom dashboards if you build it yourself; sales-scorecard and coaching platforms such as Ambition are purpose-built for weighted multi-metric scorecards pushed to Slack and wallboards; gamification tools like Spinify or Hoopla broadcast a shared leaderboard but lean toward motivation over rigorous weighting; commission platforms like QuotaPath, CaptivateIQ, and Xactly are where the scale gets teeth if you enforce fairness through pay; conversation-intelligence tools like Gong feed behavioral signal into the quality lines. Pricing across this category ranges widely — a free tier or roughly $15/user/month at the light end, mid-tens of dollars per user per month for scorecard and gamification platforms, and custom enterprise quotes for the incentive-comp systems. Confirm current pricing directly with each vendor; published tiers change frequently.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 6

Weighting without structurally favoring either team

Weights are where the politics live, and pretending otherwise is how scorecards die in month three. The weights are a statement of strategy: they say, explicitly, what the company is paying sellers to do this year. Set them in a room with sales leadership, RevOps, and finance, and write down the reasoning next to each number.

A defensible starting distribution for a mixed inside/outside team — treat this as a shape to argue with, not a prescription:

Then run the fairness test before you publish anything: score both teams retroactively on last quarter's real data and compare the median composite by motion. If inside reps median 3.6 and outside reps median 2.9, you have not discovered that inside reps are better. You have discovered that your weights favor the inside motion. Adjust and re-run until the medians land within roughly 0.2 of each other on historical data. From that neutral baseline, forward differences in composite reflect actual performance rather than design bias.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 7

Two refinements worth the effort:

Normalize within segment, compare across. Every relative KPI — deal size, cycle time, pipeline coverage — gets its baseline from the rep's own segment. The *level* is then comparable everywhere. This is what lets a level 4 mean the same thing on the floor and in the field.

Allow a motion-specific line, capped. Some teams need one line that only applies to one motion — territory coverage for outside, speed-to-lead for inside. If you must, cap the combined motion-specific weight at 10% and make sure both motions have one. Beyond that, you are back to two scorecards wearing one header.

Resist two temptations. First, do not add a KPI every time someone complains; nine lines is already at the edge of what a rep can hold in their head, and a fifteen-line matrix stops driving behavior because no single line matters enough to change a Tuesday. Second, do not re-weight mid-quarter unless the business genuinely pivoted. Weights that move constantly teach reps to wait out the change rather than chase it. Quarterly review, annual overhaul, emergency changes only for real strategy shifts.

Rolling it out so reps trust the number

A technically perfect scorecard that reps believe is rigged produces worse behavior than no scorecard at all, because now there is a specific thing to resent. Trust is built by sequence, not by explanation.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 8

Run it silently for one quarter. Score everyone, share nothing, and check three things: do the composites correlate with who leadership already believes are the top performers? Are there surprises, and when you dig into a surprise, is the scorecard right or is it broken? Do the medians hold across motions? A quarter of shadow scoring catches almost every design flaw before it costs you credibility.

Preview individually before you publish collectively. Every rep sees their own scorecard, line by line, with their manager, before any leaderboard exists. They get to challenge the data. Some of those challenges will be correct — a miscategorized opportunity, a deal credited to the wrong seat — and fixing them publicly is the fastest way to establish that the system is honest.

Publish the construction, not just the result. The weights, the level definitions, the segment baselines, and the math. Reps do not need to like the weights to trust them; they need to be able to verify that the number was computed the way you said it would be.

Show one clean cross-motion example. The comparison that proves the scale works: an inside rep at level 5 on activity quality and level 2 on deal size, an outside rep at the reverse, landing within a tenth of each other on the composite. That single example does more to sell the concept than any deck.

Give it a quarter before it touches pay. Ranking and coaching first, compensation second. Reps will pressure-test a scorecard far more aggressively once money moves, and you want the bugs found while the stakes are reputational rather than financial.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 9

Watch for three failure signals in the first two quarters. Gaming — a KPI's distribution suddenly compresses at exactly the threshold between levels, which usually means the line is measuring the wrong thing. Manager drift — two managers scoring the subjective line an average of half a level apart, which means the level definitions are too loose. And apathy — reps who cannot name their own weakest line, which means the composite is being reported but the components are not, and you have built a ranking tool instead of a coaching tool.

A buyer and build decision framework

Before evaluating any product, answer one question: where do you want the teeth? Visibility, pay, or both. That answer determines the category, and the category determines the shortlist far more reliably than any feature comparison.

Practical evaluation criteria, in the order that actually predicts success:

Do you control the weights, and can you change them without vendor services? If re-weighting requires a support ticket and a two-week turnaround, you have lost the pivot speed that justifies the whole approach.

How Do I Score My Inside and Outside Reps on the Same Scale — figure 10

Does it read your real CRM data, or does someone key it in? Manual entry survives about six weeks. Any scorecard that depends on a human remembering to update it will be stale precisely when you need it for a comp conversation.

Can reps see their own components, not just their rank? A leaderboard that shows position without showing the lines behind it drives anxiety instead of behavior change.

Does it snapshot history? You need immutable per-period records for disputes and for validating that your weights predict anything.

Does it handle segment normalization? Many tools do weighted scoring but assume one baseline for everyone, which reintroduces the exact bias you built the scale to remove. If the tool cannot do it, do the normalization upstream in your reporting layer and feed it normalized inputs.

A note on sequence: start in a spreadsheet regardless of budget. It costs nothing, it forces the KPI and weighting conversation, and it produces the shadow-scored quarter of data you need to evaluate anything else honestly. Teams that buy first tend to adopt the vendor's default metric set, which was designed for a generic sales org rather than yours, and then spend the implementation arguing about weights they would have settled in a week with a spreadsheet. Move to a platform when maintenance cost, headcount, or the need to wire the composite directly into comp makes the manual version fragile — typically somewhere north of 30 to 40 reps, or the moment a compensation dispute makes an auditable system worth paying for.

Related questions

Should inside and outside reps have the same quota?

No — quota is where territory difficulty and segment economics get encoded, so it should differ. The scorecard uses *percent of quota* rather than raw dollars precisely so that different quotas still produce comparable levels on one scale.

How often should the weights change?

Review quarterly, overhaul annually, change mid-quarter only for a genuine strategy shift such as a product moving between motions. Weights that move constantly teach reps to wait out changes instead of chasing them.

Does this work for a team under 15 reps?

Yes, and a spreadsheet is usually sufficient at that size. The value is less about ranking — you know your ten people — and more about making coaching specific and comp conversations auditable when someone challenges a decision.

What if a rep works both motions?

Score them once on the shared scale using their blended segment baselines. Hybrid reps are actually the strongest argument for a shared scale, since separate scorecards have no coherent way to represent them at all.

How do we handle a rep who is new to the territory?

Ramping reps get scored on the same lines but ranked in a separate cohort until they clear ramp — typically two to four quarters depending on cycle length. Score from day one for coaching; exclude from the stack-rank until ramp completes.

FAQ

What is the main problem with scoring inside and outside reps separately?

Two disconnected scorecards make fair comparison arithmetically impossible, so leaders end up comparing informally — by gut, tenure, or visibility. The comparison happens either way; a shared scale just makes it auditable, coachable, and defensible when someone challenges a promotion or a comp decision.

How many KPIs should the scorecard have?

Eight or nine is the practical range. Fewer than six and single-line volatility swings the composite unfairly; more than ten and no individual line carries enough weight to change a rep's behavior on a given Tuesday. Every line should be something a rep can genuinely influence through skill and effort.

How do I keep the weights from favoring one motion?

Score both teams retroactively on last quarter's real data and compare median composites by motion. If the medians differ meaningfully, your weights encode bias rather than performance. Adjust and re-run until historical medians land close together, then treat forward differences as real.

Can reps game a weighted composite?

They can game any single line, which is why the composite spreads weight across outcomes, quality, and reliability measures. Watch for distributions that compress at level thresholds — that pattern usually means a line is measuring something a rep can manipulate directly rather than something they have to earn.

Should the composite drive compensation immediately?

No. Run it for ranking and coaching for at least one quarter first. Reps pressure-test a scorecard far harder once money moves, and you want design flaws surfacing while the stakes are reputational rather than financial. Wire it to pay once the number has survived a full cycle unchanged.

What does RevOps own in this process versus sales leadership?

Sales leadership owns the weights, since weights are a strategy statement. RevOps owns the data pipeline, the segment baselines, the calculation, the per-period snapshots, and the fairness testing. Finance validates that the composite ties cleanly to the comp plan and the quota model.

Sources

flowchart TD S["How Do I Score My Inside and Outside R"] S --> N0["The job a shared rep scorecard is hire"] N0 --> N1["Choosing KPIs both motions can actuall"] N1 --> N2["How the scorecard fits the RevOps stac"] N2 --> N3["Weighting without structurally favorin"]
flowchart LR C["How Do I Score My Inside and Outside R"] C --> H0["How the scorecard fits the RevOps stac"] C --> H1["Weighting without structurally favorin"] C --> H2["Rolling it out so reps trust the numbe"] C --> H3["A buyer and build decision framework"]

Related on PULSE

Download:
Was this helpful?  
⌬ Apply this in PULSE
Pulse CheckScore reps on the metrics that matterGross Profit CalculatorModel margin per deal, per rep, per territoryHow-To · SaaS ChurnSilent revenue killer playbook