How Do I Score My Field Techs on Sales and Service?
Score field techs on a weighted multi-KPI matrix, not one number. List the KPIs that define the whole job — first-time fix rate, callback rate, customer satisfaction, average ticket, option attach, membership sales, on-time arrival — assign each a weight, score every tech 1-to-5 per line, and rank on the composite: sum of (weight × level).
Building the end-to-end scoring process
The process runs in six stages, and skipping any one of them is why most field-service scorecards die inside a quarter. Stage one is definition: sit down with your service manager and your sales lead and write out every line a complete tech produces on a job. Stage two is weighting — each KPI gets a percentage of the composite, and the weights must total 100. Stage three is instrumentation: every KPI needs a data source you trust, whether that's your field-service management platform, your dispatch board, or a survey tool. Stage four is scoring, where each tech gets a 1-to-5 level per line based on where their actual number lands against defined bands. Stage five is the composite roll-up. Stage six is the review conversation, which is the only stage that actually changes behavior.
The critical design decision happens at stage two. If you weight revenue at 60% and quality at 40%, you have built a sales org that occasionally fixes things. If you weight quality at 80% and revenue at 20%, you have built a service org that leaves money on the truck. Most fleets that run this well land somewhere near a 50/50 split between quality KPIs and revenue KPIs, with on-time arrival and other operational lines carrying the remainder. A common starting allocation looks like: first-time fix 20%, callback rate 15%, customer satisfaction 15%, average ticket 15%, option attach rate 15%, membership sales 10%, on-time arrival 10%.
Instrumentation is where most teams get stuck, and it's worth being honest about the gap. Average ticket and membership sales come straight off invoices. Option attach rate requires that your techs actually log which options were *presented*, not just which ones sold — without presentation data you can't distinguish a tech who never offers from a tech who offers to hard customers. Callback rate requires a consistent definition: most fleets count a return visit to the same equipment within 30 days as a callback, but you have to decide whether a customer-caused return counts. Write the definition down before you score anyone on it, because the first disputed callback will otherwise blow up the whole matrix's credibility.

The feedback loop at the bottom of that flow is the part people cut. A scorecard that gets calculated but never gets published to the techs is a management report, not a performance system. Publication is what turns the matrix into a motivator: a tech who can see they're a level 5 on fix rate and a level 2 on attach knows exactly what to work on next month, and knows the path to the top of the board runs through that one line.
Where the scorecard creates or leaks revenue
The revenue case for a weighted matrix is not that it makes techs sell harder. It's that it makes the *quiet losses* visible. A fleet of twenty techs where the top five have an option attach rate three times the bottom five isn't a fleet with five stars and fifteen duds — it's a fleet where fifteen techs were never told that presenting options was part of the job, and were never measured on it. The matrix converts an invisible variance into a coachable line item.

The clearest leak point is the presented-versus-sold gap. If your data shows a tech presented options on 30% of eligible jobs and closed 60% of those, the coaching problem is presentation frequency, not closing skill. If they presented on 90% and closed 15%, the problem is the pitch, the pricing, or the customer mix. One number — revenue per job — cannot tell those two techs apart, and they need opposite coaching. This is exactly the kind of diagnostic split that RevOps thinking brings to a trade: instrument the funnel stage, not just the outcome.
The second leak is the callback tax on aggressive selling. A tech who sells a big-ticket replacement and then generates a warranty callback three weeks later has consumed a second truck roll, a second labor hour, and often a chunk of goodwill. Depending on your trade and geography, a truck roll typically runs somewhere in the range of a hundred to a few hundred dollars in loaded cost before parts. If the callback also drops a five-star review to a two-star, the downstream cost lands on your cost of customer acquisition, which nobody attributes back to that tech. Weighting callback rate into the composite is how you make that cost show up in the place where the tech feels it.
The third leak is membership churn, and it's the most under-measured. Service plans are the closest thing the trades have to recurring revenue, and a tech who sells memberships to customers who don't need them is manufacturing churn that shows up two renewal cycles later. Some fleets handle this by scoring membership *retention* on the selling tech's line for the first renewal, not just the initial sale — it's harder to instrument, but it stops the "sell anything to anyone" failure mode cold.

There's also an upstream effect worth naming. When dispatch knows the scorecard, dispatch changes. A dispatcher looking at a high-value diagnostic call with a likely equipment-age replacement conversation can route it to a tech whose matrix shows strong presentation levels, while routing a straightforward repair to a tech whose strength is speed and first-time fix. That routing logic quietly lifts revenue without asking any individual tech to change behavior at all. The scorecard becomes a capacity-allocation input, not just a performance report.
Concrete numbers, bands, and benchmarks
The 1-to-5 levels only work if the bands are written down and the same for everyone. Vague bands invite favoritism, and the first time a tech thinks the levels are subjective, the matrix loses its authority. Here's how to build them.
Start with your own trailing data. Pull twelve months of per-tech numbers for each KPI. Level 3 should sit at your fleet median — not an aspirational target, the actual middle. Level 4 sits around the 75th percentile, level 5 at roughly the 90th, level 2 at the 25th, and level 1 below that. This anchoring matters: if you set level 3 at a number only two techs have ever hit, you have built a scorecard where everyone is failing, and everyone stops caring by week three.

Worked example on a seven-KPI matrix. Say a tech scores: first-time fix level 5 (weight 20), callback rate level 4 (weight 15), customer satisfaction level 4 (weight 15), average ticket level 2 (weight 15), option attach level 1 (weight 15), membership sales level 2 (weight 10), on-time arrival level 5 (weight 10). Composite = (5×20) + (4×15) + (4×15) + (2×15) + (1×15) + (2×10) + (5×10) = 100 + 60 + 60 + 30 + 15 + 20 + 50 = 335 out of a possible 500, or 67%. That's a genuinely excellent service tech sitting in the middle of the board, and the matrix tells you precisely why: two adjacent lines, attach and membership, are dragging 55 points. That's a single coaching conversation, not a performance plan.
Run the mirror case. A tech with average ticket level 5, option attach level 5, membership level 5, but callback rate level 1, satisfaction level 2, first-time fix level 2, on-time level 3 lands at (2×20) + (1×15) + (2×15) + (5×15) + (5×15) + (5×10) + (3×10) = 40 + 15 + 30 + 75 + 75 + 50 + 30 = 315, or 63%. Lower than the "weak seller." That result is the entire point of the exercise, and you should walk your leadership team through exactly this comparison before you launch, because it is the moment where they either buy in or quietly decide to keep ranking on revenue.

On review cadence: score monthly, review with each tech monthly, and re-baseline the bands quarterly or semi-annually as the fleet improves. Monthly is frequent enough that the feedback attaches to remembered jobs and infrequent enough that a bad week doesn't wreck a level. Weekly scoring on small job volumes produces noise, not signal — a tech running eight jobs a week will swing wildly on attach rate for reasons that have nothing to do with skill.
On sample-size floors: don't score a KPI for a tech below a minimum job count in the period. Twenty completed jobs is a reasonable floor for ratio metrics like attach rate and callback rate. Below that, either suppress the line and redistribute its weight proportionally across the tech's remaining KPIs, or roll a trailing 90-day window for that tech specifically. Scoring a tech's callback rate off four jobs is how you generate a grievance.
On pay: if you wire compensation to the composite, phase it. Run the matrix in visibility-only mode for one full quarter so techs can see their levels and correct before money moves. Then attach a bonus pool to composite tiers rather than converting base pay. A structure many fleets use is a flat monthly bonus at composite thresholds — say something at 70%, more at 80%, more at 90% — layered on top of existing spiffs, so the matrix adds upside instead of threatening income. Techs will tolerate a lot of measurement if it can only help their paycheck.

Pitfalls and how to avoid them
Too many KPIs. Seven is about the ceiling. Past that, individual weights get so small that no single line moves the composite, and techs stop being able to hold the matrix in their head. If a tech can't recite roughly what they're measured on, the scorecard isn't steering behavior. Cut to the lines that actually differentiate performance.
Weights that never move. The opposite failure is a matrix set once in January and untouched all year. Seasonality is real — a heating trade in July has different priorities than in January, and a matrix that can't reflect that trains techs to ignore it. Re-weighting is cheap and should be treated as a normal operating lever. Announce the change, explain the why, and let the field re-aim.

Gaming the definition. Every scored metric gets gamed at the margins. Callback rate gets gamed by logging returns as new jobs. First-time fix gets gamed by leaving a job open. Average ticket gets gamed by cherry-picking calls, which is why dispatch fairness matters — if one tech is fed all the maintenance tune-ups and another gets all the emergency replacements, their tickets aren't comparable. Either normalize the metric by job type or audit dispatch distribution monthly and correct the imbalance.
Scoring on things a tech doesn't control. On-time arrival is partly a dispatch and routing outcome. Customer satisfaction is partly a function of whether the customer got quoted a number they hated. Keep controllable lines heavy and shared-responsibility lines light, and be willing to void a line when the cause was clearly upstream. One fairly voided score buys more credibility than ten perfectly calculated ones.
Survey response bias. If satisfaction comes from post-job surveys, watch the response rate. A tech with a 12% response rate and a 4.9 average and a tech with a 60% response rate and a 4.6 average are not comparable, and the low-response tech may simply be sending surveys only to happy customers. Set a minimum response rate for the line to score, or use a blended source that includes public reviews.

Launching without the conversation. The failure mode that kills more matrices than any technical problem: publishing the scorecard cold. Techs discover they're being measured on selling, decide management has turned them into commission reps, and the best ones start taking calls from competitors. Run the launch as a meeting, show the two worked examples above, be explicit that quality carries as much weight as revenue, and take the first round of "this line is unfair" feedback seriously enough to actually change something. Adjusting one weight based on tech input at launch buys more compliance than any amount of explanation.
No coaching capacity behind it. A matrix that identifies twelve coaching opportunities and gets acted on for two is worse than no matrix, because it proves to the field that the score has no consequences. Before launch, confirm someone owns ride-alongs and has calendar room for them. The scorecard's job is to aim coaching, and if there's no coaching to aim, you've built a report.
Choosing where the scorecard lives
Once the matrix exists on paper, you're picking an execution layer, and the honest answer is that the layer matters less than the definitions. Still, the choice determines how much manual work the thing costs you every month.

A spreadsheet is free, fully transparent, and infinitely re-weightable. Every KPI is a column, every tech a row, the composite is one formula. The real cost is maintenance and staleness — the sheet works for the first two months and then someone stops updating it after a promo change. It's the right starting point precisely because it forces you to define the bands and weights before you spend money.
A field-service management platform automates measurement off the job data. ServiceTitan builds technician scorecards straight off dispatch and tracks average ticket, membership sales, conversion, and callbacks natively, and ties into good-better-best option presentation so the offer is standardized rather than improvised — it's typically custom-quoted and priced toward the enterprise end. Housecall Pro, Jobber, and ServiceM8 cover scheduling, invoicing, and reporting for smaller fleets at published subscription tiers, and they'll give you the revenue and job inputs the composite needs even though you build the weighting yourself. FieldEdge sits in a similar space with per-tech performance dashboards. The common trade: these tools give you clean inputs and weak weighting logic, so you end up exporting to compute the composite anyway.

A visibility or gamification layer — Spinify, Ambition — puts the scorecard in front of the field in real time with leaderboards and coaching cadences. Ambition in particular is built around weighted scorecards tied to coaching, which is the closest paid analogue to the method. These add motivation, not measurement; they need a source system feeding them.
A compensation engine — QuotaPath is the common one — is where the matrix grows teeth, tracking attainment across multiple plan components so a tech can see how ticket, attach, and membership each move their check. Attach this last, after the definitions have survived a quarter.
Whatever you pick, protect two properties: the weights must be yours to change without a vendor ticket, and every tech must be able to see their own levels without asking a manager. Lose either one and the matrix stops working, regardless of what you paid for it.
Related questions
How is this different from a straight commission plan?
Commission pays on one output — revenue. A weighted matrix pays on the whole role, so quality lines can offset or cancel revenue gains. Commission tends to maximize ticket size; the matrix maximizes profitable, repeatable jobs. Most fleets end up running both: base commission plus a composite-tier bonus.
Should apprentices and senior techs share one matrix?
Same KPI list, different bands. Apprentices score against apprentice-level medians and typically carry lighter revenue weights, with first-time fix and on-time arrival weighted heavier. Recalculate bands per tier so a first-year tech isn't measured against a fifteen-year veteran's ticket average.
Can the same method score dispatchers or CSRs?
Yes, with a different KPI list. Booking rate, average hold time, first-call resolution, and revenue per booked call weight the same way. The composite math is identical — it's a general performance-scoring pattern, not a trades-only tool.
How long before the scorecard changes behavior?
Expect one quarter for awareness and a second for movement. The lines that shift fastest are behavioral and fully controllable — option presentation frequency, on-time arrival. Callback rate and satisfaction lag because they reflect work done weeks earlier.
What if my data can't support half these KPIs?
Launch with the KPIs you can actually measure cleanly, even if that's four, and reallocate the missing weight. A four-line matrix with trustworthy data beats a seven-line matrix where three lines are guessed. Add the rest as instrumentation catches up.
FAQ
What is a weighted multi-KPI scorecard?
It's a system that scores field techs on several performance indicators simultaneously — first-time fix rate, customer satisfaction, callback rate, average ticket, option attach, membership sales, on-time arrival — rather than one metric. Each KPI carries a weight, each tech gets a 1-to-5 level per line, and the composite equals the sum of (weight × level). The result reflects the full role, so no one can coast on a single easy number.
How do I choose the right KPIs and weights for my team?
Work with your service and sales leads together, not separately. Pick lines that genuinely differentiate strong from weak techs and that you can measure from a trusted source. Keep the total at 100 and keep quality and revenue roughly balanced — heavily skewing either direction produces the failure mode you're trying to avoid. Start with your own trailing twelve months of data so the bands reflect reality rather than aspiration.
Why can't I just score techs on jobs closed or revenue alone?
Because a single number can't separate two techs who need opposite coaching. Revenue alone rewards the tech who sells hard and generates callbacks, and penalizes the tech who fixes everything right the first time but never offers the upgrade. The weighted matrix scores both behaviors at once, which is the only fair way to ask a tech to do both jobs on the same visit.
How often should I update the scorecard or re-weight the KPIs?
Score and review monthly. Re-weight whenever the business priority genuinely changes — a seasonal promotion, a new membership offer, a shift in margin mix — and announce the change so the field re-aims deliberately. Re-baseline the underlying 1-to-5 bands quarterly or semi-annually as fleet performance improves, otherwise everyone drifts toward level 5 and the scorecard stops discriminating.
Will techs accept being scored on multiple metrics at once?
Generally yes, if three conditions hold: the matrix is published so everyone sees their own levels, quality is weighted at least as heavily as revenue, and the first attachment to pay is upside-only. Techs object to hidden scoring and to being turned into commission reps. They rarely object to a transparent scorecard that gives them a clear path to a bigger check.
What happens if a tech excels at service but struggles with sales?
The composite drops by exactly the weight of the sales lines, and the matrix names which lines. That's the useful outcome: instead of a vague "you should sell more," you get "your presentation frequency is level 1 — ride along Thursday and we'll work the options conversation." Many strong service techs improve fast once presenting is framed as part of the job rather than as selling.
Sources
- https://www.servicetitan.com/ — field-service management platform with technician scorecards and revenue tracking
- https://www.housecallpro.com/ — field-service software for scheduling, invoicing, and reporting
- https://www.getjobber.com/ — field-service operations software for growing service teams
- https://www.servicem8.com/ — job management software for small field teams
- https://fieldedge.com/ — field-service software with per-tech performance dashboards
- https://www.quotapath.com/ — commission tracking and quota attainment
- https://spinify.com/ — sales gamification, leaderboards, and scorecards
- https://www.ambition.com/ — sales scorecards and coaching cadences
- https://hbr.org/2003/09/coming-up-short-on-nonfinancial-performance-measurement — Harvard Business Review on non-financial performance measurement
- https://www.acca.org/ — Air Conditioning Contractors of America, quality-of-installation and service standards
Related on PULSE
- [How Many Attendants Should I Schedule Each Day at My Car Wash?](/knowledge/tl0067)
- [How Many Sales Reps Do I Need to Hire for My Logistics Company?](/knowledge/tl0058)
- [How Many Salespeople Do I Need to Hire for My Car Dealership?](/knowledge/tl0052)
- [How Many Producers Do I Need to Hire for My Insurance Agency to Grow My Book?](/knowledge/tl0015)
- [How Do I Figure Out How Many People to Schedule Each Day and at What Times for My Single Store?](/knowledge/tl0002)










