How Do I Get My Loan Officers to Hit Funded-Loan Goals?
Score loan officers on funded volume and the pipeline that produces it — applications, pull-through rate, cycle time, document accuracy, and referral-partner activity — not applications alone. Build a weighted scorecard, assign each line a weight and a 1-to-5 level, publish it, and tie both coaching and variable pay to the composite.
Signals you actually need this
Most lending teams do not go looking for a scorecard. They discover they need one when a specific pattern shows up in the numbers, and the pattern is almost always the same: activity looks healthy at the top and thin at the bottom. The branch takes plenty of applications, the pipeline report looks full on Monday, and funded volume at month-end lands well under plan. When you see that gap for two consecutive months, the problem is rarely effort. It is that nobody is being measured on the part of the job between application and funding.
Here are the concrete signals worth watching for, in rough order of how often they show up.
Pull-through is drifting below your historical band. Pull-through — funded loans divided by applications taken in a period — is the single most diagnostic number in this whole conversation. Every shop has its own normal depending on channel and credit box; a retail purchase-heavy team runs a very different number than a refi-driven one, and a broker shop differs again from a bank's direct channel. What matters is your own trailing baseline. If your twelve-month rolling pull-through is running at some steady level and one officer sits fifteen or twenty points under the team, that officer is doing the visible half of the job and skipping the invisible half. They are not lazy. They are optimizing for the number they get praised for in the Monday meeting.
Cycle time varies more between officers than between loan types. You should expect a purchase file to take longer than a straightforward rate-and-term refi, and a self-employed borrower's file to take longer than a W-2 salaried one. That is normal variance driven by the work itself. What is not normal is two officers with comparable mix showing a two-week spread in average days from application to clear-to-close. That spread is a process signal — one officer is front-loading documentation and setting borrower expectations, the other is passing incomplete files downstream and letting processing chase them.

Your processors and underwriters can name the problem officers without looking anything up. This is the most reliable qualitative signal there is, and it costs nothing to collect. Ask three processors, separately, which officers submit files they have to send back. If all three name the same two people, you have your document-accuracy data before you have built a single dashboard. Operations always knows. They just have not been asked in a way that turns into a number.
Rate environment shifts and the team does not re-aim. When rates move meaningfully in either direction, the profitable work changes — refi volume appears or evaporates, purchase competition intensifies, referral partners become more or less valuable. If your officers are still running the same daily playbook six weeks after the shift, your measurement system is not communicating priority. It is measuring last quarter's business.
Turnover is concentrated in your middle performers. Top producers rarely leave over measurement; they leave over pay or over a competitor's product set. But middle performers leave when they feel the scoring is arbitrary — when they believe they are working hard on things that do not count. A published matrix fixes that specific complaint, because it makes the scoring legible before the review rather than after it.
Nobody can answer "what does a 4 look like?" Ask your branch manager to describe the difference between a solid officer and an excellent one without using the word "volume." If the answer takes more than a sentence and includes hedging, the standard exists only in that manager's head. It cannot be coached, it cannot be inherited by the next manager, and it cannot be defended in a compensation conversation.

Any two of these together justify building the scorecard. All six means you are already paying the cost of not having one — you are just paying it in fallout instead of setup time.
What good looks like versus what bad looks like
The failure mode almost every lender starts with is the single-metric scoreboard. Applications taken, posted publicly, celebrated weekly. It is easy to measure, easy to explain, and it produces exactly the behavior you would predict: officers take every application that walks through the door, including the ones with no realistic path to funding, because the application itself is the scored event.
A good system scores the chain. A bad system scores the first link and hopes the rest follows.

Concretely, here is what separates the two.
Bad: one number, weekly, public. Applications taken. The leaderboard rewards the officer who takes twenty-two applications and funds six over the officer who takes eleven and funds eight. Processing absorbs the difference as rework. The officer who is genuinely better at the job scores lower on the only visible line and eventually stops caring about the parts that are not counted.
Good: six to nine lines, weighted, published, reviewed monthly. A typical build looks like this — funded volume carrying the heaviest weight, pull-through rate next, then cycle time, document accuracy at first submission, referral-partner activity, and compliance or audit findings. Each line gets a weight reflecting how much it matters to your business this quarter. Each officer gets a 1-to-5 level on each line. The composite is the sum of weight times level across all lines. An officer who is a 5 on volume and a 1 on accuracy does not land at the top of the board, and the specific line dragging them down is visible to them without a meeting.
Bad: levels defined by vibes. "You're about a 3 on documentation" is not coachable feedback. It is an opinion the officer can dispute and the manager cannot defend.

Good: levels defined by observable thresholds. Document accuracy level 5 means the file clears initial review without a conditions request more often than not; level 3 means it usually takes one round; level 1 means it routinely takes two or more. Cycle-time levels anchor to day bands relative to your team median, not to an absolute number pulled from a trade article. The thresholds should be written down in one paragraph per line, and they should be boring enough that two different managers scoring the same officer land within one level of each other. If they cannot, the definition is too vague — tighten it.
Bad: the matrix lives in the manager's spreadsheet. If officers see their score only at the quarterly review, the scorecard is a performance-management artifact, not a behavior-change tool. Feedback delivered ninety days after the behavior changes nothing.
Good: published where officers can see their own lines and the team distribution. Not necessarily every officer's individual score — some cultures handle full transparency well, some do not — but at minimum each officer sees their own six lines, their composite, and where the team sits. The moment an officer can see that their pull-through line is the one costing them, they fix it without being told, because the path from behavior to score to paycheck is now legible.
Bad: weights set once and never revisited. The weights encode strategy. Strategy changes.

Good: weights reviewed on a fixed cadence and changed deliberately when conditions shift. Quarterly is a reasonable default. When rates move sharply, re-weight immediately and announce it — that is the whole advantage of a weighted system over a hard-coded comp plan. You can shift emphasis toward purchase-channel referral activity in one afternoon and have the floor re-aimed the next morning, without renegotiating anyone's commission agreement.
The diagram makes the structural point: fallout and rework are not neutral events that happen to a file. They are outcomes an officer influences at submission, and a scorecard that only counts applications is blind to every path on the left side of that flow.
One more distinction worth drawing. A good system separates lines the officer controls from lines they do not. Underwriting turn times, investor guideline changes, and appraisal delays are real drivers of cycle time and are not the officer's fault. If you score raw cycle time without normalizing for those, officers will correctly perceive the metric as unfair and disengage from it. Either scope the cycle-time line to the segment the officer owns — application to complete submission — or normalize against the team median for the same period so systemic delays wash out.
Real cost and ROI ranges
The honest framing here is that the scorecard itself is nearly free and the implementation is where the cost sits.

The build. A first version of a weighted matrix is a few hours of work in a spreadsheet — list the lines, agree the weights with your branch manager, write the level definitions, and score the current roster. Getting the level definitions to a state where two managers agree is the part that takes real time, usually two or three working sessions rather than one. Budget a week of part-time effort for a single branch, longer if you are standardizing across multiple branches with different local practices.
The data. This is the actual constraint. Funded volume and applications taken are trivially available from any loan origination system. Pull-through requires cohorting — you need applications taken in a period matched against fundings from that same cohort, not fundings in the period, or you get a number that swings wildly with pipeline timing. Cycle time requires clean timestamps at consistent milestones, and many teams find their milestone data is inconsistently entered. Document accuracy usually does not exist as a field at all and has to be captured, either from your LOS conditions data if it is structured, or as a lightweight manual tally by processing. Referral-partner activity is often the messiest, living in a CRM that officers update selectively.
Plan for the possibility that two of your six lines cannot be measured cleanly in month one. Launch with the four you can measure honestly and add the others once the data is trustworthy. A scorecard with four defensible lines beats one with six lines where two are guesses, because the moment an officer catches a wrong number the whole instrument loses credibility.
Tooling. There is a real range here and no need to spend at the top of it early. A spreadsheet costs nothing but carries maintenance risk and goes stale the moment the person who built it gets busy. Sales-scorecard and gamification platforms — Ambition, Spinify, Hoopla and similar — automate the visibility layer off your CRM data, generally on per-user monthly pricing that scales with headcount, with the lighter gamification tools sitting well below the coaching-platform tier. Incentive-compensation platforms — QuotaPath at the accessible end, CaptivateIQ and Xactly at the enterprise end — are what you reach for when the composite needs to drive actual commission calculation across multi-component plans. Salesforce or a comparable CRM can host the whole scorecard in custom dashboards if you are already standardized on it, but you build the matrix yourself; it is infrastructure, not a product. Pricing on the enterprise comp platforms is quote-based and varies enough by team size and plan complexity that any number quoted secondhand is not worth relying on — get your own quote.

The sequencing that wastes the least money: build the matrix in a spreadsheet, run it manually for a quarter, and only then buy tooling for whichever layer is actually hurting — visibility or comp calculation. Teams that buy first almost always end up configuring a platform around a matrix they had not thought through, then rebuilding it.
Where the return comes from. Do not model this as a productivity gain. Model it as recovered fallout and recovered capacity, because that is where the mechanism actually operates.
Recovered fallout is the cleaner of the two. Take your current pull-through rate and your current application volume. Ask what a modest improvement in pull-through — the kind that comes from officers pre-qualifying more carefully and setting borrower expectations earlier — is worth in incremental funded loans at your average revenue per funded loan. The arithmetic is straightforward with your own numbers and it is usually the largest line in the case. Crucially, this improvement requires no additional applications, no additional marketing spend, and no additional headcount. It is revenue from files you already have.
Recovered capacity is the operations side. Every conditions request is processor time, underwriter time, and borrower patience. If document accuracy improves and rework rounds drop, your existing processing team absorbs more volume without adding a head. Ask your operations manager what fraction of their week goes to chasing incomplete submissions — the number is usually higher than leadership assumes, and it converts directly into deferred hiring.

There is a third return that is harder to quantify but shows up in exit interviews: middle performers stay longer when the standard is written down. Replacing a producing loan officer is expensive in recruiting cost, ramp time, and referral relationships that walk out the door with them. A published matrix does not fix a bad comp plan, but it removes "I never knew what they wanted from me" from the list of reasons people leave.
Where it does not pay. If your funded-volume shortfall is driven by product competitiveness, pricing, or a genuinely thin lead flow, a scorecard measures the problem accurately and does not solve it. Diagnose first. If your top three officers are all missing plan by similar margins, the constraint is probably not individual behavior — it is upstream, and re-scoring the same people will just annoy them.
How it plugs into your daily and monthly workflow
A scorecard that only appears at review time is administrative overhead. One that is wired into the operating rhythm becomes the way the team talks about the work. The difference is entirely in the cadence.

Daily, at the officer level. Officers should be able to see their own lines without asking anyone. In practice this means the scorecard view is a link, not a meeting. The daily use is not score-watching — it is knowing which of the six lines is the weak one this month and letting that shape the day. An officer whose pull-through line is soft spends more time on pre-qualification conversations and less on volume-chasing. An officer whose accuracy line is soft builds a document checklist before submitting. Neither of those changes requires a manager conversation once the score is visible.
Weekly, at the pipeline meeting. Replace the applications-taken readout with a pull-through and cycle-time readout on the cohort. This is a small change that shifts the room's attention meaningfully. The question moves from "how many did you take" to "what is stuck and why." Files that have sat at the same milestone for more than a week get named. Keep this short — the weekly meeting is for unblocking, not for scoring.
Monthly, at the scoring cycle. Levels get updated. This is the one moment the matrix formally changes, and it should be predictable — same week every month, same process. Managers score, officers see their updated composite, and each officer gets one specific line identified as the focus for the coming month. One line, not three. A coaching conversation that names a single behavior with a defined threshold produces change; a conversation that names six does not.
Quarterly, at the weights review. Leadership revisits the weights against strategy. Is purchase business the priority, or portfolio retention? Is the constraint operations capacity or lead flow? The weights should visibly answer that, and when they change, the change is announced with the reasoning. Officers tolerate re-weighting well when the logic is explained and poorly when it appears to be moving the goalposts.

Who owns it. In a larger shop this sits with RevOps or sales operations — someone owns the data pipeline, the definitions, and the integrity of the numbers, separate from the manager who does the scoring. That separation matters. When the person who scores also owns the data, disputes have nowhere to go. In a smaller branch, the branch manager owns both, and the substitute for separation is publishing the raw inputs alongside the score so officers can check the arithmetic themselves.
Adjacent applications of the same structure. The mechanism generalizes further than most lending teams expect. Insurance agencies run the identical problem with producers — quotes issued versus policies bound, with retention as the pull-through analogue. Auto dealerships face it with credit applications versus delivered units. Any business where a front-line seller initiates something that a back office has to complete has the same measurement gap, and the same fix: score the chain, weight the lines, publish the levels. If you have already built the matrix for loan officers, the processing team and the referral-partner-facing roles can be scored on the same instrument with different lines and different weights.
Connecting it to pay without breaking anything. The safest sequencing is to run the scorecard for a full quarter with no compensation attached. Let officers see their scores, dispute the numbers, and force you to fix the definitions that are wrong — and some will be wrong. Only after the instrument has survived a quarter of scrutiny should the composite start driving variable pay. Wiring pay to an untested matrix guarantees that the first bad number becomes a compensation dispute, and that is the fastest way to kill the whole effort.
When you do connect it, connect a portion, not the whole thing. Keep funded volume driving the base commission structure that officers already understand, and let the composite drive a bonus layer on top. That gives the matrix real teeth without asking anyone to accept a wholesale rewrite of how they get paid.
Related questions
How is pull-through rate actually calculated?
Take applications from a defined cohort period and divide the number that eventually funded by the total taken. Measure by cohort, not by calendar period — funded loans this month came from applications taken over several prior months, so period-over-period comparison distorts the number badly.
Should individual scores be visible to the whole team?
Each officer should always see their own lines and the team distribution. Full name-by-name transparency works on floors that already run competitively and backfires on teams that do not. Start with self-view plus anonymized distribution; open it up only if the culture supports it.
How many lines should the matrix have?
Six to nine. Below five you are back to a single-metric scoreboard with extra steps. Above nine, officers cannot hold the priorities in their head and the weights become too diluted for any single line to drive behavior.
What if an officer disputes their level?
Good — that is the instrument working. Show the underlying data. If they are right, fix the number and thank them publicly. If the definition was too vague to settle the dispute, rewrite the definition. Disputes surface bad definitions faster than any review process.
Does this work for a two-person branch?
Yes, though the composite matters less than the level definitions. With two officers there is no meaningful distribution to compare against. The value is in writing down what good looks like, which makes coaching concrete and makes the next hire's onboarding faster.
FAQ
Why not just pay more commission on funded loans and skip the scorecard?
Commission on funded loans is a lagging incentive — it rewards the outcome months after the behaviors that produced it. The scorecard makes the intermediate behaviors visible while they are still changeable. Most teams need both: commission for the outcome, and a scored matrix for the leading indicators that get you there.
How long before this shows up in funded volume?
Expect a lag matching your average cycle time plus a month or two of behavior change. If your application-to-funding cycle runs several weeks, the first cohort scored under the new system does not finish funding for a while. Judge the first quarter on whether pull-through and accuracy are moving, not on funded volume.
What if my LOS cannot produce clean cycle-time data?
Launch without it. Score the lines you can measure honestly — funded volume, applications, pull-through, and a manually tallied accuracy line from processing — and add cycle time once milestone timestamps are reliable. A four-line matrix everyone trusts is worth more than a six-line matrix with two unreliable numbers.
Do the weights need to be the same for every officer?
The weights should be the same across the team for fairness; the levels are what differ per person. The exception worth making is a distinct ramp matrix for new officers, where activity and training lines carry more weight and funded volume carries less for the first several months.
Will this demotivate top producers who fund a lot but run messy files?
Some friction is the point. A high-volume officer with poor accuracy is generating real cost in operations that nobody has been charging to them. That said, weight funded volume heavily enough that a genuine top producer still lands near the top of the composite — the matrix should nudge them on accuracy, not punish them for producing.
How do referral partners fit into the scoring?
Usually as an activity line rather than a results line, because partner-sourced volume already shows up in funded volume and double-counting it distorts the weights. Score the input — meetings held, new partners activated, partners who sent at least one file this quarter — and let the output land where it already lands.
Sources
- Consumer Financial Protection Bureau — mortgage origination rules and loan originator compensation requirements: https://www.consumerfinance.gov/rules-policy/regulations/1026/36/
- Mortgage Bankers Association — industry performance and origination data: https://www.mba.org/news-and-research
- Federal Reserve — Senior Loan Officer Opinion Survey on Bank Lending Practices: https://www.federalreserve.gov/data/sloos.htm
- Fannie Mae — Selling Guide, origination and underwriting requirements: https://selling-guide.fanniemae.com/
- Freddie Mac — Seller/Servicer Guide: https://guide.freddiemac.com/
- U.S. Small Business Administration — 7(a) loan program requirements: https://www.sba.gov/funding-programs/loans/7a-loans
- Harvard Business Review — research on sales compensation and incentive design: https://hbr.org/topic/subject/sales
- Federal Financial Institutions Examination Council — Home Mortgage Disclosure Act data: https://www.ffiec.gov/hmda/
- Nationwide Multistate Licensing System — mortgage loan originator licensing and requirements: https://mortgage.nationwidelicensingsystem.org/
Related on PULSE
- [How Many Producers Do I Need to Hire for My Insurance Agency to Grow My Book?](/knowledge/tl0015)
- [How Many Salespeople Do I Need to Hire for My Car Dealership?](/knowledge/tl0052)
- [How Many Sales Reps Do I Need to Hire for My Logistics Company?](/knowledge/tl0058)
- [How Do I Figure Out How Many People to Schedule Each Day and at What Times for My Single Store?](/knowledge/tl0002)
- [How Many Attendants Should I Schedule Each Day at My Car Wash?](/knowledge/tl0067)










