How do you use Gong to coach a rep on discovery questions vs. pitch delivery in 2027?
PULSEKNOWLEDGE LIBRARY
Split Gong coaching into two scorecards: one for discovery (question count, open-ended ratio, longest-monologue, talk ratio) and one for pitch delivery (demo pacing, filler words, objection handling). Use trackers and call scores to find the specific 90-second clip, coach that clip, then re-measure the same metric two weeks later.
The scenario that forces the split
A mid-market SaaS team runs a Monday pipeline review. Two reps, same quota, same territory, wildly different outcomes. Rep A books 14 first meetings a month and converts 3 to opportunity. Rep B books 9 and converts 6. The sales manager's instinct is to coach both on "asking better questions," because that is the phrase every enablement deck uses. That instinct is wrong for at least one of them, and Gong's recorded call data is the only thing in the stack that can tell you which.
Pull the last ten calls for each rep. For Rep A, the discovery calls run 18 minutes with a talk ratio around 62% rep-side, eleven questions asked, and only two of them open-ended. The rest are confirmatory — "does that make sense," "so you're using Salesforce, right." Rep A is not doing discovery; Rep A is doing a survey with a demo bolted on the end. For Rep B, discovery calls run 34 minutes, talk ratio near 41%, twenty-two questions, nine open-ended, but the demo calls fall apart: the longest customer monologue on a demo call is 12 seconds, meaning Rep B never stops talking long enough for the buyer to react, and the deal stalls after the demo.
These are two completely different coaching problems wearing the same jersey. Rep A has a discovery deficiency. Rep B has a delivery deficiency. If you run one generic "ask better questions" session, you make Rep A marginally better and you actively waste Rep B's time — worse, you signal to Rep B that their strongest skill is the thing that needs work, which is how you get a good rep to start second-guessing the behavior that was producing conversion.
The whole point of using Gong for coaching in 2027 is that you no longer have to guess which bucket a rep is in. Every call is transcribed, every question is counted, every stretch of monologue is timestamped. The manager's job shifts from *diagnosing from memory* to *diagnosing from the call library*, and then from diagnosis to picking the one 90-second clip that proves the point. Reps do not change behavior because a manager described a pattern in a 1:1. They change because they heard themselves do the thing.

There is a second reason to formally separate the two skills. They fail differently and they recover differently. Discovery is a *volume and structure* problem — a rep asking too few questions can be fixed with a question bank and a call plan inside two weeks. Delivery is a *timing and restraint* problem — teaching a rep to shut up after a feature statement and let silence do work takes six to ten weeks and a lot of reps hearing their own filler. Mixing the two into one coaching cadence means you either abandon the delivery work too early or you drag the discovery fix out far longer than it needs.
Where RevOps enters: the split only holds if the two skills are measured on the *right call type*. Discovery metrics computed on a demo call are noise. Delivery metrics computed on a first call are noise. Somebody has to define the call-type taxonomy in Gong, make sure calls get tagged to the right stage, and keep that mapping honest as the sales process changes. That is not a manager task; that is a RevOps task, and it is the difference between a scorecard people trust and a scorecard people quietly ignore.
How the mechanism actually works
Gong's coaching value comes from four layers stacked on top of the raw recording, and you need to understand what each one actually measures before you build a coaching program on it.
Layer one — the transcript and speaker separation. Every call is transcribed and diarized, so the platform knows which words belong to the rep and which belong to the buyer. This is the substrate. Everything else is arithmetic on top of it. Transcription accuracy degrades with heavy accents, bad microphones, and multi-party rooms where two people share one line, so if a rep's numbers look bizarre, check whether the diarization split them correctly before you coach them on it.

Layer two — the interaction statistics. Talk ratio, longest monologue (rep side and customer side), patience (how long the rep waits after the buyer stops talking before speaking), question rate per hour, and interactivity/switches per minute. These are computed automatically on every call. They are the *discovery* dashboard, mostly, though longest-monologue is the single best delivery metric in the entire product.
Layer three — trackers. A tracker is a keyword or phrase rule that flags when specific language appears on a call. This is where you encode your own methodology. If you run MEDDICC, you build trackers for economic buyer language, decision criteria language, and paper process language. If you have a competitor problem, you build a tracker per competitor. Trackers are the mechanism that turns "did the rep do discovery" from a vibe into a checkbox, because you can ask: on how many of this rep's first calls did the *metrics* tracker fire?
Layer four — scorecards and call reviews. A scorecard is a manual rubric a manager fills out against a specific call. This is where human judgment lives, and it is where the discovery/delivery split becomes structural: you build two separate scorecards, not one blended one.

The mechanism only produces behavior change when those four layers feed a loop. Metrics surface a candidate problem, trackers confirm it is a methodology gap rather than a call-type artifact, the manager scores one call and clips the 90 seconds that demonstrates it, the rep watches the clip and commits to one change, and the same metric gets re-pulled two weeks later. Skip the clip and you get a conversation nobody remembers. Skip the re-measure and you never learn whether the coaching worked.
The two scorecards should not share line items. A discovery scorecard asks: did the rep establish current-state before proposing anything, did they quantify the pain in a number the buyer said out loud, did they identify who signs, did they ask at least one question the buyer visibly had to think about. A delivery scorecard asks: did the rep tie each feature to a pain the buyer already named, did they pause after the value statement, did they handle the pricing objection without discounting reflexively, did they land a specific next step with a date.
One more mechanical detail that matters more than people expect: coach on the rep's own calls, not on library exemplars, for the first three sessions. Exemplar libraries are useful for onboarding, but a rep watching a top performer learns "that person is different from me." A rep watching themselves learns "I do a thing I did not know I did." Once the rep has internalized their own baseline, exemplar calls become useful as a target.
Real numbers, ranges, and benchmarks
Treat every number below as a starting range to calibrate against your own data, not a law. The correct benchmark is always your own top quartile's actual behavior, computed from your own call library — segment, deal size, and sales motion move these numbers substantially.

Discovery-call talk ratio. The widely cited healthy band for a discovery call is roughly 40–50% rep talk time. Under 30% usually means the rep is passive and letting the buyer wander. Over 65% on a *first* call is the clearest single red flag in the entire dataset — the rep is pitching before they know what they are pitching against. Rep A in the scenario above sat at 62%; that is the number you coach.
Question count. On a 30-to-45-minute discovery call, healthy reps typically land somewhere in the 15–25 question range. Below 10 is a survey. Above 35 starts to feel like an interrogation and buyers report it as such. The more important number is the *open-ended ratio* — what fraction of questions cannot be answered with yes/no/a single noun. Aim for at least half. Rep A's 2-of-11 is the coaching target, not the raw 11.
Longest customer monologue. This is the sleeper metric. On a good discovery call you want at least one stretch where the buyer talks uninterrupted for 60+ seconds, ideally 90. That is the moment they stopped answering questions and started explaining their actual problem. If a rep's longest customer monologue across ten calls never exceeds 30 seconds, they are interrupting or over-steering — and that is a delivery habit showing up inside a discovery call.
Longest rep monologue on a demo. Keep it under about 2 minutes and 30 seconds. Past roughly three minutes of uninterrupted rep talk, buyer attention measurably drops and the odds of the buyer surfacing an objection go down — which sounds good and is actually terrible, because unsurfaced objections resurface after the call when you are not in the room.

Patience / switch rate. Reps who wait about a second after the buyer stops before responding get materially more elaboration than reps who jump in at 0.2 seconds. Interactivity — speaker switches per minute — in the range of roughly 8–12 on a discovery call generally indicates real conversation rather than alternating monologues.
Coaching volume, the operational number. A manager with eight reps should be reviewing roughly 2–3 calls per rep per month, scored, with at least one clip. That is 16–24 scored calls a month, at 20–30 minutes each including the 1:1 discussion — call it 8–12 hours of manager time monthly. Managers who commit to "I'll review everything" review nothing by week three. Cap it and protect it on the calendar.
Time to measurable change. Expect a discovery metric — question count, open-ended ratio — to move within 2–3 weeks of focused coaching, because the rep can consciously execute a question list. Expect delivery metrics — monologue length, filler words, pause discipline — to take 6–10 weeks, because they are habits, not checklists. Managers who apply the discovery timeline to a delivery problem conclude the rep is uncoachable when the rep is simply mid-curve.
Filler words. Some teams track "um/uh/like/you know" density. Useful as a delivery signal, dangerous as a scorecard line item, because it is highly visible and easy to game while producing near-zero revenue impact. If you track it, track it as context, not as a scored item.

Sample size. Do not coach off one call. Pull at least five calls of the same type before declaring a pattern. Single calls are dominated by who the buyer was that day. Five calls of the same type from the same rep with the same metric out of band is a pattern.
Trade-offs and alternatives
Every design decision in a Gong coaching program has a cost, and pretending otherwise is how these programs get quietly abandoned in month four.
Automated scores vs. manual scorecards. Automated metrics are free, consistent, and available on 100% of calls, but they measure form rather than substance. A rep can hit a perfect 45% talk ratio while asking uniformly useless questions. Manual scorecards capture substance but cost manager time and drift between managers — two managers scoring the same call routinely land two points apart on a five-point rubric. The working compromise: use automated metrics for *triage* (which calls deserve a human look) and manual scorecards for *judgment* (what actually went wrong). Calibrate manager scoring quarterly by having every manager score the same call and comparing.
Coaching discovery first vs. delivery first. If a rep is weak at both, coach discovery first. Discovery gaps poison everything downstream — a great pitch aimed at the wrong pain still loses, and you cannot even evaluate the pitch fairly because it was built on bad inputs. The exception is a rep whose discovery is adequate and whose demos are actively losing deals late-stage; there, delivery is the bottleneck and fixing discovery further returns nothing.

Scorecard breadth. A 15-line scorecard captures more nuance and gets filled out honestly for about three weeks before managers start pattern-matching to save time. A 5-line scorecard loses nuance but survives. Choose five to seven items per scorecard, all binary or three-point, and rotate the items quarterly rather than accumulating them.
Tracker precision vs. recall. Broad trackers ("budget", "price") fire constantly and become noise. Narrow trackers ("what's the approval process for this budget") miss most real instances. Build trackers narrow enough to be meaningful, then audit each one monthly by sampling ten fired calls and checking whether the fire was a true positive. A tracker with worse than roughly 70% precision should be rewritten or retired.
Peer review vs. manager review. Peer review scales — reps score each other, manager time drops sharply, and reps learn more from scoring than from being scored. But peers are conflict-averse and inflate. Use peer review for volume and manager review for the calls that matter, and never let peer scores feed a performance rating.
Buying the platform vs. building the discipline. The uncomfortable trade-off: the platform is the cheap part. A team that has Gong and does not run a scored coaching cadence gets call search and nothing else. A team with disciplined weekly call reviews on plain recordings will out-coach a team with the full platform and no cadence. Budget the manager hours before you budget the seats.

Cadence trade-off. Weekly coaching produces faster change but burns manager capacity and can feel like surveillance. Monthly is sustainable but too slow for a rep who needs to ramp. Biweekly, with one scored call and one clip each session, is the setting most teams land on after trying both extremes.
Common pitfalls and how to avoid them
Coaching the metric instead of the behavior. Telling a rep "get your talk ratio to 45%" produces a rep who pauses artificially and asks filler questions to game the number. The metric is a symptom. Coach the behavior — "after they describe the problem, ask one follow-up before you move on" — and let the ratio follow.
Comparing across call types. A rep's demo calls will always show higher talk ratio than their discovery calls, because a demo is supposed to be rep-led. If your dashboard blends them, every rep who does a lot of demos looks like a talker. Filter by call type or the data lies to you. This is the single most common analytics error in Gong coaching programs and it is a RevOps fix, not a manager fix.

Coaching from a single bad call. Managers remember the disaster call and coach it for a month. Pull five calls. If the pattern is not in four of them, it was a bad Tuesday, not a skill gap.
Using recordings punitively. The fastest way to kill a coaching program is to surface a clip in a performance conversation the rep did not see coming. Reps will start scheduling calls off-platform, taking discovery to the buyer's cell phone, and your data quality collapses. State the rule explicitly: call review is for coaching; performance conversations use pipeline and quota outcomes.
No clip, no change. A 1:1 where the manager describes the pattern verbally produces agreement and no behavior change. Play the 90 seconds. Let it be uncomfortable for eight seconds. Then ask the rep what they would do differently — the rep's own answer sticks, the manager's does not.
Coaching more than one thing. Every session, one change. A rep given four things to fix fixes zero. Write the one change down, and open the next session by re-measuring exactly that.

Letting trackers rot. Trackers are built once, during rollout, and then the product changes, competitors change, and the trackers keep firing on language nobody uses. Audit quarterly. A tracker dashboard nobody trusts is worse than no dashboard, because it produces confident wrong conclusions.
Skipping the re-measure. This is the pitfall that invalidates the entire program. If you never pull the same metric again in two weeks, you have no idea whether your coaching works, which manager is effective, or whether the rep is improving or plateauing. Put the re-measure date in the same note as the coaching commitment.
Ignoring the buyer-side signal. Most teams stare at rep metrics exclusively. The buyer's longest monologue, the buyer's question count, and how early in the call the buyer first asks a question are often more diagnostic than anything the rep did — a buyer who asks nothing for 20 minutes is a buyer who is not engaged, and that is a discovery failure regardless of how clean the rep's numbers look.
Assuming coverage equals coaching. "We reviewed 200 calls this quarter" is a vanity number. Twelve scored calls with clips and re-measures beat two hundred skimmed ones. Report on coached-and-re-measured, not on reviewed.
Related questions
How many calls should a manager review per rep each month?
Two to three of the same call type, scored with a rubric and at least one clip each. That is roughly 8–12 hours monthly for an eight-rep team. Reviewing more calls without scoring them produces coverage statistics, not behavior change.
Should discovery and pitch use the same scorecard?
No. Build two five-to-seven line scorecards with no shared items. Discovery scores question quality, pain quantification, and buyer talk time. Delivery scores feature-to-pain tie-back, pause discipline, objection handling, and next-step specificity. Blending them hides which skill is actually broken.
What talk ratio should a discovery call hit?
Roughly 40–50% rep-side is the common healthy band, but calibrate against your own top quartile. Above 65% on a first call means the rep is pitching before diagnosing. Below 30% often means the rep is passive rather than genuinely listening.
How long before coaching shows up in the metrics?
Discovery metrics — question count, open-ended ratio — typically move in 2–3 weeks because they are executable checklists. Delivery habits like monologue length and pause discipline take 6–10 weeks. Judging delivery on a discovery timeline makes coachable reps look uncoachable.
Can call recordings be used in performance reviews?
Technically yes, operationally no. The moment reps believe recordings feed performance decisions, they route important conversations off-platform and your data degrades. Keep coaching and performance evaluation formally separate and say so out loud.
FAQ
What is the single most useful Gong metric for coaching discovery?
Open-ended question ratio, with longest customer monologue as the close second. Raw question count is easy to game — a rep can fire off fifteen yes/no confirmations. The ratio of questions that require the buyer to actually think, combined with whether the buyer ever spoke uninterrupted for 60+ seconds, tells you whether real discovery happened.
How do I coach a rep whose numbers look fine but who still loses deals?
Stop looking at form metrics and read the transcripts. Good ratios with bad outcomes almost always means the rep is asking structurally correct questions about the wrong things — process questions instead of consequence questions, or discovering pain the buyer has already decided to live with. Score three calls manually on substance and the gap usually surfaces immediately.
Do trackers replace manual call review?
No. Trackers tell you whether specific language appeared; they cannot tell you whether it landed. A tracker firing on "who else is involved in this decision" proves the rep said the words, not that they got a useful answer or followed up. Use trackers to triage which calls need human review, then review them.
How do I get reps to actually watch their own clips?
Send one clip, under 90 seconds, with a single specific question attached — "what would you ask right here instead?" — and require a written one-line answer before the 1:1. Long clips and open-ended "thoughts?" prompts get ignored. The written answer is what forces the watch.
What should RevOps own in this program versus the sales manager?
RevOps owns the plumbing: call-type taxonomy, stage mapping, tracker library and its quarterly precision audit, scorecard configuration, and the reporting that separates discovery calls from demo calls. The manager owns diagnosis, clip selection, the 1:1, and the re-measure. When RevOps starts picking clips or managers start editing trackers, both jobs degrade.
Is it worth coaching filler words and speaking pace?
Marginally, and only after the structural work is done. Filler-word density is highly visible and easy to measure, which makes it seductive, but reducing "um" density rarely moves conversion. A rep who tightens their longest monologue from four minutes to two will out-earn a rep who eliminated every filler word and still never stopped talking.
Sources
- https://www.gong.io/resources/
- https://www.gong.io/blog/
- https://help.gong.io/
- https://hbr.org/2018/09/how-to-really-listen-to-your-employees
- https://www.salesforce.com/resources/articles/sales-coaching/
- https://www.linkedin.com/business/sales/blog
- https://www.forrester.com/blogs/category/sales-enablement/
- https://www.gartner.com/en/sales
Related on PULSE
- How do you build a MEDDICC scorecard that managers actually fill out?
- What talk-ratio benchmarks should you use for discovery vs. demo calls?
- How do you structure a weekly sales coaching cadence without burning manager capacity?
- How do you set up conversation-intelligence trackers that don't rot after a quarter?
- What should RevOps own in a sales enablement program vs. the frontline manager?









