The Gong Call-Review Clinic: Scoring Discovery in 60 Minutes
PULSEKNOWLEDGE LIBRARYQuality
Certified

The Gong Call-Review Clinic is a 60-minute, structured session where a sales team scores one real discovery call against a fixed rubric — typically MEDDPICC — live, as a group. The expected outcome is a shared standard for what "good discovery" sounds like, a numeric baseline score, and a repeatable weekly cadence that ties call quality to pipeline health rather than gut feel.
The outcome you should expect
Run this clinic once and you get a single data point: a scored recording and a room full of reps who now share the same vocabulary for what a strong discovery call sounds like. Run it weekly for a month and the outcome changes shape entirely — you get a trend line. Teams that commit to the cadence typically see their average MEDDPICC score move from the 7-9 range (out of 16) up into the 11-13 range within six to eight weeks, not because reps get smarter, but because the scoring exercise makes the gaps visible and repeatable in a way that a manager's private call-shadowing never does.
The first clinic almost always produces a lower score than anyone expects. A rep who "feels" like they run a tight discovery call will frequently score a 1 (vague) instead of a 2 (specific) on Metrics, Decision Process, or Champion, simply because the question was asked in a soft, open-ended way rather than pinned to a number or a name. That gap between self-perception and rubric reality is the single most valuable output of the first session — it reframes discovery from "did I ask enough questions" to "did I extract a number, a name, and a date."

The second-order outcome is organizational: a shared scoring language starts showing up in pipeline reviews. Instead of a manager asking "how's this deal feeling?", the question becomes "what's the Decision Process score?" That shift compresses forecast conversations because everyone is referencing the same 0-1-2 scale instead of adjectives. Teams that stick with the cadence for a full quarter report that forecast accuracy conversations get shorter, not because deals close faster, but because the qualification gaps get flagged three or four weeks earlier than they would have surfaced in a normal pipeline review.
What drives that outcome (mermaid)
Three mechanics drive whether this clinic actually changes behavior or just becomes a one-off training exercise that fades within a month.

Mechanic one: a fixed, shared rubric. The clinic only works if every participant scores the same call against the same criteria before comparing notes. MEDDPICC (Metrics, Economic Buyer, Decision Criteria, Decision Process, Paper Process, Identify Pain, Champion, Competition) gives eight elements, each scored 0 (not covered), 1 (vague), or 2 (specific), for a possible 16. The rubric has to be displayed and used identically every single week — swapping frameworks (BANT one week, MEDDPICC the next) destroys the trend line that makes the clinic valuable.
Mechanic two: silent independent scoring before group debate. If reps score out loud or in sequence, anchoring bias collapses the exercise into groupthink — the first person to speak sets the score for the room. The clinic has to force everyone to write down a private score during the recording, then reveal simultaneously (show of hands, sticky notes, or a shared form) before any discussion starts. The debate that follows a disagreement is where the real coaching happens, because it forces someone to articulate why they heard "specific" pain when someone else heard "vague."

Mechanic three: a closed feedback loop back into the CRM. A score that lives only in a facilitator's notebook evaporates. The score needs a landing spot — a custom field on the Opportunity record, a tracked field in the call-recording tool's own scorecard feature, or at minimum a shared spreadsheet row with the date, the rep, and the total. Without that loop, there's no way to prove the clinic is working, and it quietly gets deprioritized the first time the team gets busy.
Benchmarks and realistic ranges
Calibrate expectations with concrete ranges rather than vague improvement language. A first-time clinic on an unscored, unprepared team typically lands a total score between 6 and 9 out of 16 — that's the realistic starting point, not a failure. Teams already using some form of structured discovery training tend to start closer to 9-11.

On individual elements, Metrics and Identify Pain are usually the strongest out of the gate — reps are trained to ask about problems and impact, so these commonly score a 2. Decision Process is almost always the weakest first-session element; it's common for a room to unanimously score it 0 because nobody asked who else is involved or what the timeline looks like beyond the current conversation. Paper Process (legal, procurement, contract terms) is the second most commonly missed element, frequently scoring 0 or 1 because reps treat it as a late-stage concern rather than something to flag in discovery.
For cadence, weekly sessions for the first four to six weeks produce the fastest score movement, since the gap between sessions is short enough that a rep can apply last week's feedback to this week's calls before the pattern fades. After that initial stretch, moving to a bi-weekly or twice-monthly cadence is enough to sustain the score rather than let it erode, provided the CRM logging keeps happening every week even when the live group session doesn't.

On time allocation within the 60 minutes, the split that holds up best across repeated sessions is roughly: 10 minutes to set the rubric and surface common mistakes, 15 minutes to listen and score independently, 15 minutes for group reveal and debate, 10 minutes to rewrite the weakest questions as stronger ones, 5 minutes to set the week's action items, and 5 minutes reserved for close and buffer. Sessions that compress the independent-scoring step below 10 minutes tend to produce shallower debate, because reps haven't had time to actually commit to a score before the group conversation starts.
On measurable business impact, the metric worth watching isn't the MEDDPICC score in isolation — it's the correlation between score and stage-to-stage conversion. A deal scored 12+ at the discovery stage should convert to the next stage at a meaningfully higher rate than one scored under 8; if that correlation isn't showing up after two or three months of consistent scoring and logging, the rubric weighting or the scoring discipline (reps inflating scores to look good) needs to be re-audited before adding more sessions.

Risks, edge cases, and failure modes
The most common failure mode is turning the clinic into a one-time training event instead of a recurring habit. A single 60-minute session generates enthusiasm and a checklist, but without the weekly repetition and the CRM logging step, the scoring language fades from team vocabulary within two to three weeks and pipeline reviews revert to adjectives instead of numbers.
A second failure mode is score inflation once the exercise becomes visible to management. If reps know their manager is reading the logged scores, there's a natural pull toward rounding a 1 up to a 2 to look competent. The fix is to keep the early sessions low-stakes and coaching-oriented — no scores tied to compensation or performance reviews for at least the first month — so reps are incentivized to score honestly rather than defensively.

A third risk is picking the wrong recording. A call that's too short (under 15-20 minutes), too early-stage, or one where the prospect was clearly unqualified from the start doesn't give reps enough material to actually apply all eight rubric elements, and the group ends up debating whether an element even applied rather than how well it was executed. Choose a call from a real, still-open opportunity with a reasonably engaged prospect — anonymized public examples work in a pinch but produce weaker debate because nobody has context on what happened next.
A fourth edge case is applying one rubric weighting universally across inbound and outbound motions. An inbound call where the prospect initiated contact usually has different natural strengths (they'll volunteer pain and metrics more readily) and different natural gaps (Decision Process and Paper Process, since they haven't been asked to think that far ahead) than an outbound call, where Identify Pain and Champion typically need more deliberate probing. Running the identical rubric weight against both without adjustment produces misleading trend comparisons between inbound-heavy and outbound-heavy reps.

Finally, a legal and compliance edge case: using any call-recording platform to score and share clips requires that call recording consent and internal data-handling policies are already in place before the first clinic. Skipping this check to move fast on the training risks a compliance problem that has nothing to do with sales enablement and everything to do with how the recording was captured and who it's being shared with.
A practical rollout plan (mermaid)
Start with a single pilot clinic before committing to a permanent weekly slot. Pick one team of four to eight reps, one facilitator (usually the frontline manager), and one real call from the current pipeline. Run the full 60-minute structure once, and treat the first session's score as a baseline rather than a judgment — the goal of session one is calibration, not improvement.

For week two through six, keep the cadence weekly and keep the rubric completely fixed. Resist the urge to add new elements or switch frameworks during this stretch; the value comes from comparing identical measurements week over week. Log every score against the Opportunity record or a shared tracking sheet immediately after the session, while the debate and rationale are still fresh, since a score without the "why" behind it is much less useful two weeks later.
Around week six or seven, hold a checkpoint: pull the trend line and check whether the team's average total score has moved meaningfully (a jump of 2-3 points out of 16 is a realistic six-week target). If it hasn't moved, the likely cause is either inconsistent facilitation (different people running the debate portion differently each week) or lack of follow-through on the rewritten questions from the remediation segment — reps agreeing a stronger question exists in the room but not actually using it on live calls afterward.

From week eight onward, shift cadence to bi-weekly for the core scoring session while keeping weekly CRM logging alive independently (a rep can self-score one call a week even without the full group clinic). Introduce a second rubric variant only after the first is fully embedded — for example, a qualification-focused clinic using BANT to complement the MEDDPICC discovery clinic, run on alternating weeks so reps don't juggle two frameworks simultaneously.
Related questions
How long should a discovery call be before it's worth scoring in the clinic?
A call under 15-20 minutes rarely has enough material to apply all eight MEDDPICC elements meaningfully. Pick calls in the 20-40 minute range with a genuinely engaged prospect for the strongest scoring debate.
Who should facilitate the weekly Review — the manager or a peer?
Either works, but the manager typically gives the clinic more weight in the first month. Rotating facilitation to senior reps after the habit is established increases peer buy-in and reduces the sense that scoring is punitive.
What if two reps score the same element completely differently?
That disagreement is the most valuable moment in the session — resolve it with a 60-second debate, then have the facilitator give a ruling and explain the reasoning so the whole room recalibrates together.
Should the score affect compensation or performance reviews?
Not in the first month. Tying scores to comp too early causes reps to inflate scores defensively rather than use the exercise honestly, which quietly breaks the entire feedback loop.
Can this clinic work for a team of two or three reps instead of a larger group?
Yes, though the debate portion is less rich with fewer perspectives. Smaller teams should extend the independent-scoring window slightly and lean more heavily on the facilitator to challenge scores rather than relying on peer disagreement.
FAQ
How often should we run this clinic? Weekly for the first four to six weeks builds the shared vocabulary and habit fastest. After that, bi-weekly is usually enough to sustain the score, provided individual scoring and CRM logging continue every week even without the full group session.
What if our team doesn't use Gong specifically? The rubric and format are tool-agnostic. Any recording and transcript source works, and manual scoring from a written transcript is a completely valid substitute if a recording tool isn't in place yet.
Does this work the same way for outbound and inbound calls? The same rubric applies, but expect different natural strengths and gaps. Inbound calls tend to score higher on Identify Pain and Metrics since the prospect self-selected; outbound calls tend to need more deliberate probing on Champion and Decision Process.
What's the single biggest mistake reps make during discovery? Talking too much. A rep who dominates the conversation is pitching, not diagnosing, and it shows up directly as low scores on Decision Process and Paper Process, since those require the prospect to talk through their own internal steps.
How do we prove the clinic is actually working? Track whether deals scored higher at the discovery stage convert to the next stage at a meaningfully better rate than lower-scored deals. If that correlation isn't visible after two to three months, revisit scoring discipline before adding more sessions.
Can this format be adapted for qualification instead of discovery? Yes — swap the rubric to BANT or a qualification-specific scorecard and keep the same 60-minute structure. Many teams alternate a discovery-focused clinic with a qualification-focused one on alternating weeks once the format is established.
Sources
- Gong.io: The Science of Discovery Calls
- MEDDPICC Framework by Winning by Design
- Challenger Sale: Discovery Questions
- Salesforce: Custom Fields on Opportunities
- Clari: Forecasting and Pipeline Data
- Gartner: Sales Insights and Best Practices
- Outreach: Call Recording and Coaching
- Forrester: B2B Sales Research
Related on PULSE
- [The Customer Health Scoring Reboot — 60-Min Training](/knowledge/st208)
- [The Discovery Question Calibration Clinic — 60-Min Training](/knowledge/st0041)
- [Mastering the Discovery Call: A 30-Minute Ready-to-Run Sales Meeting Template](/knowledge/st0776)
- [Ready-to-Run Sales Discovery Workshop Template](/knowledge/st0750)
- [Discovery Call Script A/B Testing: Compare and Contrast Session](/knowledge/st0742)
This page will be disappearing soon. Save it to your device for $1 — or read it free while it is here.
@Kory-White- · if Venmo asks, the last 4 of my number are 2012









