Pulse - Value Added
Rent this Advertising Space
FRACTIONAL CRO · MARYLAND-BASED, NATIONWIDE · $0→$200M

Kory White

RevOps & Revenue Leadership

Get a 30-minute revenue checkup — Kory reviews your pipeline and forecast, then names the 1–2 fixes that move revenue fastest. 25 yrs scaling teams $0→$200M.

30-minute revenue checkup →
Hire a Fractional CROHow We Help?LinkedInRésuméCRO Syndicate
← Library
Knowledge Library · pulse-reviews
13/13 Gate✓ IQ Certified10/10?

How do you detect AE sandbagging in your 2027 forecast?

Curated by · Fractional CRO · Maryland
PULSEKNOWLEDGE LIBRARY
pulserevops.com
KnowledgeHow do you detect AE sandbagging in your 2027 forecast?
📖 3,993 words🗓️ Published Aug 22, 2026
Direct Answer

Detect AE sandbagging by comparing each rep's called probability against model-scored probability and actual outcomes across four to eight trailing quarters, then layering commit-to-pipeline ratio, late-quarter commit growth, and attainment-versus-commit history. One signal means conservatism; two or more repeating across quarters means a pattern your forecast process should flag for coaching, not punishment.

The quarter that closed at 118% and still broke the plan

Picture a $40M ARR company in the second week of Q1 2027. The CRO has a $9.6M number. The rolled-up commit from six AE teams is $7.1M, best case is $8.9M, and the board deck has already gone out saying the quarter is "tracking." Twelve weeks later the quarter closes at $11.3M — 118% of commit, 117% of plan. Everyone celebrates.

That quarter was a failure of forecasting, not a success of selling. The company hired against $9.6M. It sized SE capacity, implementation staffing, and cash-collection assumptions against a number that was wrong by nearly a quarter of itself. Two enterprise deals that closed in the final nine days had no assigned solutions engineer, so onboarding slipped six weeks and the first renewal cohort came in soft. Finance had held back a marketing spend release because the commit looked thin. The over-performance cost real money — it just didn't show up on a variance report that only flags misses.

Trace it back and the pattern is legible. Three of the eighteen AEs contributed 61% of the beat. Each of the three had the same fingerprint: a pipeline three to four times larger than their commit, a commit line that stayed nearly flat from week two through week nine, and a burst of "found" revenue in the last fourteen days. All three had beaten their commit in seven of the last eight quarters, never by less than 12%. None of them had ever missed. That last fact is the tell — genuine forecasting noise is symmetric. A rep who is honestly uncertain misses roughly as often as they beat. A rep who never misses is not calibrated; they are holding inventory.

This is what makes sandbagging structurally different from a bad forecast. A rep who is optimistic and wrong is visible immediately: they miss, everyone notices, the manager coaches. A rep who is pessimistic and wrong gets applauded. The feedback loop that corrects optimism does not exist for pessimism, which is why sandbagging compounds quietly for years while over-calling burns itself out in two quarters. The detection job is to build the missing feedback loop deliberately, because the org will not produce it on its own.

How do you detect AE sandbagging in your 2027 forecast — figure 1

The second thing the scenario reveals: sandbagging is rarely a character problem. Two of those three AEs had joined in the previous eighteen months, both from companies where a missed commit meant a territory cut. They learned a survival behavior and brought it with them. The third had a comp plan with a hard accelerator step at 100% attainment — every dollar he pushed from a strong quarter into a weak one was worth measurably more to him. He was responding rationally to the plan RevOps and finance designed. Detection that treats all three as dishonest reps will fix none of them.

How the detection mechanism actually works

The mechanism is a comparison engine, not a rule. You are looking for the persistent gap between three independently-produced numbers for the same deal: what the AE said, what a model or an objective stage-history baseline said, and what actually happened. Any single-quarter gap is noise. The signal lives in the sign of the gap holding constant over time.

Start with the raw material. For every closed opportunity in the trailing four to eight quarters you need: the AE's called category or probability as of a fixed snapshot point (week 2 and week 8 of the quarter are the two most useful), the model or baseline probability at the same snapshot, the actual close outcome, close date, amount, stage-entry timestamps, and the age of the deal at the moment it first entered commit. If your CRM does not preserve historical snapshots, this whole exercise is impossible — fixing snapshot capture is step zero and typically takes a RevOps team two to three weeks with a scheduled field-history job or a forecasting tool's built-in time-series store.

How do you detect AE sandbagging in your 2027 forecast — figure 2

With that data, four independent signals fall out.

Signal one — the calibration delta. Average the AE's called probability minus the baseline probability across all deals in the window. A well-calibrated rep lands within a few points of zero. A rep who is persistently 12–20 points below the baseline, and whose deals nonetheless close at baseline rates, is under-calling systematically. Compute it separately for deals that closed won and closed lost; a sandbagger's under-call concentrates on the won side.

Signal two — commit-to-pipeline ratio. Divide committed dollars by qualified pipeline at a fixed week. Reps carrying a healthy ratio in the mid-20s to mid-30s who suddenly show a ratio in the teens are either carrying junk pipeline or hiding real deals. The disambiguating question is what happened to that pipeline: if the non-committed tier converts at the rep's normal historical rate, those deals were real and should have been called.

Signal three — late-quarter commit accretion. Measure the share of final commit that entered commit in the last fourteen days. Some late entry is legitimate — deals genuinely accelerate. But cross-reference deal age. A deal that has been open 90 days and enters commit on day 78 of the quarter was known; the rep chose when to surface it. A deal created 22 days ago that enters commit late is a real find.

How do you detect AE sandbagging in your 2027 forecast — figure 3

Signal four — attainment-versus-commit distribution. Plot the rep's attainment against their own commit for eight quarters. You are reading the shape, not the average: tight clustering slightly above 100% with zero misses is the sandbagging signature.

The ordering in that last branch matters more than anything else in the diagram. Before you escalate a rep, you check whether the plan or the manager is producing the behavior. If three of five reps under one manager show the same pattern, you do not have three sandbaggers — you have one manager whose reps are afraid to miss.

What the numbers actually look like when you run this

Treat every threshold below as a starting point to calibrate against your own history, not a universal constant. The right cut for a company selling $30K deals on a 45-day cycle is not the right cut for one selling $600K deals on a nine-month cycle. Run each threshold backward against two years of closed data and check what it would have flagged before you turn it on.

Calibration delta. Compute the distribution across your whole sales org first. In most B2B teams the middle of the distribution sits within roughly five points of the baseline in either direction. Set your flag at the tail, not the middle — commonly around 12 to 15 points below baseline, sustained across at least three of four quarters. A single quarter at 20 points below means the rep had one weird deal.

How do you detect AE sandbagging in your 2027 forecast — figure 4

Commit-to-pipeline ratio. This one varies enormously by motion. Transactional teams with 30-day cycles routinely run commit at 40–50% of pipeline because the pipeline turns over inside the quarter. Enterprise teams with 6–9 month cycles may sit at 15% legitimately, because most of the pipeline is not for this quarter at all. Normalize by comparing each rep to their own team's median and to their own trailing history, never to a cross-industry number. A rep 40% below their team median with above-median win rates is the interesting case.

Late-quarter accretion. Establish the team baseline. Most teams find that somewhere between 15% and 30% of final commit legitimately arrives in the last two weeks — that is just how quarters end. The flag is a rep running well above their own team's baseline while the deals involved are old. Weight by deal age: commit added late from deals older than 60 days should be scrutinized; commit added late from deals under 30 days old is usually genuine.

Attainment distribution. The strongest single number here is misses per eight quarters. A calibrated rep misses their own commit somewhere between two and three times in eight. Zero misses in eight, combined with a mean attainment above roughly 110%, is the cleanest sandbagging indicator you will find, because it requires no model and no snapshot infrastructure — just closed-won history against stated commits.

Coverage of the flagged population. Expect any reasonable multi-signal rule to flag somewhere in the range of 10–25% of a sales team on first run. If it flags 60%, your thresholds are wrong or your whole comp plan is producing the behavior. If it flags 2%, you have set the bars so high you will only catch the most extreme case, which you already knew about.

How do you detect AE sandbagging in your 2027 forecast — figure 5

Time to run it. The first build is the expensive one: snapshot infrastructure, baseline model or stage-conversion table, and a scorecard is typically a two-to-six-week RevOps project depending on CRM hygiene. After that it is a scheduled job. The weekly monitoring cost to a first-line manager should be under fifteen minutes.

Expected accuracy improvement. Be honest with your leadership about this: the gain is real but bounded. Sandbagging is one contributor to forecast error alongside pipeline quality, stage definition drift, and genuine market volatility. Removing systematic under-call tightens the distribution and, more importantly, removes the *bias* — it stops your forecast from being consistently low by a knowable amount. Measure it as reduction in mean signed error, not just absolute error, because signed error is exactly what sandbagging creates.

Trade-offs: what each detection approach costs you

There is no free version of this. Every approach buys accuracy with something — manager time, rep trust, tooling spend, or false positives — and choosing badly is how detection programs die in their second quarter.

How do you detect AE sandbagging in your 2027 forecast — figure 6

Manual manager judgment. The cheapest option: first-line managers eyeball their reps' history and call it. Cost is zero, and it catches the most obvious cases. It fails on two fronts. Managers are the most conflicted possible observers — a sandbagging rep makes the manager's own roll-up look good — and judgment does not scale past the handful of reps a manager knows well. Reasonable as a stopgap, indefensible as the permanent answer.

Single-signal rules. Pick one metric, usually the calibration delta, set a threshold, flag anyone past it. Fast to build, easy to explain. The problem is false positives, and false positives here are expensive in a way they are not in most analytics: you are accusing a person. Legitimately conservative reps — often your best enterprise sellers, who under-call because they have been burned by procurement — get flagged repeatedly, learn that the system is dumb, and stop cooperating with forecasting entirely. One bad conversation with a top performer costs more than the detection saves.

Multi-signal pattern analysis. Two or more independent signals firing on the same rep across consecutive quarters. This is the defensible standard because the signals fail differently: a conservative rep will trip the calibration delta but not late-quarter accretion; a rep with junk pipeline will trip the ratio but not the attainment distribution. Requiring concurrence collapses the false-positive rate. Cost is the snapshot infrastructure and the discipline to wait a full quarter before acting.

Model-scored probability as the baseline. Buying a forecasting tool that scores every deal gives you a rep-independent second opinion, which is the cleanest possible comparison. Trade-off: cost, integration time, and a subtler problem — if the model is trained on your historical outcomes, and your reps have been sandbagging historically, the model learns their bias. Audit for that explicitly by checking whether model scores also run low against actual outcomes.

How do you detect AE sandbagging in your 2027 forecast — figure 7

Comp redesign instead of detection. The nuclear option: flatten the accelerator curve, add a forecast-accuracy component to variable comp, and remove the incentive at its source. This is frequently the highest-leverage move and the hardest to get approved, because it touches every rep's paycheck and requires the CRO and CFO to agree. It also cuts both ways — pay for accuracy and you create a new incentive to under-sell to hit a number you predicted.

The practical answer for most teams is the bottom of that diagram: run multi-signal detection now, because it is achievable this quarter, and open the comp conversation in parallel, because detection alone treats a symptom the plan keeps re-creating.

Pitfalls that turn a detection program into a trust problem

Confusing conservatism with sandbagging. Some reps under-call because they have been burned. Enterprise sellers who have watched a signed verbal die in legal three times learn to keep things out of commit until paper moves. That is judgment, not manipulation, and if your program cannot tell the difference it will punish exactly the reps whose caution is earned. The distinguishing test is what happens to the deals they hold back: a genuinely cautious rep's held-back deals slip and die at normal rates, while a sandbagger's held-back deals close at rates that were entirely predictable.

Acting on one quarter. Every threshold in this article assumes multi-quarter persistence. A rep can trip three signals in a single quarter because one $400K deal behaved strangely. Building a program that flags on a single quarter guarantees you will have the wrong conversation with someone within your first month, and word of that conversation will reach every rep on the floor before the week ends.

How do you detect AE sandbagging in your 2027 forecast — figure 8

Making it a surveillance product instead of a coaching input. The moment reps believe the forecast tool exists to catch them, the forecast degrades — they stop putting real deals in the CRM at all, and you have traded a measurable bias for an invisible one. Route flags to first-line managers, not to the CRO's dashboard. Frame the conversation around calibration as a professional skill: "your commit is a promise other teams staff against."

Skipping the manager-level analysis. Always aggregate flags by manager before you aggregate by rep. Clustering under one manager is the single most common finding and the one most often missed, because the program is designed to look at reps. A manager who publicly punishes misses will produce sandbagging across their whole team within two quarters, and coaching the reps individually will not fix it.

Exempting the top performer. Your biggest sandbagger is frequently your best closer, which is exactly why nobody wants to have the conversation. If the rule applies to everyone except the person who hit 140%, every rep learns the rule is theater. Apply it uniformly or do not build it.

Ignoring the comp plan. A steep accelerator step at 100% attainment pays a rep meaningfully more for concentrating revenue than for spreading it evenly. If your plan does that, your reps are not cheating — they are optimizing the function you handed them. Flatten the step, or add a modest accuracy component, before you attribute the behavior to character.

How do you detect AE sandbagging in your 2027 forecast — figure 9

Never closing the loop. Detection without a measurable follow-up is just a report nobody reads. When a rep is flagged and coached, set an explicit target — "commit within 10% of actual for the next two quarters" — and re-run the scorecard on schedule. If nothing changes and nothing happens, the program is dead and everyone knows it.

Where the same pattern shows up outside the AE forecast

Sandbagging is not a sales-rep phenomenon; it is what any human does when the cost of missing a number exceeds the cost of beating it. Once you have the detection machinery built for AEs, the same comparison engine transfers to several adjacent forecasts that break plans just as badly.

Renewals and customer success. CSMs forecast renewal likelihood on the same category system as AEs and face the same asymmetry — a CSM who calls a renewal at risk and then saves it is a hero, while one who calls it safe and loses it is negligent. The result is a renewals forecast that runs systematically low, which distorts net revenue retention projections and causes finance to under-forecast the exact revenue stream the board watches hardest. The calibration delta works identically here; substitute renewal outcome for close outcome.

How do you detect AE sandbagging in your 2027 forecast — figure 10

Channel and partner pipeline. Partner-sourced deals are usually forecast by a partner manager relaying a partner's own call, which means you have two layers of optimism or pessimism stacked. Partners frequently over-call to keep their tier status and under-call once a deal is registered and safe. The commit-to-pipeline ratio is nearly useless here, but late-stage accretion and attainment distribution both transfer cleanly.

Marketing-sourced pipeline commitments. Demand gen teams commit to a pipeline number the same way AEs commit to revenue, and when marketing leadership is measured on hitting that number, the incentive to set it low is identical. The tell is the same too: a marketing team that hits its pipeline commit every single quarter, never missing, is managing the commit rather than the pipeline.

Services and professional services bookings. PS forecasts feed staffing decisions with long lead times. A sandbagged PS forecast means you under-hire consultants, then miss delivery dates on the deals that "surprised" you, then take the churn hit four quarters later. This is the clearest illustration of why over-performance against a low commit is not free.

Upstream effect on capacity planning. The reason RevOps owns this rather than sales is that the forecast is an input to at least four other functions: finance's cash model, recruiting's hiring plan, SE and services capacity, and marketing's spend pacing. Each of them builds its own buffer against forecast unreliability, and those buffers stack. Removing systematic under-call does not just improve one number — it lets every downstream team plan against a tighter band, which is where the actual operating leverage lives.

Related questions

Is a rep who always beats their commit necessarily sandbagging?

Not necessarily, but it is the strongest single indicator. Genuine forecasting uncertainty is symmetric — calibrated reps miss roughly as often as they beat. Zero misses across eight quarters means either the commit is set below what the rep knows, or the rep is systematically excluded from risky deals.

How far back should the analysis window go?

Four quarters is the minimum for signal; eight is better. Shorter windows are dominated by single-deal noise, especially in enterprise motions. If your average cycle exceeds six months, weight toward eight quarters so each window contains multiple complete cycles.

Should the CRO see individual rep flags?

Generally no. Route flags to first-line managers and give leadership only the aggregate distribution and manager-level clustering. Individual flags reaching executives turns a coaching tool into a surveillance tool, and reps respond by degrading CRM data quality across the board.

Does an AI-scored probability replace the need for rep calls?

No. The model gives you an independent baseline for comparison, which is the point. Rep calls carry information no model has — a champion leaving, a budget freeze mentioned verbally. The value is in the delta between the two, not in replacing one with the other.

What if the model itself is biased low?

Audit it against actual outcomes the same way you audit reps. If your model trained on historical data from a chronically sandbagging org, it inherited the bias. Check whether model-scored probabilities also run below realized win rates before you use the model as ground truth.

FAQ

What exactly is AE sandbagging?

Sandbagging is when an account executive deliberately understates the probability or timing of deals they expect to close, keeping revenue out of the committed forecast so it can surface later. It is distinct from honest conservatism, which is symmetric and produces misses as well as beats. The defining characteristic is that the withheld revenue reliably materializes.

Why is over-performance against a low commit actually a problem?

Because the forecast is a planning input, not a scoreboard. Finance paces spend against it, recruiting sizes hiring against it, and services and solutions engineering staff against it. A quarter that lands 20% above commit means every one of those teams was sized wrong. Unstaffed late deals turn into slipped implementations and soft renewals two or three quarters downstream.

How many signals should fire before a manager acts?

Two or more, repeating across at least two consecutive quarters. Single-signal detection produces false positives at a rate that will destroy the program's credibility, and the reps most likely to be falsely flagged are experienced enterprise sellers whose caution is well-earned. Requiring concurrent, persistent signals is what makes the flag defensible in a one-on-one.

Who should own sandbagging detection?

RevOps builds and runs the analysis; sales leadership sponsors it; first-line managers act on the output. RevOps owning the mechanism matters because managers are conflicted — a sandbagging rep makes the manager's roll-up look conservative and then heroic. Separating the measurement from the person who benefits from the measurement is the whole point.

Can this be done without a dedicated forecasting tool?

Yes, partially. The attainment-versus-commit distribution needs only closed-won history against stated commits, which any CRM has. The calibration delta needs historical snapshots of rep calls, which requires either field history tracking or a scheduled export job — a few weeks of RevOps work. A commercial forecasting platform mainly buys you the model-scored baseline and the time-series storage.

What is the single fastest thing to check today?

Pull every rep's attainment against their own committed number for the last eight quarters and count misses. Any rep with zero misses and a mean above roughly 110% goes on the watchlist immediately. It takes an afternoon, needs no new tooling, and will usually identify the same people a full multi-signal build identifies three months later.

Sources

flowchart TD S["How do you detect AE sandbagging in yo"] S --> N0["The quarter that closed at 118% and st"] N0 --> N1["How the detection mechanism actually w"] N1 --> N2["What the numbers actually look like wh"] N2 --> N3["Trade-offs: what each detection approa"]
flowchart LR C["How do you detect AE sandbagging in yo"] C --> H0["What the numbers actually look like wh"] C --> H1["Trade-offs: what each detection approa"] C --> H2["Pitfalls that turn a detection program"] C --> H3["Where the same pattern shows up outsid"]

Related on PULSE

Download:
Was this helpful?  
⌬ Apply this in PULSE
Pulse CheckScore reps on the metrics that matterGross Profit CalculatorModel margin per deal, per rep, per territory