How do you actually diagnose B2B SaaS churn — and what's the framework that works in 2027?
Quality
Certified

Diagnose B2B SaaS churn by coding every lost account into one of five types — onboarding failure, value erosion, competitive switch, pricing rebellion, or genuine bad fit — then cohort those losses by tenure and segment, look for any bucket above forty percent, and ask a counter-factual for each. Blended churn rates hide everything.
The quarter the number looked fine and the business was bleeding
A DevOps tooling company sitting around twenty-five million in ARR opened its quarterly business review with a single slide: gross churn, nine percent annualized, flat against the prior quarter. Nobody argued. Nine percent is a defensible number for a company at that stage, and the board had seen worse from comparable portfolio companies. The CRO moved on to pipeline. Two quarters later the same company missed its net-new number by a wide margin, and the post-mortem found that the retention slide had been lying by omission for the better part of a year.
The lie was not in the arithmetic. It was in the aggregation. That nine percent blended a self-serve SMB tier churning north of twenty percent against a small enterprise book that barely moved. It blended first-year logos that never made it through implementation against four-year customers who had renewed three times and finally consolidated onto a competitor's platform. It blended accounts that left because a champion changed jobs against accounts that left because the product genuinely could not do what Sales said it could. Five completely different failure modes, five completely different owners, five completely different fixes — collapsed into one number that pointed at nobody and demanded nothing.
This is the ordinary condition of churn reporting in B2B SaaS. The metric exists, it gets reported, it gets watched, and it produces almost no useful action, because a rate is not a diagnosis. You cannot fix nine percent. You can fix a broken thirty-day onboarding motion. You can fix a qualification gap that lets Sales close companies outside the ideal customer profile. You can fix a renewal process that lets a finance buyer see a twelve-percent uplift for the first time forty-eight hours before the contract auto-renews. But you can only fix those things if the number decomposes into them, and blended churn never does.
The company in question eventually did the decomposition. Sixty churned logos across four quarters, each one coded to a type, each one plotted by month-of-life at the moment the contract died. Forty-two percent were onboarding failures — customers who signed, entered implementation, never reached a first meaningful outcome, and quietly declined to renew. Twenty-two percent were bad fit, meaning the account never resembled the ICP on the day it was signed. Eighteen percent were competitive switches concentrated in months twelve through twenty-four. The remainder split between value erosion and pricing rebellion. That distribution is a work plan. Nine percent is a slide.

What followed was not a customer success reorganization, which is the reflexive response and almost always the wrong one. It was an onboarding redesign that pulled time-to-value from roughly ninety days down to thirty-five, paired with tightened qualification criteria at the top of the funnel that killed a slice of deals Sales had previously been happy to close. Gross revenue retention moved from eighty-one to ninety-one percent over four quarters. Ten points of GRR at that revenue scale is worth roughly two and a half million dollars of protected recurring revenue annually — a return that no amount of CSM headcount would have produced, because headcount does not fix a motion that is structurally broken.
The framework below is that decomposition, generalized. It is not exotic. Most of it descends from work Lincoln Murphy, Nick Mehta, and the early customer success community published years ago, and the tooling to support it has existed since Gainsight and ChurnZero shipped reason-coding fields. What makes it rare in practice is not sophistication. It is the discipline to forbid the word "other."
How the five-type framework actually works
The mechanism has one non-negotiable rule: no account is allowed to die without a type code attached, and "other" is not a permitted value. Every other part of the framework is downstream of that constraint. Allow "other" and roughly a third of your losses will land there, because "other" is where a CSM puts a churn they did not see coming and does not want to explain. A third of your data disappearing into a null bucket destroys any clustering signal you might have found.
The five types are defined by where they cluster in the customer lifecycle, which is why tenure is the first axis of the analysis rather than an afterthought.

Onboarding failure clusters in months zero through six. The account signed, implementation started, and the customer never reached the event that makes the product indispensable — the first automated report that replaced a manual one, the first deploy that used your pipeline, the first deal that closed through your workflow. The signals are visible in product data well before the renewal: no key activation event inside ninety days, weekly active users flat at one or two seats when the contract covered forty, kickoff milestones sliding without escalation. Ownership sits with onboarding and product, not with the CSM who inherited a dead account in month seven.
Value erosion clusters in months six through eighteen. The customer got initial value and then plateaued. What usually happens underneath is personnel: the champion who ran the evaluation changed roles or left the company, and nobody replaced them. Signals are a champion's title change showing up on LinkedIn, a steady decline in weekly active usage against a flat seat count, and a shift in support ticket composition from "how do I do X" toward billing and administrative questions. Ownership is customer success plus marketing, because the fix is champion enablement and multi-threading — making sure at least three people in the account can articulate why the product matters.
Competitive switch clusters in months twelve through twenty-four, sometimes later in enterprise. These are customers who were reasonably happy and got a better offer, or who consolidated three tools into one platform that does eighty percent of what each did. Signals include a sudden re-evaluation request, a competitor's product name appearing in support tickets or integration questions, and an unexpected RFP at renewal. Ownership is product and sales, and the fix runs through win-loss interviews and roadmap decisions, not through CS.
Pricing rebellion clusters at the renewal date itself, regardless of tenure, which is exactly why a rolling-window dashboard tends to miss it. The signals are price-to-value language in quarterly reviews, a finance or procurement buyer joining calls late in the cycle, and downgrade requests that arrive before cancellation requests. Ownership is finance and CS jointly, because the fix is usually a pricing audit combined with a quarterly value review that documents realized ROI before the uplift conversation happens.
Genuine bad fit appears at any tenure and is the type organizations are most reluctant to code honestly. The account was never in the ICP — wrong company size, wrong vertical, wrong use case, wrong technical maturity — and no amount of customer success could have changed that. Ownership is sales and marketing. When this bucket runs above ten percent, the conversation is about MQL definitions and qualification discipline, not about CS accountability.
Running the framework is four steps, and each one exists to kill a specific class of wrong conclusion.

Step one is cohorting the churned accounts by tenure. Pull every logo lost across the last four quarters and plot them by month-of-life at churn. You are looking for visual clusters, not averages. A spike at months three through five means onboarding. A relatively flat distribution with sharp spikes in renewal months means pricing. A creeping rise beginning around month eighteen means erosion or competitive pressure. This one chart eliminates more bad theories than any other artifact a RevOps team produces, and it takes an afternoon to build if the contract dates are clean.
Step two is coding each churn to one of the five types, with "other" disallowed. The CSM, the AE of record, and the renewal owner each pick one. Where they disagree — and they will disagree, often instructively — the analyst breaks the tie using product and support signal data rather than recollection. Disagreement itself is a finding: if Sales codes an account as value erosion and CS codes the same account as bad fit, the two functions have different beliefs about who you sell to.
Step three is running the pattern test. If more than forty percent of churn lands in a single bucket, you have a systemic problem rather than a customer problem. This threshold matters because it changes the response. Forty-two percent onboarding failure does not describe forty-two unlucky customers; it describes a broken motion that every new logo will pass through next quarter. Below that threshold, causes are distributed and the correct response is tactical — tighter cohorts, more granular segmentation, continued monitoring — rather than a company-wide initiative.
Step four is the counter-factual, and it is the step most teams skip. For every churned account, ask: would changing one specific thing have saved this? Not "why did they leave," which produces a stated reason, but "what would we have had to do differently," which produces an action. If the answer is yes and the change is nameable, it becomes a project with an owner. If the answer is no, you have learned something equally valuable — the ceiling of what retention work can accomplish against that segment.
Real numbers, ranges, and what good actually looks like

Benchmarks are where churn conversations go wrong most often, because a number that is excellent in one segment is a crisis in another. Published SaaS benchmark work from sources like OpenView and Bessemer consistently shows the same shape: gross logo churn runs materially higher in SMB than in mid-market, and materially higher in mid-market than in enterprise. SMB gross churn in the fifteen to twenty-five percent annual range is common and not automatically alarming, because SMB customers go out of business, get acquired, and change direction at rates that have nothing to do with your product. Mid-market typically lands somewhere in the eight to fifteen percent range. Enterprise below ten percent, and best-in-class enterprise books run considerably tighter than that.
The practical implication is that a blended nine percent tells you nothing until you know the mix. Nine percent blended across a book that is eighty percent enterprise is bad. The same nine percent across a book that is sixty percent SMB is quite good. This is why segment-first dashboards matter more than accuracy improvements to the aggregate number: executives anchor hard on whatever figure they see first, and un-segmented retention is the single most anchoring-prone metric in the business.
Gross revenue retention and net revenue retention need to be read together and never substituted for each other. GRR measures what you kept from the existing base excluding all expansion — it can never exceed one hundred percent, and it is the honest measure of whether customers stay. NRR includes upsell, cross-sell, and seat growth, which means a strong expansion motion can mask genuinely poor retention. A company reporting one hundred and fifteen percent NRR on seventy-eight percent GRR is not a retention success story; it is a company whose growing accounts are subsidizing a leaking base, and that arrangement stops working the moment expansion slows. Watch GRR for diagnosis and NRR for growth quality.
On sample size: clustering becomes readable somewhere around twenty to thirty coded losses. Below that, one unusual account distorts the distribution badly enough to send you after the wrong fix. At fifty or more, the dominant failure mode is usually unmistakable. If your business churns fewer than twenty logos a year, widen the window to eight quarters rather than forcing conclusions from six data points, and lean more heavily on qualitative buyer interviews to compensate.
The most consequential refinement to the whole framework is weighting by revenue instead of counting accounts. A churned enterprise account at a hundred and twenty thousand dollars a year and a churned SMB account at six thousand are not equivalent events, yet standard account-count analysis treats them identically. Revenue-weighted decomposition fixes this: for each churn type, sum the ARR lost from accounts in that type, then divide by the total ARR that came up for renewal in the period. That yields the true revenue impact of each failure mode rather than its frequency.

The two views frequently disagree, and the disagreement is the point. It is common to find that pricing rebellion represents a small minority of churned accounts but a disproportionate share of churned revenue, simply because the largest accounts are the ones with procurement functions willing to fight an uplift. It is equally common to find onboarding failure dominating the account count while contributing a modest share of lost revenue, because the accounts that never activate tend to be the small ones that never expanded. If you prioritize off account count in that situation, you spend a quarter fixing onboarding and protect very little money. If you prioritize off revenue weight, you build a structured renewal process for accounts above a defined ARR threshold and protect a great deal.
Run both views every quarter and expect the weights to shift as the book matures. A problem that was small in revenue terms last year becomes large this year as those cohorts grow. The only way to catch the transition is to measure both, side by side, on the same slide.
There is also a timing benchmark worth setting expectations against. For a team with reasonably clean contract and product data, coding forty to sixty churns and completing the four-step analysis takes roughly two to four weeks of part-time analyst effort. The analysis is the fast part. Implementing the resulting fix — redesigning an onboarding sequence, restructuring a pricing tier, rebuilding qualification criteria — takes another quarter or two, and measuring whether it worked requires waiting for the affected cohort to reach the tenure window where it used to fail. Plan for a three-to-four-quarter feedback loop on any structural retention fix, and resist the urge to declare victory on leading indicators alone.
Where the framework trades off against the alternatives
The five-type framework is not the only way to attack churn, and it is worth being honest about what it costs and what it does not do.
The main alternative is predictive health scoring — a model that ingests product usage, support volume, sentiment, and engagement signals and outputs a risk score per account. Health scores are genuinely useful and they solve a different problem: they are forward-looking and operational, telling a CSM which account to call on Tuesday. The five-type framework is backward-looking and structural, telling an executive team which motion to rebuild this quarter. Teams that adopt only health scoring end up very good at saving individual accounts and no better at all at reducing the rate at which accounts need saving. Teams that adopt only the retrospective framework build good strategy and miss saves they could have made this week. The two are complements, and the sequencing matters: the framework tells you which signals actually predict loss in your business, which is what makes a health score more than a guess.

A second alternative is the pure exit-interview program — talk to everyone who leaves, aggregate the reasons, act on the themes. This fails in a specific and well-documented way. The person available for the exit conversation is usually the champion, and the champion is the least reliable narrator in the account. They liked the product, they advocated for it internally, they lost the fight. Their account of events will be sympathetic and largely useless: budget was cut, priorities shifted, it wasn't you. The economic buyer tells a different story — the ROI never materialized, or a competitor showed up with a stronger business case, or the renewal price no longer matched the perceived value. Getting to the buyer is harder and it is the only version worth acting on. Run exit interviews, but run them with the person who signed the check.
A third alternative is simply modeling churn financially — building a retention curve, projecting revenue impact, and optimizing spend against it. This is what finance tends to want and it is valuable for planning, but it produces no diagnosis whatsoever. A curve tells you the shape of the decay, not its cause, and two businesses with identical curves can require completely opposite interventions.
The costs of the five-type framework are real. It requires organizational compliance: three functions have to code every loss, consistently, including losses that reflect badly on them. Sales has to code bad-fit churn as bad fit, which is an admission that a closed deal should not have been closed. That is a compensation-adjacent conversation, and it is the most common reason implementations quietly degrade — the codes stay in the system but drift toward whichever category is least politically expensive. It also requires clean data: contract start dates, true churn dates rather than notification dates, segment tags that mean the same thing across systems. Most organizations discover during step one that their contract data is worse than they believed, and that discovery alone is worth the exercise.

Adjacent to all of this sits a workflow most RevOps teams already own and rarely connect to churn: win-loss analysis on the new-business side. The two programs ask structurally identical questions at opposite ends of the customer lifecycle, and the answers correlate more than teams expect. If competitive switches are rising in months twelve through twenty-four, the same competitor's name is almost certainly appearing more often in late-stage new-business losses. If bad-fit churn is climbing, the qualification gaps that produced it are visible in the win rates of the segments that produce it. Running the two analyses in the same review, on the same slide, with the same competitor and segment taxonomy, roughly doubles the signal from either one alone. The taxonomy has to match for that to work, which is a small governance decision with outsized returns.
The same logic extends downstream into expansion. The tenure windows where accounts churn are usually the same windows where accounts fail to expand, because both outcomes trace back to whether the customer reached and then deepened value. A team that has already built the tenure-cohort view for churn can reuse it directly to find where expansion stalls, and the mid-lifecycle intervention that reduces month-eight churn tends to be the same intervention that produces month-ten upsell. Retention and expansion programs are frequently staffed and budgeted separately when they are, mechanically, the same program measured at two different thresholds.
The pitfalls that keep good analysis from producing good decisions
The first and most damaging pitfall is reporting blended churn without segment cohorts. It has already been described, but it is worth stating as a rule rather than an anecdote: never let an un-segmented retention number reach an executive audience. Build the dashboard so segment is the default view and blended is a drill-up, not the reverse. The ordering of the slide determines the ordering of the conversation, and once a leadership team has anchored on a comfortable aggregate, every subsequent segment-level finding gets heard as an exception rather than as the actual picture.
The second is treating churn as a customer success problem when a substantial share of it is an acquisition problem. Industry work from the customer success community has argued for years that a meaningful fraction of churn — often cited in the neighborhood of thirty percent — originates before the contract is ever signed, in onboarding design and in qualification. Customer success cannot retain a customer who should never have been sold. When bad-fit churn runs above ten percent, holding the CS team accountable for it is not just unfair, it actively obscures the fix. The correct forum is a joint marketing and sales conversation about MQL definitions and qualification gates, and it should carry the same revenue weight as any pipeline discussion.

The third is over-reliance on champion exit interviews, covered above, with one addition: the timing is as wrong as the interviewee. By the time someone agrees to an exit conversation, the decision is months old and the narrative has been rehearsed internally. The higher-yield conversation happens at the moment of the first negative signal — the deferred QBR, the seat reduction, the ticket that mentions a competitor's integration — not at the moment of cancellation.
The fourth pitfall is coding secondary causes instead of primary triggers. Most churns have three or four contributing factors, and analysts trying to be thorough will code all of them, which spreads the distribution flat and destroys the clustering signal that makes step three work. Code the primary trigger: the event that moved the account from renewing to not renewing. Capture secondary factors in a free-text field for qualitative review, but keep exactly one coded type per account. Thoroughness in the coding field is the enemy of the analysis.
The fifth is using a rolling-window dashboard for a phenomenon that concentrates in renewal months. Pricing rebellion is nearly invisible in a trailing-ninety-day view because it fires in bursts tied to contract anniversaries. If your renewals cluster in January and July, a rolling view will show two unexplained bumps and no cause. Plot against contract anniversary in addition to calendar time, and the pattern resolves immediately.
The sixth is declaring a fix successful on the wrong evidence. If onboarding failure was forty-two percent of churn and you redesign onboarding, the honest test is whether the cohort that went through the new motion churns less at months three through five than the cohort that went through the old one. That takes two quarters minimum to observe. Improvements in activation rate, time-to-first-value, or CSAT are leading indicators worth watching, but they are not the result, and reporting them as the result is how retention initiatives get declared complete while the underlying rate stays flat.
The seventh is letting the coding discipline decay after the first cycle. This is the quiet failure mode. The first quarter's analysis is careful because it is a project with attention on it. By the third quarter the codes are being entered at the end of the month by whoever has time, and the distribution starts drifting toward whichever category is easiest to defend. The countermeasure is an audit: each quarter, pull a random sample of ten coded churns and independently re-code them from product and support data alone. If independent re-coding disagrees with the recorded code more than about twenty percent of the time, the data has degraded and the analysis built on it should not be trusted until it is rebuilt.

Tooling deserves a closing note precisely because it is so rarely the constraint. Gainsight, ChurnZero, and Catalyst all support reason-coding and health frameworks natively. Cohort visualization is straightforward in Looker, Sigma, or any BI layer sitting on a warehouse — and honestly, a well-built spreadsheet handles sixty churns without complaint. The bottleneck is never the platform. It is whether the organization is willing to forbid "other," code losses honestly against its own interests, and wait three quarters to find out whether the fix worked.
Related questions
How is this different from a customer health score?
A health score is predictive and per-account — it tells a CSM which customer to call this week. The five-type framework is retrospective and structural — it tells leadership which motion to rebuild this quarter. Run the framework first; its findings determine which signals belong in the health score.
What if we churn too few accounts to find a pattern?
Widen the analysis window to eight quarters instead of four, and weight the qualitative side more heavily. With fewer than twenty coded losses, lean on economic-buyer interviews and product-signal review rather than forcing statistical conclusions from a handful of data points.
Should churn reason codes affect sales compensation?
Generally no, at least not directly. The moment a bad-fit code costs someone money, coding accuracy collapses and the data becomes worthless. Use the aggregate bad-fit rate to inform qualification criteria and ICP definition instead of clawing back individual commissions.
Who should own the churn analysis itself?
RevOps or a dedicated revenue analyst — someone structurally neutral. If customer success owns the analysis, onboarding and bad-fit findings get softened; if sales owns it, bad-fit findings disappear entirely. Neutral ownership is what makes honest coding possible.
How often should this run?
Quarterly for the full four-step analysis, with monthly coding as losses occur. Coding at the moment of churn is far more accurate than reconstructing reasons three months later, but the pattern analysis needs a quarter's worth of volume to be meaningful.
FAQ

What's the difference between gross revenue retention and net revenue retention?
GRR measures revenue retained from the existing base excluding all upgrades and expansion, so it can never exceed one hundred percent. NRR includes expansion and can exceed it. Read them together: strong NRR on weak GRR means expansion is masking a leaking base, which stops working the moment growth slows. Use GRR for diagnosis, NRR for growth quality.
How many churned accounts do I need before patterns are trustworthy?
Clustering becomes readable around twenty to thirty coded losses. Below that, a single unusual account distorts the distribution enough to send you after the wrong fix. At fifty or more, the dominant failure mode is usually unmistakable. If annual volume is lower than twenty, widen the window rather than forcing a conclusion.
Why does cohorting by tenure matter so much?
Because churn causes shift by lifecycle stage. Onboarding failures dominate the first ninety days; value erosion and competitive switches appear after twelve months. Blend all tenures together and the distinct patterns cancel each other out, leaving an average that describes no actual customer and points to no actual fix.
Can one churn have multiple causes?
Almost always, but code only the primary trigger — the event that moved the account from renewing to not renewing. Capture contributing factors in free text for qualitative review. Coding every contributing cause flattens the distribution and destroys the clustering signal the whole analysis depends on.
Is it worth doing without the full framework?
Partial versions help, but they are prone to treating symptoms. The value comes from forcing measurement, categorization, and prioritization in that order. Without it, teams reliably fund the most visible problem rather than the most expensive one — usually onboarding, when the money is actually walking out at renewal.
How long before a churn fix shows up in the numbers?
Coding and analysis take two to four weeks with clean data. Implementing a structural fix takes a quarter or two. Verifying it requires waiting for the affected cohort to reach the tenure window where it previously failed — so plan on a three-to-four-quarter feedback loop before the rate itself moves.
Sources
- https://openviewpartners.com/saas-benchmarks/
- https://www.bvp.com/atlas/state-of-the-cloud
- https://www.gainsight.com/blog/
- https://www.churnzero.com/blog/
- https://sixteenventures.com/
- https://hbr.org/2014/10/the-value-of-keeping-the-right-customers
- https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights
- https://a16z.com/saas-metrics/
- https://www.klipfolio.com/resources/kpi-examples/saas/gross-revenue-retention
Related on PULSE
- How do you diagnose whether your churn is a product problem or a customer-success problem?
- What is an inbound qualification framework, and which one actually works (BANT, MEDDPICC, Sandler, etc.)?
- How do you actually diagnose stuck deals in your pipeline?
- Top 10 questions to diagnose a stalled deal in the sales cycle
- What question would you ask a rep who consistently loses deals at the proposal stage to diagnose the real issue?
- Top 10 questions to diagnose why a deal is stuck in negotiation
This page will be disappearing soon. Save it to your device for $1 — or read it free while it is here.
@Kory-White- · if Venmo asks, the last 4 of my number are 2012
This page is gone.
This one is off the shelf now. $1 keeps it on your phone for good — the whole page, pictures and diagrams included.










