How are 2027 RevOps teams calculating ROI on AI SDRs when humans are still closing?
PULSEKNOWLEDGE LIBRARYQuality
Certified

RevOps teams are calculating AI SDR ROI by attributing only the sourcing and early-stage portion of a deal to the agent, then dividing that weighted revenue by fully-loaded cost including oversight. Because humans are still closing, the honest measures are cost-per-qualified-meeting, stage-one conversion, and pipeline velocity — never raw closed-won credit.
The scenario that breaks the old spreadsheet
Picture a mid-market SaaS team heading into a 2027 planning cycle. They ran AI SDR agents for three quarters. The agent platform costs roughly $1,500–$2,500 per seat per month, data enrichment adds a few thousand per seat per year, and a RevOps analyst spends about half their week tuning sequences and reviewing transcripts. Outbound meeting volume roughly tripled. Closed-won revenue rose 18%.
The CFO asks a simple question: what was the return? The sales leader opens a dashboard that says "AI-sourced pipeline: $9.4M" and "AI-influenced closed-won: $2.1M," divides by the software bill, and reports 9x. Finance rejects it in under a minute, and correctly so, because that number embeds three unforced errors.
The first error is influence stacking. In most CRMs, an opportunity is flagged "AI-influenced" if any AI-generated activity exists anywhere on the record. When the agent handles top-of-funnel across the whole territory, nearly every deal picks up that flag — including deals that came from an inbound demo request, an existing customer expansion, or a warm partner referral where the agent's contribution was one ignored follow-up email. Influence-flag coverage above roughly 70% of all closed-won is a strong tell that the flag is measuring reach, not contribution.

The second error is the denominator. Seat cost is usually less than two-thirds of true cost. The missing third is human: the analyst hours, the deliverability and domain warming work, the quarterly prompt and messaging refresh, the compliance review in regulated verticals, and — most expensive of all — AE time spent sitting in meetings that never should have been booked. If AEs run 30 AI-booked meetings a month at an hour each including prep and notes, and a third are unqualified, that is ten hours of loaded AE capacity burned monthly per rep.
The third error is timing. Deals sourced in Q1 close in Q3 or Q4 on a six-to-nine-month enterprise cycle. Comparing this quarter's software spend to this quarter's closed-won compares two unrelated cohorts. In a growing program the spend is always ahead of the revenue, so the ratio understates a good program during ramp and overstates it the moment spend flattens.
The fix that finance actually accepts is a cohort-and-weight model: freeze a cohort of opportunities by the month the agent first created contact, wait for the cohort's natural cycle length, apply an explicit attribution weight to the sourcing and early-nurture stages only, and divide by fully-loaded cost for that same cohort period. It reports later and lower than the dashboard number, and it survives audit.

How the attribution mechanism actually works
The mechanism has three moving parts: a clean sourcing definition, a stage-scoped weight, and a cohorted denominator. Get any one wrong and the output is decorative.
Sourcing definition. Draw a hard line between *sourced* and *touched*. Sourced means the agent made first contact with an account or buying-group member who had no prior open opportunity and no inbound activity in a defined lookback window — 90 days is the common choice. Everything else is touched. In practice, teams implement this as a field written once, at opportunity creation, and locked thereafter. Locking matters: if the field recalculates, the agent quietly accrues credit on deals it joined late. Sourced-only crediting typically produces a much smaller pool than the influence flag — it is not unusual for a "70% influenced" territory to be 20–30% genuinely sourced.
Stage-scoped weight. The agent's causal contribution is concentrated where it acts: identifying the account, reaching a person, earning the first conversation, and keeping cadence through early evaluation. It contributes essentially nothing to security review, procurement, mutual action plans, or negotiation — humans are still closing those. So the weight should be applied to stages, not to the deal as a whole. A defensible starting split for a five-stage funnel gives the agent a majority share of the value created in stage one, a modest share through stage two, and zero from the point the deal enters formal evaluation. Translated to whole-deal terms, that usually lands the agent between 15% and 35% of a sourced deal's revenue — and 0% of a deal it merely touched.

Cohorted denominator. Sum every cost incurred during the cohort's sourcing month, not the closing month: platform seats, enrichment and verification credits, sending infrastructure and domain costs, the fractional analyst FTE, the amortized quarterly tuning, and a loaded hourly charge for AE time spent in agent-booked meetings that were disqualified. That last line is the one teams skip, and it is often the difference between a program that looks healthy and one that isn't.
Two implementation details decide whether this survives contact with a real CRM. First, write the sourcing tag and the stage-entry timestamps as CRM fields, not as report filters — filters drift, fields are auditable. Second, snapshot the weight table with a version number and an effective date. When you retune weights next year, you want last year's reported ROI to remain reproducible rather than silently rewritten by the new model.
The numbers that actually move, and the ones that don't
Stop reporting a single ROI multiple. Report a small panel, because different metrics fail in different directions and the panel makes the failure visible.

Cost per qualified meeting. Not cost per meeting booked — cost per meeting an AE accepted after the call. Compute it as fully-loaded cohort cost divided by accepted meetings. This is the cleanest efficiency measure the program has, because it lands inside the window where the agent is genuinely causal. Track it monthly and watch the trend, not the absolute: a rising cost-per-qualified-meeting while volume climbs means you are buying reach at declining quality.
Meeting acceptance rate. Of meetings the agent books, what share does the AE keep after the first call? This is the single most diagnostic number in the whole program. A low acceptance rate does not merely waste AE hours — it corrupts every downstream metric, because unaccepted meetings still inflate the "meetings booked" figure the vendor dashboard shows. Set a floor, and treat a breach as a targeting problem rather than a volume problem.
Stage-one-to-stage-two conversion, split by source. Run agent-sourced and human-sourced opportunities as separate cohorts through the same funnel. If agent-sourced deals convert materially worse at the first real gate, the agent is booking on curiosity rather than intent, and the volume advantage is partly illusory. A modest gap is normal and acceptable when the volume multiple is large enough to overcome it; a large gap is a targeting defect.

Win rate and average deal size, split by source. These are your guardrails, and the expectation is that they should be roughly *flat* between cohorts. AI SDRs do not make deals bigger and do not make humans close better — they change how many at-bats exist. If agent-sourced deals show a sharply lower win rate or smaller average size, the program is shifting the mix toward easier-to-reach, lower-fit accounts, and pure volume growth is masking a quality decline.
Pipeline velocity, decomposed. Velocity is opportunities times win rate times average deal size divided by cycle length. Decompose it rather than reporting the composite, because the agent almost exclusively moves the first term. Time-to-first-meeting is where the compression shows up — the agent works the whole list simultaneously and follows up on schedule indefinitely, so the wait between list load and conversation shrinks substantially. Late-stage cycle length typically does *not* improve and can lengthen slightly, since agent-sourced buyers often start colder and need more human work to build consensus. If your composite velocity number is up but stage-three-onward duration is up too, you accelerated the cheap half of the funnel.
Deliverability health as a cost driver. Bounce rate, spam-complaint rate, and domain reputation belong on the ROI panel, not a separate deliverability report. Poor list hygiene raises enrichment spend, burns sending domains that cost real money and weeks to replace, and depresses reply rates across every campaign including the human ones. A degrading reputation is a future cost increase that shows up in the ROI ratio a quarter late.
Payback period. For a program in ramp, the more useful figure than a multiple is: how many months from a cohort's sourcing month until that cohort's weighted revenue exceeds its loaded cost? Report the observed payback for closed cohorts and a modeled payback for open ones. Finance is far more comfortable with a payback month than with a multiple whose denominator they can't inspect.

One number to actively distrust: total messages sent, sequences launched, or "activities logged." Agent throughput is effectively unbounded and therefore carries no information about value. Any metric that scales linearly with how hard you turn the volume dial is a vanity metric by construction.
Trade-offs between the models you could pick
There is no single correct attribution model; there are four common ones, each wrong in a predictable direction. Choose deliberately, document why, and hold it constant long enough to compare periods.
Sourced-only, binary credit. The agent gets 100% of revenue on deals it sourced, 0% on everything else. Enormously simple, easy to audit, and defensible to finance. It overstates in one specific way: a sourced deal that a human then rescued through six months of relationship work still credits fully to the agent. Best fit for high-velocity mid-market motions with short cycles and small buying groups, where the gap between sourcing and closing is genuinely narrow.

Stage-weighted. The model described above — the agent earns a defined share of the value created in stages one and two, and nothing after. It is the most causally accurate option and the most work to build, because it requires reliable stage-entry timestamps and a versioned weight table. Best fit for enterprise motions with long cycles, where the human contribution after first meeting genuinely dominates.
Time-decay. Credit weights each touch by recency to close. Popular in marketing attribution, and actively misleading here: it systematically penalizes the agent, whose entire contribution is furthest from the close, while rewarding whoever happened to send an email during the negotiation. Avoid it for SDR programs specifically.
Holdout testing. Withhold agent outreach from a matched slice of the territory — a randomly assigned account list, matched on segment and size — and compare pipeline generated per account against the treated slice. This is the only model that measures *incremental* rather than *attributed* value, and it answers the question the CFO is actually asking: what would have happened anyway? The costs are real: you sacrifice coverage on the holdout, you need enough accounts for the difference to be readable above noise, and you have to wait a full cycle. Most teams run one holdout per year at planning time to calibrate, then use a cheaper model for monthly reporting.

The meta trade-off worth naming: precision costs credibility when nobody can explain the model. A stage-weighted model with a versioned table, reviewed with finance once, beats a bespoke seventeen-input model that only one analyst understands. If the person presenting the number cannot reconstruct it on a whiteboard, it will not survive the next budget cycle regardless of how correct it is.
Pitfalls that quietly inflate the number
Counting booked meetings instead of held-and-accepted meetings. No-shows can run high on cold agent-sourced outreach, and vendor dashboards typically count the booking. Instrument held and accepted separately, and reconcile against calendar data rather than trusting the sequence tool's own tally.
Letting the influence flag drift into the ROI model. Someone will eventually build a report on "AI-influenced pipeline" because it is the bigger, better-looking number, and it will end up in a board deck. Kill the metric at the source: either remove the flag from reportable fields or rename it to something that cannot be mistaken for credit.

Double-counting across the stack. If your marketing attribution model already credits a nurture program, and the agent's sequence credits the same opportunity, the same dollar appears in two ROI calculations presented to the same executive. Reconcile at the deal level once a quarter and confirm total attributed revenue across all programs does not exceed total closed-won.
Comparing agent output to a fully-ramped human rep. The honest comparison is against the marginal human SDR you would otherwise hire — including ramp time, tooling, management overhead, and attrition — not against your best tenured rep's best quarter. Ramp alone typically means a new human rep produces little for the first several months, and leaving that out understates the agent's genuine advantage.
Ignoring the AE-hours line. This is the pitfall with the largest dollar impact. Disqualified meetings consume the most expensive capacity in the org. Price AE time at a loaded rate, multiply by disqualified meetings, and put it in the denominator every period. Teams that add this line often discover their ROI drops by a quarter or more — and that the fastest way to raise it is tightening targeting, not buying seats.

Never revisiting the weights. The weight table encodes an assumption about where the agent contributes. As the agents handle more of early evaluation, that assumption ages. Re-derive the weights from a holdout annually, version the change, and restate forward only.
Attributing away a channel conflict. If agent outreach is hitting accounts that inbound or partner motions were already working, you may be reattributing existing pipeline rather than creating it. The 90-day lookback in the sourcing definition catches most of this; a quarterly overlap audit against inbound and partner-sourced lists catches the rest.
Reporting a multiple with no confidence interval. Small cohorts produce wild ratios. If a cohort has fewer than roughly thirty closed opportunities, report the panel — cost per qualified meeting, acceptance rate, conversion by source — and explicitly decline to report a single ROI figure until the sample supports it.
Related questions
Should the AI SDR be measured against a quota like a human rep?
No. Quota assumes finite capacity and personal accountability; agent capacity is elastic and its output quality is a function of targeting and messaging you control. Set thresholds on efficiency and quality — cost per qualified meeting, acceptance rate — and let volume float.
Who owns the AI SDR number in the org?
RevOps owns the model and the reporting; sales leadership owns the acceptance bar and targeting; marketing owns messaging alignment. Split ownership fails when the vendor dashboard becomes the source of truth. The number must be reproducible from CRM data by someone with no vendor login.
How long before a program's ROI is readable?
At minimum one full sales cycle plus a quarter — so six to nine months for enterprise, less for velocity motions. Efficiency metrics like cost per qualified meeting are readable within about six weeks; revenue-based ROI simply is not, and reporting it early is guessing.
Does the model change if agents start running discovery calls?
Yes, materially. If agents conduct qualifying conversations rather than only booking them, their contribution extends into stage two, and the weight table should shift accordingly. Re-derive from a holdout rather than adjusting by intuition, and version the change with an effective date.
What if we can't get clean stage timestamps?
Fix that before building any weighted model. Use sourced-only binary credit in the meantime — it is cruder but honest. A weighted model built on unreliable stage data produces precise-looking numbers with no underlying accuracy, which is worse than a simple model everyone knows is approximate.
FAQ
Why not just compare closed-won before and after deploying AI SDRs?
Because too much changes simultaneously — headcount, territory, pricing, market conditions, competitive dynamics. A before/after comparison attributes every one of those shifts to the agent. A matched holdout run concurrently controls for all of them, which is why it remains the gold standard even though it is more expensive to run.
Is a lower win rate on agent-sourced deals automatically a failure?
Not automatically. If the agent produces several times the opportunity volume at a fraction of the cost, a moderately lower win rate can still net more revenue per dollar. It becomes a failure when the gap is large enough that AE capacity is being consumed by deals that were never winnable — at which point the constraint is human closing capacity, and adding volume actively destroys value.
How should we handle deals where the agent sourced the account but a human sourced the actual champion?
Split at the buying-group level rather than the account level. Credit the agent for contacts it genuinely created and the human for contacts they created, then weight by which contact drove the opportunity forward. If your CRM cannot express contact-level sourcing, default to crediting whoever created the contact that appears on the first held meeting.
What is the minimum instrumentation needed to start calculating this properly?
Four things: a locked sourcing field written at opportunity creation, stage-entry timestamps, a meeting-accepted flag set by the AE after the first call, and a cost ledger that includes human oversight hours. Without the accepted flag you cannot measure quality; without the cost ledger you cannot measure cost. Everything else is refinement.
Should the same model apply across segments?
The framework should be identical; the weights should not. Mid-market deals with small buying groups sit closer to the sourcing event, so the agent's share is legitimately higher. Enterprise deals with large committees and long procurement cycles concentrate value in the human close. Use one model with segment-specific weight tables, versioned together.
How do we present this to a skeptical CFO?
Lead with the denominator. Show the fully-loaded cost ledger line by line — including AE hours on disqualified meetings — before showing any revenue figure. A model that opens by admitting its costs earns the credibility to make a revenue claim; one that opens with a 9x multiple invites an audit of exactly the line you left out.
Sources
- Gartner — Sales research and insights
- Forrester — B2B sales and marketing research
- HubSpot Knowledge Base — Attribution reporting
- Salesforce Help — Campaign influence and attribution
- Gong Labs — Sales research and data
- McKinsey — Growth, marketing and sales insights
- Harvard Business Review — Sales and marketing topic hub
- SaaStr — SaaS go-to-market and sales content
- Google — Measuring incrementality and holdout experiments
Related on PULSE
- How should RevOps redesign the 2027 pipeline review cadence when AI predicts stage duration better than humans?
- What happens to pipeline coverage ratio when 2027 AI agents auto-remove stale deals 3x faster than humans?
- Top 10 Closing Coaching Techniques for SDRs
- When should we hire our first account executive if revenue is $5M ARR and the founder is still closing?
- How do longer sales cycles in 2027 change the role of customer references in deal closing?
- Should I Hire a Fractional CRO If I Am the Founder Still Closing Every Big Deal?
This page will be disappearing soon. Save it to your device for $1 — or read it free while it is here.
@Kory-White- · if Venmo asks, the last 4 of my number are 2012
This page is gone.
This one is off the shelf now. $1 keeps it on your phone for good — the whole page, pictures and diagrams included.









