Pulse - Value Added
FRACTIONAL CRO · MARYLAND-BASED, NATIONWIDE · $0→$200M

Kory White

RevOps & Revenue Leadership

Get a 30-minute revenue checkup — Kory reviews your pipeline and forecast, then names the 1–2 fixes that move revenue fastest. 25 yrs scaling teams $0→$200M.

30-minute revenue checkup →
Hire a Fractional CROHow We Help?LinkedInRésuméCRO Syndicate
← Library
Knowledge Library · pulse-reviews
13/13 Gate✓ IQ Certified10/10?

How do you establish a cross-functional data dictionary for revenue metrics before an IPO?

PULSEKNOWLEDGE LIBRARY
pulserevops.com
KnowledgeHow do you establish a cross-functional data dictionary for revenue metrics before an IPO?
📖 3,372 words🗓️ Published Aug 16, 2026
Direct Answer

Establish a cross-functional data dictionary by naming one owner, freezing the 15–20 metrics that appear in your S-1, and writing each definition with its exact calculation logic, source system, exclusions, and time convention. Get Finance, Data Engineering, and RevOps to co-sign every entry, store it under version control, and enforce it through a semantic layer.

The scenario that forces the issue

A late-stage SaaS company sits eight months from a targeted filing. The CFO's board deck says ARR is one number. The CRO's Salesforce dashboard says something 4% higher. The product team's usage-based revenue report says something different again, because it counts overage in the month consumed while billing counts it in the month invoiced. Nobody is lying. Three teams are computing three defensible things and calling all of them "revenue."

That gap is survivable while private. It is not survivable in an S-1, where the same metric appears in the prospectus, the MD&A, the investor deck, and — after listing — every quarterly release. Underwriters' counsel and the audit team will ask you to trace a reported number backward to a source table and forward to a disclosure. If the trace fails, or if two internal systems disagree, the response is not a polite note. It is a diligence request that consumes weeks of your finance team's calendar at exactly the moment they have none to spare.

The failure almost never begins in the data warehouse. It begins in language. Someone in 2022 built a Salesforce field called Closed_Won_Revenue__c. It was accurate then, when the company sold one annual subscription SKU. Then the company added implementation services, then a usage tier, then multi-year contracts with ramped pricing, then a reseller channel with net-of-fee recognition. Each addition was written into the billing system correctly and into the CRM field carelessly. Four years later the field name still says "revenue" and the field contents mean roughly six things depending on the deal shape.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 1

A data dictionary is the artifact that ends that ambiguity. It is not documentation in the decorative sense — a wiki page nobody opens. It is a governed contract stating that when any employee, dashboard, board deck, or SEC filing says "Net Revenue Retention," there is exactly one calculation behind it, one source of record, one set of exclusions, and one named human accountable for changing it. Everything else in this answer is mechanics in service of that single property.

The useful framing: you are not documenting your metrics. You are reducing the number of legitimate answers to each metric question from many to one, and then making the reduction stick under pressure. Documentation is the byproduct. Consensus enforced by tooling is the product.

How the mechanism actually works

The workable pattern has four layers, and skipping any one of them is where most pre-IPO efforts quietly fail.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 2

Layer one — the metric inventory. Start from the output, not the input. Pull your most recent board deck, your last three investor updates, and the standard S-1 metric set for companies in your category. That intersection is your scope. For most B2B SaaS businesses it lands between 15 and 25 metrics: ARR or ARR-equivalent, net new ARR decomposed into new, expansion, contraction, and churn; gross revenue retention; net revenue retention; customer count by definition-of-customer; average contract value; CAC and CAC payback; gross margin by revenue stream; revenue by segment; revenue by geography; deferred revenue and remaining performance obligation; and whatever product-specific metric your category expects. Resist the urge to catalog all 400 Salesforce fields. Field-level documentation is a data-engineering hygiene project; a pre-IPO dictionary is a disclosure-integrity project, and conflating them is the most common way to spend six months and ship nothing usable.

Layer two — the definition record. Each metric gets a structured entry, not a paragraph. Minimum fields: plain-language definition a new board member could read; the exact SQL or formula, not a description of it; the system of record; the specific tables and columns; every exclusion stated affirmatively ("excludes professional services revenue," "excludes contracts in a trial-conversion state," "excludes reseller gross-up"); the time convention — point-in-time versus average-over-period, and which calendar boundary; the currency treatment and FX rate source for multi-currency revenue; the named owner; the approver set; the effective date; and a note listing teams that have historically misread it. That last field feels soft and is the highest-value one in the record, because it converts institutional folklore into written warning.

Layer three — the enforcement surface. A definition that lives only in a document will be re-derived by an analyst under deadline pressure within a quarter. The fix is a semantic layer: dbt metrics, a LookML or Looker model, Cube, or an equivalent, where the calculation is defined once and every dashboard, extract, and notebook consumes it by reference. When a VP asks for "NRR by segment," the analyst does not write a query — they call a metric. This is the single change that most reliably prevents recurrence, because it makes the governed path cheaper than the ungoverned one. Governance that fights convenience loses.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 3

Layer four — change control. Store dictionary entries as files in a version-controlled repository, and require a pull request with approval from at least two of Finance, Data Engineering, and RevOps for any change. The commit history becomes your audit trail. When an underwriter asks whether your NRR definition changed between the periods presented, you show a diff with a date and two approver names instead of reconstructing a story from Slack.

The loop matters more than the boxes. A dictionary without a periodic trace audit degrades silently — someone builds a one-off extract for a customer QBR, it gets forwarded, it becomes a recurring report, and now an ungoverned number circulates with your logo on it.

Real numbers, ranges, and benchmarks

Concrete planning figures, drawn from how this work actually sequences rather than from any single vendor's methodology.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 4

Timeline. From kickoff to enforced adoption, budget three to six months. Weeks one and two: scope the metric list and pull the three-way comparison of how each metric is currently computed by Finance, RevOps, and Product. Weeks three through six: draft definitions for the top ten metrics — the ones on the cover of the board deck — and run them through co-sign. Weeks seven through twelve: implement in the semantic layer and repoint the top dashboards. Month four onward: extend to the long tail, run the first trace audit, and start the monthly council cadence. Teams that try to compress this below eight weeks generally produce a document without an enforcement layer, which is the failure mode that looks like success for one quarter.

Effort. Expect roughly 0.5 FTE from RevOps as the driving owner, 0.25 FTE from a finance manager, and 0.25–0.5 FTE from a data engineer for the semantic-layer implementation. The owner needs write access to the semantic layer and to CRM validation rules; an owner with responsibility and no write access produces meeting notes, not definitions.

Discrepancy expectations. When you first run the three-way comparison, plan for a meaningful fraction of your metrics to disagree across teams — in practice the disputes cluster in predictable places rather than spreading evenly. The recurring offenders: multi-year ramped contracts (does year-one ARR use the current rate or the annualized average?), mid-term upgrades (does expansion count on the amendment date or the next renewal boundary?), usage overage (consumed period or invoiced period?), the definition of a customer for retention math (legal entity, billing account, or parent hierarchy?), and FX (spot at period end, average for the period, or constant currency?). Write those five down first. They will account for most of your reconciliation pain.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 5

Audit dry run. Schedule a pre-audit trace with your external audit firm six to nine months before the target filing date. Give them ten to fifteen key metrics and ask them to walk each one from the definition to the source system to the reported number in your most recent board deck. Two things surface reliably: manual adjustments living in someone's spreadsheet between the warehouse and the deck, and timezone handling for global revenue that shifts a small but nonzero amount of bookings across period boundaries.

Cadence. Monthly council meetings of sixty minutes through the pre-filing period; quarterly after listing, with the same rigor. Post-IPO, a metric definition change is a disclosure question, not an internal one — the SEC's guidance on key performance indicators expects companies to disclose when and why a KPI definition changes and, where material, to recast prior periods. Treat every proposed change after your first public filing as something your disclosure committee reviews.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 6

Scope discipline. Fifteen to twenty-five metrics governed properly beats two hundred documented loosely, every time. The long tail can wait; the cover-of-the-deck metrics cannot.

Trade-offs and the alternatives people actually consider

There are three realistic implementation paths, and the right one depends on your data-team maturity more than your revenue scale.

Path one — spreadsheet plus wiki. Fastest to start, zero tooling cost, and genuinely adequate for a company still six or more quarters from filing. The trade-off is that it has no enforcement surface: nothing stops an analyst from writing a different query, and drift is invisible until someone notices two decks disagree. Most companies start here and should — just recognize it as a staging ground, not a destination. The signal that you have outgrown it is the first time two teams present conflicting numbers in the same meeting.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 7

Path two — version-controlled markdown plus a semantic layer. Definitions live as files in a repository with pull-request review; calculations live in dbt metrics, LookML, or an equivalent. This is the sweet spot for most pre-IPO companies. You get a real audit trail, a real enforcement mechanism, and no new procurement cycle if you already run dbt or Looker. The cost is that it requires a data engineer's ongoing participation and it puts the dictionary in a tool that finance stakeholders may not open, which means you also need a rendered, readable view for non-technical reviewers.

Path three — a dedicated data catalog. Purpose-built cataloging tools give you lineage visualization, business glossary features, and access controls that Finance and Legal will find familiar. The trade-off is procurement time, cost, and — critically — that a catalog documents lineage but does not by itself enforce calculation. A catalog without a semantic layer still permits two analysts to compute NRR differently; it just makes the divergence easier to spot afterward. Buy the catalog for discoverability and governance workflow, not as a substitute for the semantic layer.

A cross-cutting trade-off worth naming: strictness versus adoption. A dictionary that blocks every ad-hoc query will be routed around within weeks — analysts will export to a spreadsheet and work there, and you will have lost visibility entirely. The more durable posture is to make the governed metric trivially easy to call, allow ad-hoc exploration freely, and draw the hard line only at externally-facing numbers: anything in a board deck, investor update, press release, or filing must come from the governed definition. That boundary is defensible, enforceable, and does not make your analysts hate the system.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 8

The same logic applies to the adjacent workflows this project inevitably touches. Sales compensation runs on bookings definitions that often differ from finance's revenue definitions — and legitimately so, since commission on a signed multi-year deal is not the same event as recognized revenue. Do not force those into one number. Do define both in the dictionary, state explicitly that they differ, and document the reconciliation between them. The same goes for pipeline and forecast metrics, which live in the CRM and use stage-based definitions that never appear in a filing but absolutely appear in the diligence conversation about how you manage the business. Marketing's pipeline-sourced attribution is another neighbor: it will never be an S-1 metric, but if your investor narrative claims a payback period, the CAC inputs behind it need the same definitional rigor as ARR.

Common pitfalls and how to avoid them

Trusting the field label over the calculation. A CRM field named "Closed Won Revenue" may bundle one-time fees, recurring subscription value, and implementation services — three things that must be distinct in an IPO-ready dictionary. The rule: every definition record cites formula logic, not a field name. If the entry could be satisfied by pointing at a field, it is not a definition.

Versioning chaos. Three teams maintaining three spreadsheets is the default entropy state, and the copies drift within a quarter. The fix is structural, not behavioral: one repository, pull-request review, two-of-three approval. Asking people to "keep it in sync" without a mechanism is asking them to lose.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 9

Over-defining too early. A dictionary attempting to catalog every field the sales team has ever created will not finish before the filing window closes. Freeze scope at the S-1 metric set, ship it, then extend.

Skipping the time convention. "Q4 ARR" is ambiguous until you state point-in-time versus average-over-period, and which timezone closes the quarter. Prospectuses require consistent period-over-period comparison; an inconsistent time convention produces a restatement risk that surfaces late and costs disproportionately.

Documenting without enforcing. The most common expensive failure. A well-written dictionary that no dashboard reads is a historical artifact. Pair every governed definition with a semantic-layer implementation in the same work cycle, or expect drift within one quarter.

How do you establish a cross-functional data dictionary for revenue metrics before an IPO — figure 10

Leaving manual adjustments undocumented. Nearly every company has at least one spreadsheet step between the warehouse and the reported number. Auditors will find it. Document it in the mapping — system of record, transformation steps, manual adjustments, responsible party — rather than letting it be discovered.

No version history. Maintain an appendix documenting every definition change over the trailing 24 months with date, rationale, and impact on reported numbers. This is what demonstrates governance rigor to underwriters and what prevents a last-minute restatement conversation. It costs almost nothing to maintain from day one and is painful to reconstruct retroactively.

Assuming the work ends at listing. Public-company metric definitions are disclosure. Changing one after listing without clear disclosure invites scrutiny. Carry the council forward at quarterly cadence and route material changes through your disclosure committee.

Related questions

Who should own the data dictionary?

A RevOps lead or data governance manager owns the artifact and drives the cadence. Ownership without co-signature fails — Finance and Data Engineering must approve every entry. The owner needs write access to the semantic layer, not just meeting-scheduling authority.

What if two teams genuinely disagree on a definition?

Escalate to the Metric Council, decide, and record both the decision and the rejected alternative with rationale. Some disputes are legitimate differences in purpose — bookings versus recognized revenue, for example. Define both separately and document the reconciliation rather than forcing a false merge.

Do we need a dedicated tool?

Not initially. A version-controlled repository plus a semantic layer covers most pre-IPO needs. A dedicated catalog adds discoverability and governance workflow, but does not by itself enforce calculation consistency — that remains the semantic layer's job.

How does this connect to sales compensation?

Comp plans run on bookings definitions that legitimately differ from recognized revenue. Include both in the dictionary, state the difference explicitly, and document the reconciliation. Silent divergence between comp and finance numbers is a recurring source of trust erosion.

What changes after the IPO?

Cadence relaxes to quarterly, rigor does not. Metric definition changes become disclosure events subject to SEC KPI guidance, so route material changes through the disclosure committee and consider whether prior periods need recasting.

FAQ

What exactly is a cross-functional data dictionary for revenue metrics?

A governed record defining every revenue metric — ARR, NRR, gross retention, CAC payback — with its exact calculation, source system, exclusions, time convention, and named owner, co-signed by Finance, Data Engineering, and RevOps. Its purpose is to reduce each metric to exactly one legitimate answer across the company.

How long does this take before an IPO?

Three to six months from kickoff to enforced adoption. The first two weeks scope the metric list and surface disagreements; weeks three through twelve draft and implement the top ten metrics; month four onward extends coverage and starts the audit cadence. Compressing below eight weeks typically yields a document with no enforcement layer.

Which metrics should we define first?

The ones on the cover of your board deck and in your expected S-1 disclosure set — typically 15 to 25. ARR and its decomposition, gross and net retention, customer count, CAC payback, revenue by segment and geography, deferred revenue and remaining performance obligation. Field-level cataloging of the full CRM is a separate, later project.

What causes most of the reconciliation pain?

Five recurring disputes: ramped multi-year contracts, mid-term upgrade timing, usage overage period assignment, the definition of a customer for retention math, and FX treatment. Resolve those five explicitly and most other disagreements resolve with them.

How do we stop the dictionary from drifting after we publish it?

Enforcement, not discipline. Implement each definition in a semantic layer so dashboards consume it by reference, require pull-request approval for changes, and run a quarterly trace audit walking key metrics from definition to source to reported number. Drift is a tooling problem disguised as a behavior problem.

Should auditors see the dictionary before we file?

Yes. Schedule a pre-audit trace six to nine months before your target filing date and ask the audit team to walk ten to fifteen metrics end to end. It reliably surfaces undocumented manual adjustments and timezone or period-boundary issues while there is still time to fix them cheaply.

Sources

flowchart TD S["How do you establish a cross-functiona"] S --> N0["The scenario that forces the issue"] N0 --> N1["How the mechanism actually works"] N1 --> N2["Real numbers, ranges, and benchmarks"] N2 --> N3["Trade-offs and the alternatives people"]
flowchart LR C["How do you establish a cross-functiona"] C --> H0["How the mechanism actually works"] C --> H1["Real numbers, ranges, and benchmarks"] C --> H2["Trade-offs and the alternatives people"] C --> H3["Common pitfalls and how to avoid them"]

Related on PULSE

Download:
Was this helpful?  
Sources cited
Pulse RevOps operational practicePulse RevOps operational practice