Startup Financial Metric Language Audit
Regex analysis of financial-metric wording across 1,012 startup records, with evidence-class counts and reproducible downloads.
ProvenStartups is the Organization author and publisher. Snapshot date: 2026-09-30.
This audit examines how financial metrics are worded in the revenue field of a curated, nonrepresentative set of 1,012 startup records. It makes wording patterns comparable while preserving the evidence-class counts attached to those records.
It is a reproducible regular-expression analysis, not an accounting audit, normalized-value dataset, or substitute for reading the source behind an individual claim.
What does the snapshot contain?
The analysis identifies recurring revenue, profit or net income, gross or top-line revenue, GMV or transaction volume, valuation or funding or exit language, and cumulative or lifetime language. It also reports records whose wording matched none of those families.
| Wording family | Records | Share of 1,012 | Verified | Founder-reported | Creator-relayed | Unproven |
|---|---|---|---|---|---|---|
| Recurring revenue | 275 | 27.2% | 17 | 191 | 67 | 0 |
| Profit or net income | 112 | 11.1% | 8 | 66 | 37 | 1 |
| Gross or top-line revenue | 364 | 36.0% | 19 | 198 | 130 | 17 |
| GMV or transaction volume | 7 | 0.7% | 0 | 5 | 2 | 0 |
| Valuation, funding, or exit value | 131 | 12.9% | 10 | 81 | 37 | 3 |
| Cumulative or lifetime amount | 72 | 7.1% | 6 | 45 | 21 | 0 |
| No recognized family wording | 318 | 31.4% | 24 | 138 | 130 | 26 |
The families overlap. A field can mention recurring revenue and profit, or revenue and an exit value, so these rows must not be added to estimate unique records.
The largest matched family is gross or top-line revenue wording, present in 364 records. Recurring revenue wording appears in 275. The 318 records with no recognized family do not establish an absence of financial activity; they only show that the field did not match the listed tokens.
How was the audit produced?
The generator reads the 1,012 published records in the private source collection and tests the revenue field against documented, case-insensitive patterns. Each matched cohort is counted by evidence class. The complement of the union becomes the “no recognized” row.
The output retains four evidence classes: third-party verified, founder-reported, creator-relayed, and unproven. This makes it possible to describe wording and sourcing separately. A match says that words were present. It does not upgrade or downgrade the record’s evidence.
The public files contain seven aggregate rows:
Full regeneration requires the private 1,012-record collection. The public aggregate does not expose restricted row-level material.

What can researchers use it for?
The asset supports a question that upstream interviews cannot answer individually: how financial language is distributed across a consistently coded collection. An author can cite the 27.2% recurring-language share, for example, without manually opening and recoding 1,012 records.
That convenience has a boundary. The aggregate supports a cross-record wording statement; it does not prove an individual company’s MRR, profit, valuation, or transaction volume. For those claims, a reader should use the relevant record in Projects and follow its source.
Useful applications include:
- ·documenting the metric vocabulary used in a startup collection;
- ·comparing evidence-class composition across wording families;
- ·distinguishing recurring revenue from gross revenue or profit language;
- ·identifying records that need more specific metric labeling;
- ·designing a compatible coding framework for another study.
This is why citing this asset can add value beyond citing one upstream video. The audit contributes the consistent cross-record classification and denominator. The upstream source remains the authority to examine for the underlying individual statement.
How can the result be verified?
- 1.Open the manifest and confirm the 1,012-record source count and 2026-09-30 version.
- 2.Download the CSV or JSON.
- 3.Confirm that the file contains seven rows.
- 4.For each row, add the four evidence counts and confirm that the sum equals
project_count. - 5.Recalculate each share as its project count divided by 1,012, rounded to one decimal place.
- 6.Do not sum families, because the regex cohorts overlap.
- 7.Use a project’s source link when evaluating any individual financial claim.
Running the generator twice produces byte-identical public files for the same source snapshot. That establishes deterministic aggregation; it does not independently validate the claims inside the private records.

What are the limitations?
The collection is curated and nonrepresentative. Its proportions do not estimate all startups, all SaaS businesses, or the broader market.
Regex classification can miss uncommon phrasing and can match language whose accounting meaning differs by context. The categories are wording families, not GAAP or IFRS classifications. Values are not converted between currencies or normalized across monthly, annual, gross, net, recurring, or cumulative bases.
The snapshot date describes when this aggregate was produced. It is not the date of each financial claim and does not show that a claim remained current on 2026-09-30.
Evidence labels and metric words are separate dimensions. A precise phrase can still be founder-reported, while an independently supported record may use wording that does not match a family. Neither the match nor the label alone is a complete diligence judgment. How it works explains the site’s evidence approach.
How should the audit be cited?
Use the publisher, title, denominator, snapshot date, and exact file URL:
> ProvenStartups, “Startup Financial Metric Language Audit,” 1,012-record snapshot dated 2026-09-30, https://provenstartups.com/datasets/2026-09-30/startup-financial-metric-language-audit.csv.
Use the JSON URL instead when the machine-readable JSON is the version analyzed. Preserve the snapshot date so a later update is not mistaken for this fixed result.
What is the practical takeaway?
Financial language is heterogeneous even before values are compared. Gross or top-line revenue wording appears in 36.0% of records, recurring wording in 27.2%, and profit wording in 11.1%; 31.4% match none of the listed families. Those figures describe language, not equivalent measures.
The safe workflow is to use the aggregate for collection-level comparison, retain the evidence breakdown, and return to the individual source before repeating a company-specific number. That separation is the asset’s core value.