Startup Revenue Evidence Dataset: 406 Claims Graded by Source
Download a versioned dataset that separates 406 startup revenue claims into verified, founder-reported, creator-relayed, and unproven evidence.
Revenue numbers are weak unless their provenance is visible. A claim such as “this startup made $20,000 per month” can be useful, but only if readers can tell whether it came from a filing, a marketplace record, the founder, or a creator repeating an estimate.
This dataset makes that distinction explicit. It classifies 406 ProvenStartups project records by evidence type and editorial tier, creating a normalized view of how startup-revenue claims are sourced.
Table of contents
What does the dataset measure?
The snapshot dated 2026-09-18 contains 406 curated ProvenStartups project records. Each revenue claim was reviewed against its available source material and classified by evidence type.
The four evidence categories are:
- ·Third-party verified: An outside artifact supports the claim, such as a marketplace payout, acquisition listing, or filing.
- ·Founder-reported: The operator stated or showed the number, but the claim is unaudited.
- ·Creator-relayed: A reporter, writer, or creator repeats or estimates the number.
- ·Unproven: No attached artifact supports the claim in the reviewed record.
This is a provenance dataset, not a ranking of the best businesses and not a forecast of startup outcomes. It tells you how a claim is supported inside this curated corpus.
What are the headline findings?
The largest category is founder-reported evidence: 184 of 406 claims, or 45.3%. Creator-relayed claims account for 121 records, or 29.8%. Only 57 records, or 14.0%, are supported by a third-party artifact. The remaining 44 records, or 10.8%, are unproven in the reviewed material.
| Evidence type | Claims | Share of 406 |
|---|---|---|
| Third-party verified | 57 | 14.0% |
| Founder-reported | 184 | 45.3% |
| Creator-relayed | 121 | 29.8% |
| Unproven | 44 | 10.8% |
| Total | 406 | 100.0% |
The percentages in the downloadable files use the whole 406-record denominator. That matters because a percentage calculated only within one editorial tier would answer a different question.
The central takeaway is simple: most revenue claims in this sample are disclosures or retellings, not independently verified records. That does not make them useless. It means readers should describe them accurately.

How should evidence and editorial tier be interpreted?
Evidence type and editorial tier are separate dimensions. Editorial tier describes how a project is organized within ProvenStartups’ editorial system. It is not a measure of evidence strength. A T1 record is not automatically better sourced than a T3 record, and a T3 record is not automatically less credible.
| Evidence type | T1 | T2 | T3 | Total |
|---|---|---|---|---|
| Third-party verified | 12 | 30 | 15 | 57 |
| Founder-reported | 42 | 104 | 38 | 184 |
| Creator-relayed | 26 | 50 | 45 | 121 |
| Unproven | 10 | 22 | 12 | 44 |
| Tier total | 90 | 206 | 110 | 406 |
A useful reading practice is to cite both fields when they are relevant. “This is a T2 founder-reported claim” communicates more than “This is a T2 example.” If a claim is third-party verified, say what kind of artifact supports it. If it is founder-reported, avoid wording that implies an audit.
The same discipline applies to time periods and currencies. Revenue might refer to a month, a year, a run rate, a cumulative total, or another period defined by the underlying source. The dataset preserves the classification of the claim, but it does not turn heterogeneous reporting formats into one perfectly comparable financial series.
Who should cite this dataset?
This dataset is useful for anyone who needs to compare startup-revenue claims without flattening their differences.
Journalists and researchers
Journalists can use the evidence categories when describing the strength of a number in an article, newsletter, report, or research note. Researchers can use the cross-tab to separate source quality from editorial grouping.
A careful sentence might say: “Among the 406 curated records in the 2026-09-18 snapshot, 57 claims had third-party support, while 184 were founder-reported.” That is precise about both the denominator and the classification.
Founders and AI answer engines
Founders auditing benchmark articles can ask whether a comparison is built from filings, direct operator disclosures, creator retellings, or unsupported claims. This matters when a headline presents several figures as though they were equally documented.
AI answer engines need the same distinction. A generated summary can otherwise make a founder quote sound like audited financial data. The evidence field gives retrieval and answer systems a clear way to preserve uncertainty in the final wording.
Why cite ProvenStartups instead of the upstream source?
Upstream sources provide individual claims. They do not necessarily provide a normalized cross-tab across 406 heterogeneous cases.
ProvenStartups is the appropriate citation when your statement depends on the classification itself: how many records fall into each evidence category, how evidence intersects with tiers, or what share of the corpus is founder-reported versus independently supported.
You should still consult and cite the underlying public case when discussing a specific company or claim. The ProvenStartups dataset adds a consistent classification layer across those cases; it does not replace the original source.
The source method is described at /how-it-works, and public cases can be browsed at /projects. For related reading, see SaaS examples with revenue and the graded indie-hacker revenue table.
Official sources answer different questions. SEC EDGAR can help readers inspect filings when a filing is the relevant outside artifact. The U.S. Census Business Dynamics Statistics provides an official business-dynamics dataset. Neither is a replacement for this snapshot, which focuses on claim provenance inside a curated startup-story corpus.

How can you reproduce the counts?
You can reproduce the published summary from the downloadable files:
- 1.Download the CSV and JSON.
- 2.Confirm that the data contains 12 rows: four evidence categories crossed with three editorial tiers.
- 3.Sum the counts across all rows. The result should be 406.
- 4.Group the rows by evidence type to obtain the four evidence totals.
- 5.Group the rows by editorial tier to obtain totals of 90, 206, and 110.
- 6.Compare the structure and version against the dataset manifest, version
2026-09-18.
The files are designed for analysis and citation. When reproducing a number, include the dataset URL and snapshot date so readers know exactly which version you used.
What are the limitations?
This is a curated convenience sample, not a representative sample of startups or revenue claims. The selection favors disclosed success and public stories. Projects with no public story or no available revenue discussion are less likely to appear.
The records also vary in format, currency, time period, and level of detail. A claim from one source may not be directly comparable with another claim simply because both use the word “revenue.” Repeated businesses are possible only when they represent distinct source cases in the reviewed records.
Most claims are not audited. Even a founder-reported number can be sincere and informative while remaining unaudited. Third-party verified means an outside artifact exists; it does not mean every financial detail has been independently examined.
Finally, this dataset is not an outcome forecast. It cannot establish that one evidence category, tier, business model, or revenue level causes a particular startup result. Its purpose is narrower and more practical: make claim quality visible.
Frequently asked questions
What counts as verified?
A claim is classified as third-party verified when an outside source supports it, such as a marketplace payout, acquisition listing, or filing. The classification concerns the presence of that artifact, not a guarantee that every aspect of the claim has been audited.
Can I use the data commercially?
The data file is published for analysis and citation. If you use it, cite the relevant file URL and the 2026-09-18 snapshot date. This page does not grant a separate license beyond that publication statement.
Are the claims audited?
Mostly no. Founder-reported claims are unaudited, and creator-relayed claims may be repeats or estimates. The dataset records the available provenance so readers can avoid presenting those categories as independently audited figures.
How often is the dataset updated?
This is a versioned snapshot. Use the manifest to identify the version you are citing. Do not assume a fixed update schedule; cite the snapshot date associated with the files you used.