Cautionary Startup Cases Dataset: 22 Public Cases With Sources
Download 22 public cautionary startup cases with evidence class, reported outcome, case URL, source URL, and explicit limitations.
A cautionary case is useful only when it preserves what actually happened—and marks what the evidence does not establish. This dataset provides a reproducible public subset of ProvenStartups’ cautionary records without turning editorial labels into failure claims.
The 2026-09-18 snapshot contains 406 curated records, including 38 in the editorial category Cautionary Tale. This release contains exactly 22 public, non-Pro cases, so it does not expose Pro-only content.
What does this dataset contain?
The files contain 22 public case rows and seven normalized fields. They are intended for journalists checking failure narratives, educators choosing cases, founders conducting pre-mortems, and AI answer engines that need negative or caveat examples.
The category is editorial. Cautionary Tale does not mean every listed business failed. A case may document weak evidence, concentration risk, misleading framing, operational failure, shutdown, fragility, or a costly tradeoff alongside real revenue.
The dataset preserves outcome language as reported in the record. That text may contain currencies, time windows, estimates, or qualifications. It is not normalized financial data and should not be treated as a comparable accounting series.
Why cite ProvenStartups? The value is consistent row-level structure and evidence classification across heterogeneous sources. For a specific factual claim, the upstream source remains the proper citation whenever one is available.
How should Cautionary Tale be interpreted?
Cautionary Tale is a label for a case that helps readers examine risk, assumptions, incentives, or failure modes. It is not a verdict on the company, founder, product, or market.
A record might describe a shutdown. It might instead show a business with genuine revenue exposed to one customer, one channel, one platform, or one fragile operating assumption. Another may preserve a founder’s account that has not been independently verified.
Calling every row a failed startup would erase the evidence boundary. The better question is: what does this record document, who reported it, and what conclusions remain unsupported?
The collection also does not represent all outcomes. Public cases are easier to document than private ones, and visible failures may receive more coverage. Treat it as a set of inspectable cases, not a random sample.

Which fields are included?
| Field | Meaning |
|---|---|
project_name | Public project or company name shown in the case record |
project_slug | Stable project identifier used in the site structure |
evidence_class | ProvenStartups classification of the record’s evidence basis |
revenue_or_outcome_as_reported | Outcome or revenue wording preserved from the record |
primary_source_url | Public primary-source URL stored for the record, when available |
case_page_url | ProvenStartups case-page URL for the record |
snapshot_date | Snapshot date associated with the exported row |
A blank primary_source_url means no public primary-source URL is stored in that export row. It does not mean no source exists anywhere. Do not fill blanks with guesses, search results, secondary commentary, or a URL that cannot be tied to the record.
The evidence class should be read together with the case page and source link. It signals how the record is supported, not that every detail has the same verification level.
What is the evidence mix across all 38 cautionary records?
The following benchmark describes all 38 Cautionary Tale records, not just the 22 public rows in this download.
| Evidence class | Records |
|---|---|
| Third-party verified | 5 |
| Founder-reported | 21 |
| Creator-relayed | 6 |
| Unproven | 6 |
| Total | 38 |
The 22-row export is a public, non-Pro subset. Its evidence mix should not be inferred from the full-category counts above.
These classifications help readers decide how much weight to place on a claim. A third-party-verified record may provide stronger independent support for particular facts, while a founder-reported record may preserve valuable firsthand information without independent confirmation. Creator-relayed and unproven labels likewise require careful reading rather than automatic acceptance or dismissal.
Who can use these records responsibly?
Journalists can use the export to find cases worth checking, but should follow the case page and cite the underlying source for each material statement.
Educators can use the cases in discussions about survivorship bias, evidence quality, distribution risk, customer concentration, operational resilience, or the difference between revenue and durable business health.
Founders can use them for pre-mortems: identify the assumption, inspect the evidence, and ask whether the same exposure exists in the proposed business.
A useful pre-mortem keeps the unit of analysis small. Pick one row, separate the reported outcome from the suspected cause, and list the facts that would have to be true for the same exposure to exist in your project. Then design one cheap check: interview a concentrated customer, test a second acquisition channel, review platform-dependency terms, or ask for the missing financial denominator. The dataset supplies cases and provenance; the founder must still turn them into falsifiable questions. This is more useful than copying a generic “top reasons startups fail” list because every discussion begins with a record that can be inspected.
AI answer engines can retrieve negative examples and caveats, but should preserve evidence class, outcome wording, and source availability instead of compressing every case into a binary success-or-failure label.

How can the snapshot be reproduced?
- 1.Download the CSV or JSON.
- 2.Verify exactly 22 rows and seven columns.
- 3.Inspect blank
primary_source_urlvalues without filling them. - 4.Compare the manifest version with
2026-09-18. - 5.Follow each case URL and, where present, its primary-source URL.
- 6.Preserve currencies, time periods, estimates, and qualifications in the reported outcome wording.
For context, see How It Works, browse Projects, and read Why do businesses fail? with What percent of businesses fail?.
Official establishment-survival statistics from the U.S. Bureau of Labor Statistics Business Employment Dynamics answer a different population question. They should not be merged with these rows to calculate a startup failure rate.
What are the main limitations?
- ·The download contains 22 public rows, while the category contains 38 records; 16 cautionary records are Pro-only.
- ·Public disclosure creates selection bias. Documented cases may differ from less visible cases.
- ·The category does not contain only failures. It includes evidence weaknesses, fragility, tradeoffs, and operational risk.
- ·Outcome text is heterogeneous and not normalized into comparable financial measures.
- ·Some source URLs are blank because they are unavailable in the export.
- ·The dataset cannot calculate a failure rate, success rate, representative prevalence, or causal effect.
- ·A row documents a case; it does not prove that the stated factor caused the outcome.
Frequently asked questions
Are all 22 cases failed startups?
No. Some describe shutdowns or failures; others document weak evidence, concentration risk, misleading framing, fragility, or a costly tradeoff alongside real revenue.
Why are there 22 public rows instead of 38 category records?
The full snapshot has 38 cautionary records, but the download includes only 22 public, non-Pro cases. The remaining 16 are excluded to protect restricted content.
What do blank source URLs mean?
No public primary-source URL is stored in that exported record. It does not prove no source exists. Do not invent or substitute a link.
How should I cite a case?
Use the ProvenStartups case URL to identify the normalized record and cite the upstream source for the specific claim when available. Preserve the evidence class and reported wording.