Startup Source Coverage Audit: 406 Claims Checked for Traceability
Download a source-coverage audit of 406 startup records, grouped by evidence class and attached structured-source count.
ProvenStartups is the Organization author and publisher of this audit. The snapshot covers 406 startup claims and measures how many structured sources are attached to each claim. The main result is concentrated coverage: 350 claims have one attached source, 44 have none, and 12 have two or more.
This is a descriptive convenience sample, not a success forecast. Source presence does not establish that a claim is true, that a source is high quality, or that a startup will succeed.
Where should you start?
What does the audit measure?
The audit checks traceability at the claim level. Each row has an evidence label and a structured sources array. The source count is the length of that array.
The four evidence labels are:
- ·Verified: 57 claims
- ·Founder: 184 claims
- ·Creator: 121 claims
- ·Unproven: 44 claims
The source-count buckets are zero sources, one source, and two or more sources. Prose links are excluded from this count. A claim with zero sources therefore has no source attached structurally in the audited record; it does not necessarily mean that no relevant information exists anywhere.
This distinction matters because source presence is not the same as corroboration. One attached source may be useful, but the audit does not assess whether it independently confirms a claim. It also does not rank source quality.
How is source coverage distributed?
Across the full snapshot, the distribution is:
| Evidence label | Claims | 0 sources | 1 source | 2+ sources |
|---|---|---|---|---|
| Verified | 57 | 17 (29.8%) | 38 (66.7%) | 2 (3.5%) |
| Founder | 184 | 12 (6.5%) | 164 (89.1%) | 8 (4.3%) |
| Creator | 121 | 6 (5.0%) | 113 (93.4%) | 2 (1.7%) |
| Unproven | 44 | 9 (20.5%) | 35 (79.5%) | 0 (0%) |
| Overall | 406 | 44 | 350 | 12 |
The overall pattern is clear: most claims have exactly one structured source attached. Two-plus-source coverage is uncommon, representing 12 of 406 claims. Zero-source rows account for 44 claims.
The evidence labels show a second pattern. Founder and creator claims have the highest share of one-source attachment: 89.1% and 93.4%, respectively. Verified claims are more distributed, with 29.8% having zero sources and 3.5% having two or more. Unproven claims also have a relatively large zero-source share at 20.5%, while none has two or more attached sources.
These figures describe how the dataset is documented. They do not determine whether one evidence label is more accurate than another.

What is the evidence-label/source-attachment gap?
Evidence labels and source attachment are related fields, but they answer different questions.
The evidence label classifies the claim according to the audit’s editorial categories. The source count records how many items are present in the structured source array. Neither field substitutes for the other.
The gap is visible in both directions. Among 57 verified claims, 17 have zero attached structured sources. At the same time, 35 of the 44 unproven claims have one attached source. That means a label should not be read as a direct proxy for the number of sources attached to a row.
A source count of one also should not be read as corroboration. It indicates one structured attachment, not two independent confirmations. Likewise, a zero count indicates a structural absence in this snapshot, not proof that the claim has no supporting material outside the audited field.
For readers using the dataset, the safest interpretation is to examine the evidence label and the attached sources together, then apply the stated limitations.
How was the audit performed?
The audit used the 2026-09-22 snapshot with a total sample size of 406 rows.
For each row, the audit retained its evidence label and counted the entries in its structured sources array. The count was then assigned to one of three length buckets:
- 1.Zero sources
- 2.One source
- 3.Two or more sources
The grouped results were calculated separately for each evidence label and for the full sample. The overall totals reconcile to 406 claims:
- ·44 rows in the zero-source bucket
- ·350 rows in the one-source bucket
- ·12 rows in the two-plus-source bucket
The label totals also reconcile to the same snapshot:
- ·57 verified
- ·184 founder
- ·121 creator
- ·44 unproven
This method is intentionally narrow. It measures structured source attachment and does not attempt to judge the underlying claim, review the full provenance of a source, or infer whether multiple sources are independent.
How can readers verify the results?
Reproduction requires the snapshot files and the manifest. Download the exact machine-readable files here:
The public files let readers verify every published aggregate row and its snapshot metadata. Full regeneration requires the 406 coded source records: take each record's evidence label and structured sources length, place the length into the zero, one, or two-plus bucket, then group by evidence label.
Do not count ordinary prose links unless they are also represented in the structured sources array. That exclusion is necessary to match the audit definition.
Readers can also review the broader Projects area and the How it works explanation for surrounding context. The related asset is the Startup revenue evidence dataset.

What are the audit’s limits?
This is a curated sample rather than a census of startup claims. Its results describe the selected rows in this snapshot and should not be generalized to all startups, all evidence practices, or all claims.
The evidence classes are editorial categories. They are useful for organizing the audit, but they are not independent scientific measurements or guarantees of accuracy.
The snapshot is mutable. Future versions may change rows, labels, source attachments, or totals. Readers should cite the dated files and manifest used for their analysis.
The audit also does not publish a restricted row-level export. The downloadable CSV and JSON contain aggregate results, so they support output verification but not independent row-level regeneration.
Finally, traceability is not the same as truth, quality, or business performance. A claim can have an attached source that is incomplete or weak, while a claim with no structured attachment may have context elsewhere. Nothing in this audit predicts startup success, revenue, durability, or investment outcomes.
What should readers cite?
For a reproducible citation, reference the dated snapshot and identify the format used. The CSV is convenient for tabular analysis, while the JSON preserves the structured record format. The manifest provides the snapshot-level reference.
Use the exact links rather than substituting an undated page:
- ·Data table: startup-source-coverage-audit.csv
- ·Structured records: startup-source-coverage-audit.json
- ·Snapshot reference: manifest.json
The appropriate claim is that, in the 2026-09-22 snapshot of 406 curated claims, 44 had zero structured sources, 350 had one, and 12 had two or more. Broader conclusions require additional data and independent validation.
What do readers ask most often?
Does one source mean a claim is corroborated?
No. One source means that one item appears in the row’s structured sources array. The audit does not assess independence, quality, completeness, or corroboration.
Does zero sources mean the claim has no evidence?
No. Zero means that no source is attached structurally in the audited record. It does not establish that no relevant material exists elsewhere.
Is this audit a prediction of startup success?
No. It is a descriptive traceability audit of a curated convenience sample and is not a success forecast.