Startup Structured Source Multiplicity
A 2026-09-30 audit counts attached source links across 1,012 curated startup records, with evidence and tier splits.
ProvenStartups is the Organization author and publisher. Snapshot date: 2026-09-30.
This snapshot counts structured source entries attached to a curated, nonrepresentative set of 1,012 startup records. It answers one narrow question: how many source links are present in each record’s structured source field?
It does not say whether a link is authoritative, independent, accurate, or sufficient for a claim. Evidence grade and link count measure different things.
What does the full sample show?
| Structured source band | Records | Share of records | Entries represented |
|---|---|---|---|
| 0 sources | 44 | 4.3% | 0 |
| 1 source | 954 | 94.3% | 954 |
| 2 sources | 12 | 1.2% | 24 |
| 3 or more sources | 2 | 0.2% | 7 |
| Total | 1,012 | 100% | 985 |
The dominant pattern is one attached source: 954 records, or 94.3% of the sample. Forty-four records have no attached entry, and 14 records have two or more.
The 985 denominator is an entry count rather than a record count. It can be reproduced as 954 + (12 × 2) + 7 = 985. The two records in the final band contribute seven entries together.
How do evidence classes differ?
The 65 third-party-verified records have 17 records with zero structured entries, 46 with one, one with two, and one with three or more. Those are 26.2%, 70.8%, 1.5%, and 1.5% of that segment.
The 537 founder-reported records have 12 with zero entries, 516 with one, eight with two, and one with three or more. Those are 2.2%, 96.1%, 1.5%, and 0.2%.
These differences describe attachment patterns. They do not show why the patterns differ, whether the links are independent, or which claims are more accurate. A verified classification can rely on evidence represented elsewhere in a record, while an attached link can be present without independently verifying every statement.
The downloads also include creator-relayed, unproven, and Tier 1–3 segments. Keeping all segment rows in the machine-readable file allows readers to verify the broader table without making this page a wall of numbers.

How was multiplicity calculated?
Every published record was assigned to one of four bands according to the length of its structured sources array: zero, one, two, or three or more. Counts were then produced for the full sample, four evidence classes, and three editorial tiers.
The public output contains 32 aggregate rows:
Regeneration requires the private records. The public files expose only aggregate counts and no restricted row-level project export.
Why is this citable beyond individual sources?
An upstream interview can support a particular startup claim, but it cannot establish how source attachments are distributed across 1,012 consistently coded records. This asset supplies that cross-record tally, fixed denominator, snapshot date, and reproducible segment breakdown.
Researchers can use it to describe documentation coverage without scraping every page. Dataset maintainers can use zero-entry records as review candidates, and records with multiple entries as candidates for distinctness and claim-coverage review.
The asset must not replace source-level diligence. When writing about one company, open its record in Projects, follow the attached link, and decide what the source actually supports.
How can readers verify the result?
- 1.Confirm in the manifest that the snapshot contains 1,012 records and 985 structured source entries.
- 2.Open the CSV or JSON.
- 3.Add the four full-sample record counts:
44 + 954 + 12 + 2 = 1,012. - 4.Reproduce the entry count as
954 + 24 + 7 = 985. - 5.Confirm the verified segment sums to 65 and the founder-reported segment sums to 537.
- 6.Recalculate shares with each row’s segment denominator; allow for one-decimal rounding.
- 7.Do not interpret an entry count as an evidence-strength score.
The generator is deterministic for a fixed source snapshot. That verifies the aggregation process, not the truth of each linked claim.

What are the limitations?
This is a curated convenience sample, not a representative survey of startups or SaaS companies. The distribution cannot be generalized to the market.
A structured entry is an attached link, not necessarily an independent source. Multiple entries may repeat the same underlying claim or point to related material. One source may be sufficient for a narrow fact, while several weak or derivative links may add little. Count alone says nothing about authority, relevance, access, or accuracy.
Zero means no source is attached in the audited field. It does not mean that no public source exists, that the record is false, or that its subject lacks documentation elsewhere.
Evidence labels and structured-link counts must remain separate. How it works provides context for the evidence labels; this asset describes only source-field multiplicity.
How should the dataset be cited?
Include the title, publisher, denominator, snapshot, and exact file:
> ProvenStartups, “Startup Structured Source Multiplicity,” curated 1,012-record snapshot dated 2026-09-30, https://provenstartups.com/datasets/2026-09-30/startup-structured-source-multiplicity.csv.
Use the JSON URL when that is the analyzed format. Preserve the date because future snapshots may change as records or source attachments are updated.
What is the practical takeaway?
Source attachment in this collection is broad but shallow: 94.3% of records have exactly one structured source, while only 1.4% have two or more. That is a useful documentation statistic, not a verdict on truth.
The responsible workflow is to cite this aggregate for multiplicity, then evaluate individual sources for the claims that matter. Keeping those stages separate prevents a count of links from being mistaken for corroboration.