Startup Source-Domain Concentration: 985 Links Across 1,012 Records
Download a host-level provenance audit of 985 structured source links attached to 1,012 startup records.
YouTube accounts for 917 of 985 structured source entries—93.1%—in this 2026-09-26 collection of 1,012 records. It is the dominant destination by a wide margin: 908 projects include a YouTube source, and YouTube is the primary source for 904 projects. Bilibili follows with 51 entries, or 5.2%.
This is a provenance-concentration finding, not a judgment about source quality, independence, ownership, or accuracy. The collection is a curated convenience sample built heavily from video founder interviews and creator breakdowns; it is not startup research generally.
ProvenStartups is the Organization author and publisher.
Contents
- ·What is the headline finding?
- ·How are the structured source links distributed?
- ·What does primary source mean here?
- ·How was the concentration calculated?
- ·How can the result be reproduced?
- ·What are the limits of this snapshot?
- ·Why does this matter for secondary analysis?
- ·How should this snapshot be cited?
- ·What are the frequently asked questions?
What is the headline finding?
The structured source data is highly concentrated in a small number of destinations. YouTube alone represents 93.1% of all structured source entries. Adding Bilibili brings the two largest hosts to 968 of 985 entries, or approximately 98.3% of the total.
The distribution is therefore better understood as a narrow source-domain profile than as a broad survey of the web. Most links point to video or creator-oriented platforms, with YouTube providing the overwhelming majority of the observed source trail.
The dataset contains 1,012 records and 985 structured source entries. These are different units. A record can contain more than one structured entry, and multiple entries may belong to one project. Forty-four projects lack a structured source. That absence does not mean those projects have no prose, no other source, or no relevant evidence elsewhere.
How are the structured source links distributed?
The normalized host counts are:
| Normalized host | Structured entries | Share of 985 entries | Projects represented | Primary for projects |
|---|---|---|---|---|
| youtube.com | 917 | 93.1% | 908 | 904 |
| bilibili.com | 51 | 5.2% | 50 | 50 |
| douyin.com | 14 | 1.4% | 14 | 14 |
| elevenlabs.io | 2 | 0.2% | 1 | 0 |
| stripe.com | 1 | 0.1% | 1 | 0 |
Percentages are rounded to one decimal place, so displayed shares may not sum perfectly to 100%. The table lists destination hosts only. A shared host does not establish that links came from the same publisher, channel, organization, or independent source.
The project figures also should not be added together as though they were mutually exclusive categories. A project may have entries on more than one host. Similarly, the number of structured entries is not a count of distinct projects.

What does primary source mean here?
“Primary” means the first item in a project’s structured source array. It does not mean the strongest, most authoritative, most independent, or most accurate source.
Under that definition, YouTube is primary for 904 projects, Bilibili for 50, and Douyin for 14. ElevenLabs and Stripe each appear as structured destinations but are primary for zero projects.
This distinction matters when interpreting the dataset. Array order records how the source list is structured; it does not create a ranking of evidentiary strength. A later source may be more useful for a particular question even when it is not listed first.
How was the concentration calculated?
The calculation uses the 985 structured source entries available in the collection as of 2026-09-26.
Hosts were normalized to lowercase. www and m prefixes were removed. youtu.be was mapped to youtube.com, so YouTube links using either the short-link or main-domain format were counted together.
After normalization, each destination host was counted across the 985 structured entries. The percentage for a host is:
host entry count ÷ 985 × 100
Project counts identify how many projects are represented by entries on each host. Primary counts identify how many projects list that host’s entry first in the structured source array. These measures answer different questions and should remain separate.
The method counts destinations, not publishers or channels. It also does not infer whether two links on the same host refer to the same creator, company, interview series, or underlying evidence.
How can the result be reproduced?
The public files contain six aggregate output rows. They verify the host totals; full regeneration requires the 1,012 coded source records in the private repository.
- 1.Confirm the manifest reports 985 structured entries across 1,012 records.
- 1.With the coded source records, extract each structured source URL.
- 1.Normalize hosts by lowercasing, removing
wwwandmprefixes, and mappingyoutu.betoyoutube.com.
- 1.Group hosts, count entries and unique projects, and separately count first-array-item hosts.
- 1.Divide each host count by 985 and round percentages to one decimal place.
The manifest provides the associated dataset reference. Related context is available in the source coverage audit, startup revenue evidence dataset, method page, and projects index.

What are the limits of this snapshot?
The first limitation is scope. This is a curated convenience sample, not a representative sample of startups, startup media, founder content, or online evidence. It was built heavily from video founder interviews and creator breakdowns. The observed concentration may therefore reflect how this collection was assembled rather than the distribution of startup sources generally.
The second limitation is the unit of analysis. There are 1,012 records but 985 structured source entries, and some projects have multiple entries. Host counts are counts of destinations, not counts of unique publishers, channels, organizations, or independent reporting efforts.
The third limitation is semantic. The same host can contain materially different kinds of pages and creators. A YouTube count does not establish common ownership, common authorship, common methodology, or independence. It also does not establish source quality or accuracy.
The fourth limitation concerns missingness. Forty-four projects lack a structured source. That does not prove that no prose, transcript, document, or other source exists for those projects. It only describes the absence of a structured source in this collection.
Finally, “primary” is a positional label. It identifies the first array item, not the source that should receive the greatest evidentiary weight.
Why does this matter for secondary analysis?
A secondary analysis that treats the collection as broadly and independently sourced could overstate the diversity of its provenance. The host distribution shows that most structured source links resolve to YouTube, with Bilibili and Douyin accounting for most of the remainder.
That does not make the entries unusable. It changes what can responsibly be inferred from them. Analyses should preserve the distinction between many links and many independent sources. They should also disclose the sample’s concentration and its video-heavy construction when summarizing findings.
The practical value of this audit is that it quantifies provenance concentration before downstream analysis begins. That helps prevent a secondary analysis from assuming broad independent sourcing when the structured links are actually clustered among a few destinations.
How should this snapshot be cited?
Startup Source-Domain Concentration: 985 Links Across 1,012 Records. 985 structured source entries across 1,012 records. Dated 2026-09-26. Data files: CSV, JSON, and manifest.
What are the frequently asked questions?
Does 93.1% mean that 93.1% of projects came from YouTube?
No. It means 917 of 985 structured source entries normalized to youtube.com. The project count is 908, and a project may have multiple entries. Entry share and project share are different measures.
Does the result show that YouTube is the most reliable source?
No. The calculation measures destination concentration only. It does not evaluate reliability, quality, accuracy, independence, ownership, or publisher identity.
Why are there 985 entries but 1,012 records?
The collection contains 1,012 records and 985 structured source entries, while 44 projects lack a structured source. Records can contain multiple entries, and the record, project, and entry units are not interchangeable.
Can this be treated as research on startup sources generally?
No. The collection is a curated convenience sample built heavily from video founder interviews and creator breakdowns. It describes this collection’s structured-source profile, not startup research generally.