Startup Region Disclosure Audit: What 1,012 Records Actually Say
Download a 1,012-record audit showing how often startup location is stated, missing, and distributed across the explicit subset.
ProvenStartups is the Organization author and publisher. Snapshot date: 2026-09-26
Which sections are included?
What is the answer?
The 1,012-record snapshot shows substantial geographic missingness. Only 75 records explicitly state a region or country. The remaining 937 records are marked “Not stated,” equal to 92.6% of the full set.
Among the 75 records with an explicit country label, the United States appears most often, with 29 records. That represents 38.7% of the explicit subset and 2.9% of all 1,012 records.
The central finding is therefore about disclosure, not geographic dominance. This audit documents how often location is stated before attempting any geographic comparison. It does not establish where most startups are located, which market is largest, or how the broader startup ecosystem is distributed.
“Not stated” is not a country. It is a disclosure category indicating that the record does not provide an explicit region in the audited snapshot.
What does the explicit subset show?
The explicit subset contains 75 records across 23 country labels. The United States accounts for 29 records, followed by India with 5. Japan and the Netherlands have 4 records each. Canada, Kenya, and Turkey have 3 each.
Eight country labels appear twice: Australia, France, Germany, Nigeria, South Korea, Spain, Thailand, and the UK. Eight more appear once: Argentina, Armenia, Mexico, Pakistan, Poland, Portugal, Sweden, and Vietnam.
The distribution is useful as a description of the disclosed subset. It must not be treated as a distribution of all 1,012 records. The denominator changes the meaning of every percentage: a country can represent a sizeable share of explicit disclosures while still representing a small share of the full snapshot.
| Country label or labels | Records | Share of explicit subset | Share of all records |
|---|---|---|---|
| United States | 29 | 38.7% | 2.9% |
| India | 5 | 6.7% | 0.5% |
| Japan; Netherlands | 4 each | 5.3% each | 0.4% each |
| Canada; Kenya; Turkey | 3 each | 4.0% each | 0.3% each |
| Australia; France; Germany; Nigeria; South Korea; Spain; Thailand; UK | 2 each | 2.7% each | 0.2% each |
| Argentina; Armenia; Mexico; Pakistan; Poland; Portugal; Sweden; Vietnam | 1 each | 1.3% each | 0.1% each |
The 29 United States rows do not prove United States dominance. They show that the United States is the most frequent explicit label in this curated set. That is a narrower and more defensible statement.

How should the results be read?
Read the audit as an accounting of stated information. It answers: how many records disclose a location, and what labels occur among those disclosures?
It does not answer: where are all the startups located? It also does not estimate the geographic composition of the records marked “Not stated.” Those records remain unresolved within this snapshot.
Missing location must not be inferred from a startup name, language, currency, platform, audience, or any other indirect signal. A familiar name may suggest a location, but suggestion is not explicit disclosure. The same caution applies to language, payment currency, distribution platform, customer audience, or apparent market focus.
The correct unit of interpretation is the aggregate. Individual rows should not be reclassified based on assumptions, and the explicit subset should not be presented as representative of the full collection or of the market.
How was the audit conducted?
The audit uses the 1,012-record snapshot dated 2026-09-26. Each record is treated according to whether its region is explicitly stated in the source data.
Records with an explicit country label are counted in the disclosed subset. Records without an explicit region are counted as “Not stated.” No missing location is supplied through inference.
The reported percentages use two denominators:
- ·The explicit-subset percentage is the country count divided by 75.
- ·The full-snapshot percentage is the country count divided by 1,012.
Percentages are rounded to one decimal place. Thus, the United States is 29 out of 75, or 38.7%, and 29 out of 1,012, or 2.9%. The same approach produces 92.6% for the 937 records marked “Not stated.”
This method preserves the distinction between disclosure frequency and geographic prevalence. It also makes the result easier to inspect: readers can see both the number of explicit records and the much larger number for which no region is stated.
How can readers verify the result?
The public downloads contain the 24 aggregate output rows. They verify the published counts; full regeneration requires the 1,012 coded source records in the private repository.
- 1.Open the CSV dataset or JSON dataset.
- 1.Confirm in the manifest that the snapshot has 1,012 records and is dated 2026-09-26.
- 1.Confirm that the output totals 75 explicit records plus 937 “Not stated” records.
- 1.Confirm that 23 explicit country labels match the displayed counts.
- 1.Recalculate each explicit-subset share using 75 and each full-snapshot share using 1,012.
- 1.Preserve both denominators and cite the exact file URL, publisher, record count, and date.
The Projects page provides the broader project context, while How it works provides additional context for the underlying collection.

What are the limits?
This is a curated, nonrepresentative snapshot. It should not be generalized to all startups, all records outside the dataset, or the startup market as a whole.
The audit measures what is explicitly disclosed in the records available for this snapshot. It does not measure the true location of every organization represented. The 937 “Not stated” records may contain startups from many regions, but this audit does not assign them countries.
The 75 explicit records should not be used as a proxy for the 1,012-record population. Likewise, the country ranking within the explicit subset should not be interpreted as a market ranking.
Finally, the percentages are descriptive and rounded. Small differences may arise if readers calculate with unrounded values or apply a different denominator. Reproduction should therefore retain both raw counts and the denominator used.
How should this audit be cited?
Cite the publisher, title, record count, snapshot date, and exact file URL. A concise citation can follow this form:
> ProvenStartups, “Startup Region Disclosure Audit: What 1,012 Records Actually Say,” 1,012-record snapshot dated 2026-09-26, CSV.
For analyses based on the machine-readable version, cite the JSON file instead. The manifest can be included when documenting the snapshot identity or reproducibility details.
Citation should make clear that the source supports disclosure statistics. It should not imply that the audit establishes the location of all 1,012 records or proves market-wide geographic shares.
What is the practical takeaway?
Start with missingness. Before comparing countries, regions, or markets, establish how many records provide an explicit geographic label.
Here, the answer is 75 of 1,012. That means geographic comparison applies directly to only 7.4% of the snapshot, while 92.6% remains “Not stated.” The most responsible headline is therefore that location disclosure is sparse.
The United States is the leading explicit label in the available subset, with 29 records, but that is not evidence of United States dominance across the full set. It is evidence that the United States is the most frequently disclosed country among the records that disclose one.
For downstream analysis, preserve the two denominators, keep “Not stated” visible, and avoid filling gaps with assumptions. The unique value of this audit is that it documents missingness before geographic comparison.
What are the four frequently asked questions?
Is “Not stated” the same as an unknown country?
Yes, in the limited sense that the country is not explicitly provided in the audited record. It is not a country label and should remain a separate disclosure category rather than being assigned to any geography.
Does the United States have the largest share of all 1,012 records?
No such conclusion is supported. The United States has 29 records, equal to 38.7% of the 75 explicit records and 2.9% of all 1,012 records. The audit does not classify the 937 records marked “Not stated.”
Can location be inferred from language, currency, or audience?
No. Missing location must not be inferred from a name, language, currency, platform, audience, or similar indirect signal. Only explicit region disclosure belongs in the geographic counts.
What is the safest way to use these results?
Use them as aggregate disclosure statistics from a curated, nonrepresentative snapshot. Report the raw counts, identify the denominator, preserve the “Not stated” category, and cite the exact ProvenStartups file URL and snapshot date.