Startup Source Fetch Freshness Audit
A 2026-09-30 snapshot of fetch age across 985 structured source entries, with evidence-class cuts and reproducible downloads.
ProvenStartups is the Organization author and publisher. Snapshot date: 2026-09-30.
This audit describes how recently structured sources were collected for a curated, nonrepresentative set of 1,012 startup records. The analysis uses 985 structured source entries as its denominator.
Fetch age is the difference between a source entry’s fetched date and the snapshot date. It measures collection recency, not the claim date, publication date, event date, accuracy, or independent verification.
What does the full snapshot show?
| Fetch-age band | Source entries | Share of 985 | Unique projects |
|---|---|---|---|
| 0–30 days | 611 | 62.0% | 608 |
| 31–90 days | 213 | 21.6% | 206 |
| 91–180 days | 161 | 16.3% | 155 |
| More than 180 days | 0 | 0% | 0 |
| Missing or invalid date | 0 | 0% | 0 |
The entry counts sum to 985. Displayed shares total about 99.9% because each value is rounded to one decimal place.
Unique project counts must not be added across age bands. A project can have multiple sources collected at different times, so band-level project counts are not mutually exclusive.
What do the evidence-class cuts show?
The third-party-verified segment contains 51 source entries. Ten were fetched within 30 days (19.6%), seven within 31–90 days (13.7%), and 34 within 91–180 days (66.7%). Those entries occur across 10, seven, and 31 unique projects respectively.
The founder-reported segment contains 536 source entries. Its 0–30-day band contains 355 entries (66.2%) across 353 projects. The full downloadable file provides all five age bands for every evidence class rather than asking readers to infer missing values from the headline.
These distributions describe collection operations. The verified segment’s older fetch profile does not show that its claims are less accurate, and the founder-reported segment’s newer profile does not make its claims independently verified.

How was fetch age calculated?
For every structured source entry, the generator parses the explicit fetched date and subtracts it from 2026-09-30. It then assigns the result to 0–30, 31–90, 91–180, more than 180, or missing/invalid.
The public output contains 25 aggregate rows: five bands for the full sample and five for each of four evidence classes.
Regeneration requires the private record collection. The public output contains aggregate results rather than restricted row-level data.
Why might an author cite this asset?
Individual source pages can show when one claim was published or retrieved, but they do not establish collection-recency coverage across 985 consistently coded entries. This audit provides the fixed denominator, common date calculation, evidence-class cuts, and downloadable result.
It is useful when a methodology section needs to disclose how recently sources entered a dataset. It can also identify segments that merit source-level review. The correct next step is always to open the individual source and determine what it says, when its claim applies, and whether later evidence supersedes it.
Readers can explore the relevant records in Projects. The site’s evidence framework is described in How it works.
How can readers verify the result?
- 1.Confirm in the manifest that there are 1,012 source records and 985 structured source entries.
- 2.Open the CSV or JSON.
- 3.Confirm that the full-sample band counts are 611, 213, 161, zero, and zero.
- 4.Add
611 + 213 + 161and confirm the 985-entry denominator. - 5.For the verified segment, add
10 + 7 + 34and confirm its 51 entries. - 6.Keep source-entry and unique-project counts separate.
- 7.Treat the one-decimal share total of approximately 99.9% as rounding, not a missing entry.
The same input produces byte-identical downloads when the generator is rerun. That verifies the transformation, not the current truth of the underlying claims.

What are the limits?
The 1,012 records are curated and nonrepresentative. The age distribution cannot be generalized to all startup research, companies, or web sources.
Fetch recency is not claim freshness. A page fetched yesterday can repeat a claim from years ago. A source’s fetch date also does not establish its publication date, the period covered by a financial result, or the date on which an event occurred.
No age band evaluates accuracy, independence, accessibility, or evidentiary strength. Evidence class and fetch age are retained as separate fields because collapsing them would imply a relationship the data does not establish.
The absence of entries older than 180 days or with invalid dates applies to this fixed snapshot only. Later snapshots can change as sources are added or refreshed.
How should the audit be cited?
The citation should identify collection recency rather than imply that the claims themselves are fresh:
> ProvenStartups, “Startup Source Fetch Freshness Audit,” 985 structured source entries, snapshot dated 2026-09-30, https://provenstartups.com/datasets/2026-09-30/startup-source-fetch-freshness-audit.csv.
Use the exact JSON URL when analyzing the JSON version. Include the 985-entry denominator, because 1,012 is the number of records rather than the unit used for age shares.
What is the practical takeaway?
Within this snapshot, 62.0% of source entries were collected in the preceding 30 days and every dated entry was collected within 180 days. That is a statement about the data-collection timeline only.
The defensible workflow is to disclose this coverage, then assess claim dates and evidence at the individual-source level. Keeping collection time separate from claim time prevents a recent fetch from being mistaken for a current fact.