Startup Difficulty Distribution: 406 Projects by Evidence and Tier
Download quartiles and high-difficulty shares for 406 startup cases, segmented by evidence class and editorial tier.
Published by ProvenStartups
The 2026-09-22 snapshot shows a broad middle range of startup difficulty: across 406 projects, the composite mean is 2.80 on a 1-to-5 scale, the median is 2.80, and 17.0% score at or above 3.5. Tier 3 is the clearest high-difficulty group, with a mean of 3.48 and 49.1% of projects at or above 3.5. Tier 1 is the clearest low-difficulty group, with a mean of 2.22 and no projects at or above 3.5.
These results describe this dataset. They do not predict which projects will succeed. This is a descriptive convenience sample, not a success forecast.
Where should you start?
What does the distribution show?
The overall distribution is centered near the midpoint of the scale rather than at either extreme. The all-projects segment has a 25th percentile of 2.20, a median of 2.80, and a 75th percentile of 3.20. That means the central half of the snapshot falls between 2.20 and 3.20.
The segment averages separate in useful but limited ways:
| Segment | n | Mean | P25 | Median | P75 | Share ≥3.5 |
|---|---|---|---|---|---|---|
| All projects | 406 | 2.80 | 2.20 | 2.80 | 3.20 | 17.0% |
| Verified | 57 | 3.13 | 2.60 | 3.00 | 3.80 | 33.3% |
| Founder | 184 | 2.83 | 2.35 | 2.80 | 3.25 | 18.5% |
| Creator | 121 | 2.66 | 2.20 | 2.60 | 3.00 | 9.1% |
| Unproven | 44 | 2.67 | 2.20 | 2.50 | 3.05 | 11.4% |
| Tier 1 | 90 | 2.22 | 2.00 | 2.20 | 2.40 | 0% |
| Tier 2 | 206 | 2.70 | 2.20 | 2.60 | 3.00 | 7.3% |
| Tier 3 | 110 | 3.48 | 3.00 | 3.40 | 3.95 | 49.1% |
The tier pattern is the strongest separation in the snapshot. Tier 1 has a narrow central range from 2.00 to 2.40. Tier 2 spans a higher range, with a median of 2.60 and a 75th percentile of 3.00. Tier 3 is higher still: its middle half runs from 3.00 to 3.95.
The evidence segments overlap more substantially. Verified projects have the highest mean among the named evidence groups at 3.13, while founder projects have a mean of 2.83, unproven projects 2.67, and creator projects 2.66.
Is verified evidence associated with higher difficulty?
In this snapshot, verified projects appear harder on the composite scale. Their mean is 3.13, compared with 2.80 across all 406 projects. One-third of verified projects, 33.3%, meet or exceed the 3.5 threshold, versus 17.0% overall.
That association should not be interpreted as evidence causing difficulty. The table is descriptive: it reports how the labels and scores occur together in this sample. It does not establish a causal relationship, and it does not show that verification makes a project harder.
The verified segment also has a wider upper range than several other groups. Its 75th percentile is 3.80, compared with 3.25 for founder projects, 3.00 for creator projects, and 3.05 for unproven projects. Those differences are part of the snapshot’s observed structure, not a general rule about evidence.

How do the tiers differ?
Tier 1 is the lowest-scoring group in the snapshot. Its mean is 2.22, its median is 2.20, and its 75th percentile is 2.40. No Tier 1 project reaches the 3.5 threshold.
Tier 2 is the largest tier, with 206 projects. Its mean is 2.70, median 2.60, and 7.3% share at or above 3.5. It sits between Tier 1 and Tier 3 on every reported summary measure.
Tier 3 contains 110 projects and has the highest mean, median, upper quartile, and threshold share. Its mean is 3.48, its median is 3.40, and its 75th percentile is 3.95. Nearly half of its projects, 49.1%, reach or exceed 3.5.
The tier separation is partly editorial and definitional. The tiers are not independent physical measurements of difficulty. They are groupings used in the dataset, and the relationship between tier labels and the composite score reflects how those labels and scoring rules were defined.
What is the difficulty score measuring?
Each project receives five editorial scores:
- ·technology difficulty
- ·acquisition difficulty
- ·capital difficulty
- ·competition difficulty
- ·validation difficulty
Each component uses an ordinal scale from 1, easier, to 5, harder. The composite score is the arithmetic mean of those five component scores.
For example, the composite is not a separate observed outcome. It is a summary calculated from the five editorial dimensions. A higher composite means that the project received higher average ratings across those dimensions.
The scale should therefore be read comparatively within this snapshot. A score of 3.20 is not a probability, time estimate, revenue forecast, or guarantee of execution effort. It is the arithmetic average of five ordinal editorial ratings.
How can the results be verified?
The downloadable files contain the eight published summary rows:
Full regeneration requires the 406 coded source records. The procedure is:
- 1.Start with the 406 project records in the 2026-09-22 snapshot.
- 2.For each project, average its five component scores: technology, acquisition, capital, competition, and validation.
- 3.Assign each project to the relevant evidence segment and tier using the labels in the dataset.
- 4.For each segment, count the records and calculate the arithmetic mean of the project-level composites.
- 5.Sort the composite scores in ascending order.
- 6.Calculate quartiles using linear interpolation at position
(n - 1)q, whereqis 0.25, 0.50, or 0.75. - 7.Count the projects with composite scores greater than or equal to 3.5, then divide by the segment size.
The median is the q = 0.50 case of the same quartile procedure. The threshold share is reported as a percentage of the segment, not as a separate score.

What are the main limitations?
The first limitation is the sample itself. This is a convenience sample of 406 projects, so its composition may not represent all startup projects or all possible startup ideas.
The second limitation is measurement. The five inputs are editorial, ordinal scores. Averaging ordinal subjective scores produces a useful compact index for comparison, but it does not make the underlying judgments objective or interval-scaled.
The third limitation is interpretation. A higher score indicates greater rated difficulty across the selected dimensions. It does not indicate a lower chance of success, a longer time to launch, a larger funding requirement, or a particular business outcome.
The fourth limitation is grouping. Evidence labels and tiers summarize how projects were categorized for this dataset. The observed differences between groups may reflect editorial or definitional choices, as well as differences among the projects themselves.
Finally, the snapshot is cross-sectional. It describes the projects and labels available on 2026-09-22. It does not track how scores, evidence, tiers, or outcomes change over time.
The public CSV and JSON verify the published output, but they do not contain a restricted project-level export. They therefore cannot independently regenerate the distributions without the underlying coded records.
Where can I explore related material?
Use ProvenStartups Projects to browse the project collection and How It Works for the broader explanation of the scoring and presentation approach.
For a focused comparison, read the related asset: AI Startup Difficulty Benchmark.
FAQs
Is a higher score better?
No. A higher score means higher rated difficulty on the 1-to-5 composite scale. It is not a quality ranking or success prediction.
What does “verified” mean here?
“Verified” is an evidence segment label used in this snapshot. The distribution shows that verified projects have a higher average score, but it does not show that verification causes difficulty.
Why are the quartiles reproducible?
The quartiles use sorted composite scores and linear interpolation at (n - 1)q, with q values of 0.25, 0.50, and 0.75.
Can this dataset forecast startup success?
No. It is a descriptive convenience sample and should not be treated as a success forecast.