AI Startup Difficulty Benchmark: Five Scores Across 13 Models
Download five-axis difficulty scores for 406 curated startup cases, grouped across 13 models with explicit scoring limits.
The easiest category to build is not automatically the easiest category to sell. In this 406-project snapshot, Digital Publishing has the lowest all-five mean at 2.23, yet its competition score is 3.55. Scale Reference is hardest overall at 3.88, with capital at 4.28.
These are editorial comparison scores, not forecasts of startup success, measured costs, or outcome probabilities. They help founders compare burdens before deciding what to build, validate, or study.
What will you find on this page?
What does the benchmark measure?
ProvenStartups coded 406 curated project records in a snapshot dated 2026-09-18. Each record receives five editorial scores:
- ·Tech: implementation and technical complexity.
- ·Acquisition: difficulty of reaching and converting users.
- ·Capital: expected funding, infrastructure, or operating-resource burden.
- ·Competition: intensity of competing products and alternatives.
- ·Validation: difficulty of testing demand credibly.
Every score uses a 1–5 scale, where 1 means easier and 5 means harder. The scores are normalized editorial comparisons across source cases. They are not direct measurements of dollars, engineering hours, customer counts, or probability of success.
Category rows use arithmetic means. There is no weighting beyond one record per project. The displayed all-five category mean is the mean of the five displayed category means.

Which categories look easier or harder?
| Category | Count | Tech | Acquire | Capital | Compete | Validate | Mean |
|---|---|---|---|---|---|---|---|
| AI Content | 25 | 1.84 | 2.84 | 1.88 | 3.40 | 2.60 | 2.51 |
| AI E-commerce | 10 | 1.90 | 2.70 | 1.90 | 3.50 | 2.30 | 2.46 |
| AI Service | 48 | 2.15 | 2.98 | 2.04 | 3.04 | 2.33 | 2.51 |
| AI Website | 15 | 2.33 | 3.07 | 2.67 | 3.47 | 3.20 | 2.95 |
| Consumer App | 60 | 2.48 | 3.18 | 2.20 | 3.55 | 2.63 | 2.81 |
| Digital Publishing | 31 | 1.19 | 2.97 | 1.13 | 3.55 | 2.32 | 2.23 |
| Directory Site | 12 | 2.17 | 3.25 | 2.08 | 2.83 | 4.00 | 2.87 |
| Ecosystem Tool | 9 | 2.67 | 2.89 | 1.78 | 3.11 | 2.44 | 2.58 |
| Platform Plugin | 14 | 2.43 | 2.64 | 1.64 | 3.14 | 2.29 | 2.43 |
| SaaS | 93 | 2.94 | 3.01 | 2.20 | 3.49 | 2.83 | 2.89 |
| Scale Reference | 36 | 3.92 | 3.47 | 4.28 | 4.00 | 3.72 | 3.88 |
| Simple Tool | 15 | 2.00 | 2.80 | 1.60 | 3.27 | 2.80 | 2.49 |
| Cautionary Tale | 38 | 2.63 | 3.21 | 2.26 | 3.82 | 2.71 | 2.93 |
Digital Publishing ranks lowest by all-five mean at 2.23, followed by Platform Plugin at 2.43 and AI E-commerce at 2.46. That describes relative burden under this rubric. It does not mean those categories are more likely to succeed.
Scale Reference ranks highest at 3.88. Its capital score of 4.28 is also the highest displayed score. The projects grouped there carry heavier resource demands across the coded cases, especially around capital and technical execution.
Some categories show why a single average misleads. Directory Site has a relatively low tech score of 2.17 but a validation score of 4.00. Building the directory may be manageable; proving that it will become useful, trusted, and habit-forming can be much harder.
SaaS shows a balanced burden: tech is 2.94, acquisition is 3.01, and competition is 3.49. A SaaS product may be straightforward to prototype while still requiring sustained distribution and differentiation.
Why does easy-to-code still mean hard-to-sell?
A low tech score answers only one question: how difficult does implementation appear relative to the other coded projects? It does not answer whether people will notice the product, trust it, pay for it, keep using it, or recommend it.
That distinction matters because distribution often becomes the binding constraint. A founder can launch a functional tool quickly and still face expensive acquisition, crowded alternatives, unclear positioning, or weak evidence that the problem is urgent. Digital Publishing illustrates the pattern: very low tech and capital scores coexist with a 3.55 competition score.
The practical reading is not “choose the lowest-scoring category.” Identify the constraint most likely to dominate your situation:
- ·If engineering capacity is limited, inspect tech.
- ·If you already have an audience, acquisition may be less burdensome for you than the category mean.
- ·If you are bootstrapping, capital deserves extra weight.
- ·If incumbents are strong, competition may dominate the plan.
- ·If buyer pain is uncertain, validation should come before extensive implementation.
These scores are starting points for questions, not substitutes for customer conversations or experiments.
Who should cite this benchmark?
Founders can cite it when comparing execution burdens across several product directions. A founder evaluating an AI service against a platform plugin can use the five axes to structure a pre-build discussion.
Researchers can cite it when discussing how technical difficulty differs from acquisition, capital, competition, and validation burdens in startup examples. Writers can use category-level scores when explaining why a project may be easy to prototype but difficult to distribute or validate.
Why cite ProvenStartups? The organization applied a consistent five-axis coding scheme across 406 source cases. Individual interviews and project descriptions generally do not provide this cross-case benchmark. The value is normalization and aggregation, with limits disclosed.
For context, see How It Works, browse Projects, and compare no-code SaaS examples with one-person company guidance.

How can you reproduce the category rows?
- 1.Download the CSV or JSON.
- 2.Confirm that the 13 category counts sum to 406.
- 3.Check that every axis uses 1 for easier and 5 for harder.
- 4.Recompute each category mean from underlying records when available.
- 5.Recompute the all-five value as the arithmetic mean of the five displayed category means.
- 6.Compare the version with the dataset manifest.
The results should match after normal rounding. Reproducing the arithmetic does not turn editorial scores into measured market facts.
What are the limitations?
The rubric is subjective. Two careful reviewers could assign different scores to the same project, especially where category boundaries or operating assumptions are unclear.
Category breadth varies. SaaS, AI Service, and Cautionary Tale can contain materially different business designs. A category mean compresses that variation into one number. Means also hide spread: 2.50 could describe tightly clustered records or a mixture of easy and difficult cases.
The curated sample is not representative of the full market. Scores can also age as tools, infrastructure prices, platform policies, buyer expectations, and competitive landscapes change. The snapshot date is part of the result.
For broader planning context, the SBA business guide is a general reference. It does not validate any ProvenStartups score.
Frequently asked questions
What does the 1–5 scale mean?
One means easier relative to the other coded projects, while five means harder. It is an editorial comparison, not a percentage, dollar estimate, or probability.
Which category is easiest?
Within this snapshot and rubric, Digital Publishing has the lowest all-five mean at 2.23. Its 3.55 competition score shows why that answer needs interpretation.
Can these scores predict startup success?
No. They describe comparative burdens in curated records and establish neither causation nor a probability of success.
How should I cite this benchmark?
Cite ProvenStartups, the title, snapshot date 2026-09-18, and the relevant CSV, JSON, or manifest. State that the figures are editorial category means across 406 curated records.