In this piece
Products We Shipped: What Outcome Benchmarks Can Prove
The earlier version of this article gave approximate percentages for products that scaled, remained modest, were discontinued within 18 months, or never reached MVP. Wavect does not have a complete, independently audited longitudinal dataset that supports those rates. They were internal recollections, not a reproducible benchmark, and the claimed match to unnamed industry reports was unsupported. We have removed the percentages rather than present precision the evidence cannot carry.
What remains useful is the question founders were trying to answer: what happened after delivery, how was the outcome classified, and what did the engineering partner actually control? This guide provides a method for asking that question without confusing product outcomes, company survival, and delivery quality.
Building a Software Product?
Book Free ConsultationWhy are the old percentages not a benchmark?
A rate needs a defined numerator, denominator, cohort, observation window, and collection method. The old article did not publish an engagement list, a fixed observation date, documentary evidence for each status, rules for products with multiple iterations, or a treatment for missing follow-up. Its outcome bands could also overlap: revenue, fundraising, user growth, continued operation, and product discontinuation are different events.
An agency's engagements are not a random sample of startups or software products. Selection changes with geography, budget, sector, stage, referral network, and the services the agency accepts. Even a perfectly maintained internal register would describe that portfolio, not the market. Wavect experience may generate hypotheses and questions, but it should be labeled as first-party experience and not generalized into universal failure rates.
Define the cohort before calculating anything
Freeze the cohort and the rules before reviewing outcomes. Record at least:
- the inclusion rule, such as contracts signed, discoveries completed, or production releases;
- the unit of analysis, distinguishing a company, SaaS or other product, application, major iteration, and engagement;
- the first and last observation dates and the minimum follow-up required;
- the outcome event and evidence needed to classify it;
- unknown, confidential, merged, sold, pivoted, and still-observed cases;
- whether the result is a count, a point-in-time status, a time-to-event estimate, or a qualitative review.
Do not silently treat missing follow-up as survival or failure. A recently launched product has had less time to discontinue or scale than an older one. Re-run the analysis at a stated date and retain the earlier snapshot so readers can see classifications change.
Use an outcome register, not a universal ranking
The following categories are a proposed internal record structure, not observed market shares:
| Status | Minimum evidence | Important qualification |
|---|---|---|
| Stopped before production release | Dated project record and release history | Can reflect deliberate discovery, a pivot, funding, or delivery problems; it is not automatically a failed product |
| Released, later discontinued | Release evidence plus a dated owner or public confirmation | Record the observed date and stated reason without inferring a single cause |
| Operating at review date | Current production or owner evidence | Operation alone does not prove profitability, growth, user value, or product-market fit |
| Met a defined growth or impact threshold | Metric, threshold, period, source, and owner approval | Use the product's agreed objective; fundraising is not a substitute for product performance |
| Outcome unknown | Documented follow-up attempts | Keep unknowns in the denominator disclosure instead of guessing |
Terms such as "scaled," "steady," "sunset," and product-market fit need operational definitions. For example, a threshold might use retained customers, successful task completion, recurring gross margin, or an agreed public-service outcome. The right metric depends on the product.
Why company-survival statistics do not validate product rates
Official business-demography data are useful for enterprise research, but their unit and definitions differ from this article's subject. Eurostat defines enterprise births and deaths around combinations of production factors and defines survival through continued activity measured by employment, turnover, or investment. Its business-demography metadata describes those concepts and the harmonized method.
Statistics Austria reports enterprise births, deaths, active enterprises, and up to five-year survival for enterprise cohorts using registers and administrative data. Its current general business-demography page reports 2024 results and explains the population and sources. A company can survive after discontinuing one product, and a product can continue after a sale, merger, or change of operator. These statistics therefore cannot corroborate a percentage of agency-shipped software products.
How should success be measured after launch?
Define success before delivery and connect it to the problem the product is meant to solve. Measures can include task completion, retention, user satisfaction, cost to serve, reliability, accessibility, security, revenue quality, or a domain-specific outcome. Collect from more than analytics alone and state the source, owner, cadence, and decision threshold.
The UK Government Service Manual recommends combining performance metrics with user research and choosing the method according to the service and question. Its official guidance on measuring service success also points teams to user feedback, support data, and financial information rather than relying only on digital analytics. That is a sound measurement principle, not evidence for a universal commercial-product success rate.
How should causes of discontinuation be reported?
Do not rank causes from memory. Product demand, founder availability, distribution, regulation, capital, scope, team conflict, reliability, security, and delivery quality can all matter, interact, or change over time. Unless the evidence supports causal attribution, record these as reported reasons or review hypotheses.
A useful case note separates facts from interpretation: what changed, when it changed, who reported it, what data was available, and what alternative explanations remain. Customer confidentiality may limit disclosure. In that case, report the limitation instead of filling the gap with an anecdote or a precise percentage.
What can a delivery partner reasonably own?
There is no supportable universal percentage for engineering's share of a product outcome. Responsibility depends on contract, authority, access, team composition, product stage, and the failure in question. A delivery partner can usually be assessed against agreed engineering responsibilities such as discovery evidence, architecture decisions, code quality, security controls, accessibility, reliability, software-delivery discipline, release readiness, documentation, observability, handover, and response to accepted defects.
Market demand, financing, pricing, distribution, executive decisions, regulatory strategy, and continuing product ownership may sit elsewhere, or may be shared. Write a responsibility map before the work begins. If the partner also owns product discovery, growth, or operations, define the relevant outcomes and decision rights explicitly.
How should founders evaluate an agency's track record?
- Ask for the denominator. How many engagements were eligible, and which were excluded?
- Ask for dates and definitions. When were outcomes checked, and what counted as release, operation, discontinuation, or growth?
- Ask how unknowns were handled. Missing follow-up should be visible.
- Separate delivery from company outcomes. Review acceptance, incidents, maintainability, handover, and later product evidence independently.
- Check the candidate's role. A portfolio logo does not show who made the decisions or wrote the system.
- Speak to references. Ask about difficult trade-offs, post-launch ownership, defects, documentation, and transition.
- Treat curated case studies as case evidence. Published pages for Hyperstate AI, Offlinery, Scramble Pay, and LivLive describe selected work; they are not a complete outcome registry.
- Apply the same standard to alternative partner models. An agency, consultant, or fractional cofounder should distinguish verified delivery evidence from broader company outcomes.
What would Wavect need to publish a defensible portfolio rate?
We would need a versioned engagement register, stable inclusion rules, client-approved status evidence, a fixed observation date, treatment of confidential and lost-to-follow-up cases, outcome definitions set before classification, and enough context to prevent readers from treating the portfolio as representative of the wider market. Independent review would make the result stronger.
Until that exists, the honest statement is narrower: Wavect has seen products stop before release, discontinue after release, continue at modest scale, and meet material growth or impact goals. That observation is qualitative and first-party. It does not establish the probability of any outcome for a new founder.
Final thoughts
A shipped-product outcome rate is only as credible as its cohort, definitions, evidence, follow-up, and observation date. Wavect's earlier percentages did not meet that standard, so we removed them. Official enterprise-survival statistics answer a different question and cannot validate an agency portfolio.
Founders should ask delivery partners for a transparent denominator, dated classifications, unknown cases, and evidence of the work the partner actually controlled. Teams should define product success before launch, combine performance data with research and operational evidence, and preserve a responsibility map. That produces a useful review without pretending that one agency's history is a universal forecast.
