Blog / App Development
· App Development

Why software projects fail, according to the data rather than the anecdotes

Themba Mahlangu · 6 min read

Explanations for software project failure are usually anecdotes with a moral attached. The measured picture is more useful and points somewhere specific.

The Standish Group's CHAOS research reports 31 percent of software projects successful, 50 percent challenged, meaning late, over budget or delivered with less than promised, and 19 percent failed outright. Underneath the headline is a much larger effect: small projects succeeded around 90 percent of the time, while large projects succeeded less than 10 percent of the time.

That gap is bigger than the difference produced by any process choice measured in the same research, and it is a variable you control directly.

Size is the dominant variable

A ninefold difference in success odds, driven by project size alone, should reorder how these decisions are made. It suggests that the most valuable thing you can do to protect a software investment is to make each committed piece smaller, before choosing a methodology, a technology or a supplier.

The mechanism is not mysterious. A larger project has more requirements to be wrong about, more people whose availability has to hold, more time for the business to change underneath it, and more distance between a decision and the evidence that it was wrong. Each of those compounds rather than adds.

The practical version is that a programme of five sequential small projects, each producing something usable, has a materially different risk profile from one project of the same total size, even with the same people building the same thing.

The three named success factors

The same research names user involvement, executive support and a clear statement of requirements as the leading factors in project success. All three are on the client side, and none of them are technical.

User involvement fails quietly. Projects specified by managers describing how the work should happen, rather than by the people doing it, produce software that models an idealized process. The gap surfaces at launch, when the people expected to use it explain what actually happens.

Executive support matters most for the unglamorous decisions: releasing people's time for review, granting access to systems, and choosing when two departments disagree. Its absence looks like a project that is technically on track and cannot get decisions.

A clear statement of requirements is the one most often claimed and least often present. The test is whether the document says what is excluded. A requirements document with no boundary has not been finished, it has been stopped.

Where AI projects fail differently

AI projects add a failure mode of their own, and it has been measured. MIT's Project NANDA studied more than 300 enterprise deployments and found that 95 percent of enterprise generative AI pilots produced no measurable financial return, attributing the cause to integration and organizational learning rather than to model capability.

McKinsey's State of AI survey shows the same shape from the adoption side: 88 percent of organizations use AI in at least one function, around a third have begun scaling it, 7 percent describe it as fully scaled, and 39 percent can attribute any enterprise-level EBIT impact to it, with most of those below 5 percent.

The specific failure is a pilot arranged to avoid the difficult data work. It succeeds on that basis, and then cannot be scaled because the difficult work was never done. The model was rarely the constraint.

The failures that show up in practice

The estimate priced a different project. Two quotes for the same brief differing by a factor of three almost always means the firms assumed different things. The assumptions were never compared, and the gap surfaces as change requests.

Nobody drew the data model. Cost and risk live in the entities and their relationships. A project scoped in screens discovers the real complexity during the build.

The first deployment happened late. Environment, access, permission and integration problems exist in every project. Meeting them in the final week is what turns a small delay into a large one.

Data migration was assumed to be simple. Real systems accumulate records that violate their own rules. Nobody knows how many until someone counts, and the count changes the plan more than any other single activity.

A second user role appeared mid-project. Roles multiply against every screen, endpoint and report. This is the most common source of a build quietly doubling.

Nothing was ever cut. A scope that only grew was never really managed. The absence of anything removed is a reliable indicator that the project is heading for the challenged category.

What actually reduces the risk

Make each commitment small enough that being wrong is survivable, and sequence them so each produces something usable. Get software into a real environment in the first week so the environmental problems arrive early. Involve the people who will use it, not only the people who commissioned it. Insist that the scope document says what is excluded. Look at the real data before pricing anything that touches it.

None of that is novel and all of it is regularly skipped, usually under time pressure created by the desire to start building.

Our own structure follows from this. Discovery runs five to ten days and ends with a statement of work carrying fixed pricing, assumptions, exclusions and risks. Builds run two to four weeks with staging in the first week. Each phase is separately priced, so each commitment stays in the range where the odds are good.

Frequently asked questions

What percentage of software projects fail?

Standish's CHAOS research puts it at 19 percent failing outright and 50 percent challenged, meaning late, over budget or reduced in scope. Only 31 percent are fully successful. Success rates vary sharply with project size.

Does agile reduce failure?

Less than reducing project size does, on the same data. Both agile and traditional approaches work when the scope is small and users are involved, and neither rescues a project that is too large.

Why do AI projects fail more often?

MIT's 2025 research found 95 percent of enterprise generative AI pilots produced no measurable return, caused by integration and organizational factors rather than model quality. The common pattern is a pilot that avoids the hard data access and therefore cannot be scaled.

What is the single best predictor of success?

Project size, by a wide margin, followed by whether the people who will use the software were genuinely involved in defining it.

Can a failing project be recovered?

Usually, by cutting it into something small enough to finish. That normally means abandoning part of the planned scope and shipping a narrower version, which is politically difficult and technically the straightforward answer.

Where to start

If a project is being planned and looks large, the useful step is cutting it into pieces that can each succeed on their own. That is what our scoping produces, starting at $3,000 and credited toward the build.

Book an AI audit.