Engage
← All insights
Insight · AI capability building

The AI Factory: a disciplined route from AI ambition to working software.

At a glance
  • In 2025, 42 percent of organisations abandoned most of their AI initiatives, up from 17 percent a year earlier. The causes are structural rather than technical.
  • Initiatives that reach production typically begin with a working prototype tested against a live business problem, followed by an explicit decision on whether to scale.
  • A small team of senior engineers, embedded in the client organisation and using agentic development methods, can deliver in weeks what conventional programmes deliver in quarters.
  • The model rests on four disciplines: forward deployment, working software as the unit of progress, staged commercial commitment, and independent quality assurance.

AI investment is running ahead of AI value

The evidence on enterprise AI delivery is consistent and unflattering. S&P Global Market Intelligence found that 42 percent of organisations abandoned most of their AI initiatives in 2025, up from 17 percent a year earlier, and that the average organisation discontinued 46 percent of its AI proofs of concept before they reached production. BCG reports that only 26 percent of companies have developed the capabilities required to move beyond proofs of concept and generate measurable value. Gartner reached a similar conclusion, predicting that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value.

What is notable in those findings is the absence of the technology itself among the causes of failure. The models are rarely the constraint. The constraint is the delivery system around them: long programmes, large blended teams, requirements gathered at a distance from the people who own the problem, and value that remains theoretical until the budget is exhausted.

Organisations that do reach production tend to share a common starting point. They build a working prototype against a live business problem, place it in the hands of the relevant business owners within weeks, and take an explicit, evidence-based decision on whether to scale. The differentiator is execution speed and decision discipline, not the sophistication of the strategy.

What we mean by an AI Factory

The term is chosen deliberately. A factory is not a project. It is a repeatable capability that converts inputs, in this case business problems, data and domain knowledge, into working and testable software on a fixed cadence and at a predictable cost. Built well, it becomes a durable asset: each prototype improves the organisation's understanding of its data, its integration points and its appetite for change, and each cycle costs less than the one before.

Four disciplines define the model.

1. Forward deployment

The engineers work inside the client team rather than alongside it. They join the client's stand-ups, share its channels and report to a nominated technical lead through the discovery period. The first week is spent on site with the people who own the problem, and it ends with a build-ready specification rather than a workshop summary. Proximity improves the quality of the questions asked and shortens the life of incorrect assumptions.

2. Working software as the unit of progress

Progress is measured in commits, not documents. Every week ends with a demonstration of the running prototype on current data. If the prototype is wrong, the prototype is changed. This reverses the sequence of conventional advisory work, in which analysis hardens into recommendations before anything has been tested against operational reality.

3. Staged commercial commitment

The engagement carries an explicit go or no-go decision at the end of discovery. If the first week does not produce a clearly scoped prototype worth building, the work stops at a small fraction of the total fee. This is incentive design rather than generosity. A team that must earn continuation in week one scopes honestly, and a sponsor who can stop at low cost decides quickly.

4. Independent quality assurance

Agentic development produces volume; senior review produces trust. Architecture, security posture and code quality are reviewed weekly by a senior engineer who sits outside the build team and is accountable directly to the executive sponsor. Every prototype passes through that review before handover, assessed against agreed criteria: security, maintainability, integration fit and clarity of the path to scale.

The economics of agentic engineering

The commercial logic of this model rests on a measurable change in how software is built. In a controlled experiment published by researchers at Microsoft, MIT and GitHub, developers using an AI pair programmer completed a defined implementation task 55.8 percent faster than a control group. That study is now several years old, and current agentic tooling, in which engineers direct multiple development threads in parallel, has extended the effect considerably.

In practice, one senior engineer can now direct several parallel work streams covering scaffolding, integration and test coverage, while reserving personal judgement for the decisions that carry risk: architecture, data contracts and security boundaries. Two senior engineers working in this way approach the delivery throughput of a conventional six-person team, without the coordination overhead that additional headcount brings.

The limits deserve equal attention. Agentic development is applied to well-defined tasks under senior architectural direction. Judgement about what to build, how components compose and where security boundaries sit is never delegated to tooling. That is why independent quality assurance is a structural feature of the model rather than an option: speed without adversarial review accumulates technical debt at the same accelerated rate.

Evidence from delivery

The same small-team model has now been proven across sectors. A global freight operation, following an unsuccessful two-year implementation by a major consultancy, received a working replacement system in three months; the platform now manages approximately $1 billion in annual freight transfers, integrated with the ERP and vessel-management estate. A multi-tenant scheduling and optimisation platform for the energy sector, backed by a national energy company joint venture, moved from kick-off to its first live international customers in under five months. An established billing software vendor reached a production-ready platform running multiple autonomous customer journeys within two months of project start.

The industries differ; the pattern does not. Senior engineers close to the problem, working software demonstrated weekly, and a decision gate that keeps the sponsor in control of the economics.

Managing the principal risks

  • Scope ambiguity after discovery. The decision gate is the primary control. If week one does not produce a clearly scoped prototype, the engagement closes at minimal cost.
  • Integration constraints. Discovery includes system walk-throughs and access provisioning. Where access is blocked, prototypes are scoped to replicated data with a defined path to live integration.
  • Remote delivery alignment. Daily stand-ups, a lead engineer aligned to the client's business hours, and an optional in-person sprint reserved for decisions that benefit from presence.
  • AI output quality. Agentic development under senior direction, with every prototype passing independent review before handover.
  • Stakeholder availability. A named executive sponsor and a single technical lead from signature, with demonstration slots protected in calendars from the outset. Delivery speed depends on the client's availability as much as on the engineering team, and no methodology removes that dependency.

The boundaries of an eight-week sprint

An eight-week sprint produces working software and an evidence-based scale-up decision. It does not produce a hardened production platform, a mature MLOps capability or an adopted way of working. Those require the next stage of investment, and it serves no one to suggest otherwise. The purpose of the factory is to make that next decision fast, inexpensive and grounded in something the organisation can use, test and challenge. The scale-up stage is the subject of From proof of concept to production.

About the author. Matthew Timms is CEO and Founding Partner of Third Horizon. He has led enterprise technology organisations across energy, banking and government, including a global digital, data and technology function of more than 3,000 people.