Engage
← All insights
Insight · AI industrialisation

From proof of concept to production: industrialising AI use cases.

At a glance
  • The average organisation discontinues 46 percent of its AI proofs of concept before production. A validated model marks the beginning of the engineering, not the end.
  • Build the platform rather than the project: a shared AI foundation makes each subsequent use case an increment rather than a new programme.
  • MLOps is an operating discipline rather than a toolset. Drift monitoring, versioned registries and automated retraining keep a model reliable long after launch.
  • Adoption warrants its own team, baseline and KPIs. A model that commercial teams do not act on returns nothing, whatever its accuracy.

Most AI value is lost between pilot and production

Most large organisations have now demonstrated that AI works on their data. Considerably fewer have made it work in their business. S&P Global Market Intelligence found that the average organisation discontinued 46 percent of its AI proofs of concept before they reached production. Gartner predicted that at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025. BCG found that only 26 percent of companies have built the capabilities to move beyond pilots and generate measurable value.

The gap is rarely the model. A typical commercial AI portfolio, covering customer segmentation, basket and assortment analysis and churn prediction, can be validated on sampled data in a matter of weeks. What separates a pilot from a production system is everything around the model: the full customer and product universe rather than a sample; hardened, monitored data pipelines rather than a data scientist's working environment; single sign-on, role-based access and private endpoints rather than a shared demonstration login; and several hundred commercial users applying the output in daily decisions rather than a steering committee reviewing it quarterly.

The root error is to treat that gap as a deployment task. It is a build in its own right, with its own architecture, economics and adoption challenge.

Build the platform, not the project

The instinctive way to scale three use cases is to run three projects. Each acquires its own pipelines, deployment scripts and security review; each duplicates well over half of the engineering of the others; and the estate fragments before it exists.

The alternative is to treat the first wave of use cases as the occasion to build a shared AI platform: one data foundation, one model registry, one CI/CD and MLOps pipeline, and one governed route to production that every model, current and future, travels. The economics compound. The first use case carries the platform cost; the second and third become increments; by the fourth, deploying a new model is measured in weeks rather than quarters, and the platform has become a durable asset of the group rather than an artefact of a single project. Data residency and sovereignty requirements are addressed once, in the platform's regional architecture, rather than renegotiated for each initiative.

MLOps is what keeps a model reliable

A model that is accurate at launch and unmonitored thereafter is a growing operational risk. Customer behaviour shifts, product ranges change and pipelines degrade quietly. A segmentation model trained on last year's patterns will continue to produce confident and increasingly inaccurate output unless something is watching.

MLOps is that discipline, and it is a cycle rather than a toolchain: data preparation and feature engineering; training; validation covering accuracy, drift and bias; a versioned registry; deployment to batch and real-time endpoints; continuous monitoring of model and infrastructure; and retraining triggered by new data or performance decay. Environments promote cleanly from development through QA to production, infrastructure is provisioned as code, and releases are triggered by code changes, data changes and model updates alike. None of this is novel. What distinguishes organisations that sustain AI in production is not access to the tooling but the discipline of treating the cycle as mandatory for every model, with no manually operated exceptions surviving beyond launch.

Adoption determines the return

The most demanding work in an AI scaling programme is not in the pipeline. It is in the moment a commercial manager decides whether to act on a churn score. A model that several hundred users quietly disregard returns nothing, whatever the quality of its architecture, and this factor, more than any technical one, is where AI programmes fail.

Adoption therefore warrants the same structure as the build: its own team, its own baseline and its own KPIs, running in parallel from the first day.

  • Start from the workflow, not the tool. Map user journeys and personas early. Identify where in the existing sales and marketing workflow each insight lands and which decision it should inform, then embed the output there rather than in a separate portal users must remember to visit.
  • Baseline before launch. Measure current workflow efficiency and decision outcomes before rollout. Without a baseline, the value discussion in month six cannot be resolved.
  • Coach by persona. Different roles require different fluency. Persona-specific learning journeys, embedded coaching and train-the-trainer models scale adoption; a recorded training session does not.
  • Instrument usage and iterate. Usage analytics showing who acts on which insights, and with what result, should feed the backlog. Prioritise visible early wins to establish confidence, then extend what demonstrably works.
  • Transfer capability deliberately. The end state is the client's own team operating the platform, the models and the adoption programme, with knowledge transfer planned from the outset rather than arranged at handover.

Three questions for leadership

Three tests indicate whether an AI scaling programme is positioned to deliver. Does the architecture make the next use case cheaper than the current one? Does a defined, automated cycle keep every model reliable after launch? And is a named individual accountable, with a baseline and a target, for whether commercial teams change how they decide? Programmes that satisfy all three tend to compound in value. Programmes that satisfy none contribute to the abandonment statistics cited above.

How the first working use cases are built quickly, before this scaling investment is committed, is addressed in The AI Factory. Where the resulting capability should sit within the organisation is addressed in Rethinking the IT operating model.

About the author. Matthew Timms is CEO and Founding Partner of Third Horizon, advising boards and investors on AI, technology and transformation across the Gulf, the UK and Europe.