Skip to content

Taking a machine learning capability from experiment to production service

A well-performing model that no part of the organisation was structured to run. What it took to close that gap, and what it cost.

Summary. A capable model with no route into production, in an organisation that had never put one there. Building the route turned out to be the more valuable half of the work.

Context

A large regulated bank with a capable data science function and a newly established shared machine learning platform. Analytical work was strong. Production machine learning was, at that point, largely theoretical: models were built, demonstrated, and did not move.

I owned the platform and the product path through it.

The problem

A transaction categorisation capability existed as working analytical work. It performed well. It had obvious downstream value in more than one part of the business. And it had no route into production, because no route existed for anything of its kind.

The framing that mattered was recognising that this was not a model delivery problem. It was the first instance of a class of problem, and whatever we built to solve it would either become the path everything else used, or would be thrown away.

My role

Product owner for the platform and for this use case's path through it. I did not build the model. I owned whether it could ever run, what it would run on, what it had to satisfy, who would operate it, and who would consume it.

Constraints

  • A regulated environment with several control functions holding genuine veto, none of which had previously assessed a production machine learning service.
  • No established precedent, so every requirement had to be discovered rather than looked up.
  • A live banking environment, where the cost of a production incident is not measured in engineering time.
  • A small squad with concurrent platform commitments, so this could not consume everything.

Stakeholders

Data science, who built the model and had a reasonable expectation it would ship. Platform engineering, who had to run it. Risk, privacy, security, data governance, architecture and model risk, each holding an approval. An operations function that would eventually carry it. And business areas downstream who would consume the output and whose processes would have to change to use it.

The interesting stakeholder problem was that none of them had done this before either, so several were being asked for decisions they had no precedent for.

Decisions

Treat it as two products, not one. The use case, and the path. Sequenced so that the path work was reusable, and accepting that this made the first delivery slower than it needed to be.

Engage control functions during build rather than after. More elapsed meetings, far less rework, and it converted the control functions from reviewers into people with context.

Design for reuse before there was a second consumer. Difficult to justify at the time, on the argument that a single-consumer service is a project and a multi-consumer service is a capability. This proved to be the decision that mattered most.

Refuse to route production operations through the squad. Operational ownership had to sit with a team structured to hold it, even though absorbing it ourselves would have been faster in the short term and was repeatedly the path of least resistance.

Approach

Incremental delivery, with the control conversation running as a parallel track rather than as a phase. Platform capabilities built as general capabilities that this workload used, not as bespoke pieces for it. Explicit sequencing of the gates so there was never ambiguity about what came next.

Outcome

The capability moved into governed production and now operates at meaningful daily transaction volume, with downstream consumption in more than one business area rather than the single area that originally sponsored it.

The larger outcome was the path itself. Subsequent workloads did not repeat the discovery, because the route existed and was documented.

Lessons

The first one is infrastructure, whatever it looks like on the plan. The first instance of a class of work is where the organisation decides how that work happens. Scoping it as a single delivery is the most common and most expensive mistake available, because you get the delivery and not the capability.

Design for the second consumer before you have one. It is nearly impossible to justify and it is where the return actually comes from.

Do not absorb operational ownership to move faster. It works, briefly, and then the platform team is running production services it is not structured to run, and cannot build anything new.

Elapsed time and wasted time are different numbers. Involving control functions early made the calendar longer and the total effort substantially smaller. Leadership conversations go better when you separate those explicitly rather than defending the calendar.

What I would do differently

I would put the operational ownership conversation first, before the build rather than during it. We had it late, under time pressure, which made it a negotiation rather than a design decision, and the answer we reached was right but harder-won than it needed to be.

I would also have written the path down earlier. It existed in my head and in a sequence of conversations for longer than it should have, which meant the reuse benefit arrived later than it could have.

Confidentiality

This is a generalised account based on professional experience. Specific employer systems, internal names, architecture, customer information, proprietary metrics and operational details have been omitted. It does not represent the views of any current or former employer.