Designing a production gate across six control functions
Summary. Six legitimate control processes and no map between them. The work was building the map, and the design constraint was that nobody could lose a decision they already owned.
Context
A shared machine learning platform serving several data science teams inside a regulated bank. Demand was growing. Each team was independently working out how to get work approved for production, and each was arriving at a slightly different answer.
The problem
Six control functions applied: risk, privacy, security, data governance, architecture and model risk. Each had a legitimate process. None of them was the problem.
The problem was that no artefact described the combined route. So every team discovered it, in a different order, largely by finding out what they had missed. That rediscovery cost was paid repeatedly, appeared in no plan, and was attributed to governance being slow.
Two failure modes followed. Teams were surprised late, which is expensive. And because the standard was implicit, what actually got approved depended partly on who reviewed it, which is a control weakness in its own right.
My role
I owned the platform and the design of its production promotion gate. I did not own any control decision, and it was important that this stayed true. Specialist approval authority for privacy, model risk and security remained with those functions throughout.
Constraints
- No authority to change any control function's process, nor any wish to.
- Could not slow existing delivery while building this.
- Had to work for workloads of varying risk profile, so a single heavyweight path for everything would have been worse than the status quo.
- Had to survive staff turnover, which meant it had to live in tooling rather than in anyone's head.
Stakeholders
Six control functions, each with its own priorities and no obligation to help document their process. Data science and engineering teams, who wanted it faster. Platform engineering, who would operate the gate. And leadership, who wanted a number for how much quicker things would get, which was the hardest conversation of the lot.
Decisions
Separate control authority from path ownership, and say so explicitly to each function early. This was the decision that made cooperation possible. Nobody was being asked to give anything up.
Map the existing route before designing anything. Trace what actually happened to work that had already been through, including the loops. The real path was not the documented one, and the gap between them was the design brief.
Specify artefacts, not just approvals. Most recoverable time was being lost to teams producing the wrong evidence and producing it again. Naming exactly what each gate needed removed more delay than any sequencing change.
Make the platform the enforcement point, not the decision point. The platform confirms that the defined approvals are complete before promotion. It does not assess whether they should have been given. This distinction is the whole design.
Put it in the tool teams already used rather than in a document, so following the path was the default action rather than an act of diligence.
Approach
Built incrementally alongside real workloads rather than as a standalone programme, so it was validated against actual work continuously. Each control function reviewed its own section separately, which surfaced corrections that a combined workshop would have smoothed over.
Outcome
A repeatable lifecycle with defined gates, specified artefacts and templated workflows, operating as the platform-side promotion gate for machine learning workloads.
The clearest effect was on variance rather than duration. Teams stopped being surprised. Onboarding a new team onto the platform went from a multi-week exercise to a matter of days, because most of what made it slow was rediscovery.
It also held up under a real test: when the platform extended into generative AI on a second cloud provider, the same lifecycle structure adapted rather than being rebuilt, which suggested the model was principled rather than specific to the first case.
Lessons
Name the authority split out loud, early. Governance work stalls when control functions suspect a platform team is trying to acquire their decisions. Saying plainly that you want to own the route and not the ruling changes the entire conversation.
Map before you design. The documented process is not the process, and you cannot see the gap from a policy document.
Specify the evidence, not just the steps. Knowing that a review is required is worth much less than knowing exactly what to bring to it.
Expect the map to be politically visible. Writing down the route makes clear where the delays are and which requirements are ambiguous.
Budget for maintenance or do not start. A route that is not kept current is worse than none, because people follow it and are wrong.
What I would do differently
I would have set the leadership expectation about elapsed time much earlier and much more bluntly. The honest claim was always about variance and rediscovery cost, not about halving review durations, and letting an optimistic framing sit unchallenged for a while made the eventual conversation harder than it should have been.
I would also have built the maintenance mechanism into the first version rather than adding it once decay became visible.
Confidentiality
This is a generalised account based on professional experience. Control functions are described by generic category, which is standard vocabulary in every regulated organisation and specific to none. Specific employer systems, internal names, policies, artefacts, customer information and operational details have been omitted. It does not represent the views of any current or former employer.