Improving platform onboarding as a product problem
Summary. Onboarding a team onto a shared platform took weeks, and the assumption was that this was a documentation gap. It was not. It was an unowned product.
Context
A shared enterprise machine learning platform with a growing number of internal consumer teams. Adoption was the stated goal and the platform was funded on that basis.
The problem
Getting a new team productive took a period best measured in weeks. Everyone described this as a documentation problem, and several attempts had been made to fix it by writing more documentation, which had not worked.
Tracing what actually happened to an arriving team showed why. The elapsed time was not going into environment setup, which was quick. It was going into: working out which access requests were needed and from whom, discovering which control requirements applied to their particular workload, waiting on approvals that could have been requested a fortnight earlier, and asking a platform engineer questions that had been asked by every previous team.
None of that is a documentation problem. It is a sequencing and ownership problem wearing a documentation costume.
My role
Platform product owner. I owned the decision to treat onboarding as a product with a defined outcome rather than as a support activity absorbed by whoever was free.
Constraints
- The platform team was small and already committed. Any solution that required a dedicated onboarding person was not available.
- Consumer teams varied in technical maturity, so a single rigid path would fail at both ends.
- Access and approval steps sat with other functions and could not be shortcut, only sequenced better.
- The improvement had to survive the departure of any individual, including me.
Stakeholders
Incoming data science teams, who experienced the cost. Platform engineers, who were absorbing the repeated questions and losing build time. Access management and the control functions, whose steps were in the critical path. And leadership, for whom onboarding time was a visible proxy for whether the platform investment was working.
Decisions
Define the outcome precisely. "Onboarded" was previously vague. We defined it as a named team with an environment, the right access, a known control path for their intended workload, and a first piece of work running. Vague outcomes cannot be improved because nobody can tell when they are reached.
Move the discovery work to the front and make it a form, not a conversation. Most of the delay was caused by information being gathered late. Asking a small set of structured questions on day one let every downstream request be raised in parallel rather than in sequence.
Standardise the reusable parts and refuse to standardise the rest. Environment shape, access patterns and the control path were made identical for everyone. What teams actually built stayed entirely theirs. This is the responsibility line applied to onboarding.
Put it in the workflow tool teams already used so the state of an onboarding was visible to everyone involved without anyone having to ask.
Assign it an owner. Not a dedicated person, an accountable one. Unowned processes decay quietly and nobody notices until it is a complaint.
Approach
Instrumented first: traced two recent onboardings end to end and recorded where the days actually went, rather than where people believed they went. The gap between those two things was the argument for changing anything.
Then built the standardised path against the next real team arriving, adjusted it against the one after, and only then documented it.
Outcome
Onboarding moved from a multi-week exercise to a matter of days. The reduction came almost entirely from parallelising requests that had previously been discovered and raised sequentially, not from any technical change.
A secondary outcome mattered more over time: platform engineers stopped losing build capacity to repeated first-week questions.
Lessons
Instrument before improving. The team's belief about where the time went was wrong, and every previous fix had been aimed at the wrong target. Two traced examples were enough to redirect the whole effort.
Most internal-platform delay is sequencing, not duration. Steps that must happen serially in people's heads can very often happen in parallel in reality.
"Write better documentation" is what an organisation says when a process has no owner. It is almost never the fix, and it is always the cheapest-sounding one.
Onboarding is a product with users, an outcome and a funnel. Treating it as a support activity guarantees it stays bad, because support activities are absorbed rather than designed.
What I would do differently
I would have measured it from the beginning rather than after the complaints reached a threshold. The instrumentation was cheap and would have been just as informative a year earlier.
I would also have resisted the first instinct to solve it with documentation for longer than I did. We spent effort there before tracing the actual path, and that effort produced very little.
Confidentiality
This is a generalised account based on professional experience. Specific employer systems, internal names, tooling, timeframes, customer information and operational details have been omitted. It does not represent the views of any current or former employer.