For a while the problem looked like arithmetic. The organization needed more research coverage than a small central team could give it, product teams wanted answers faster, designers wanted to talk to users without waiting in a queue, and demand kept climbing. The obvious answer was democratization. Train more people, hand them tools, let them run studies, move on.
As we started actually mapping what that would look like across dozens of pods, it became clear democratization was not the thing we were solving for. The real question was how to preserve rigor while participation went up, and those two things pull against each other in ways a training plan does not fix.
Research is not a pile of activities
Most organizations think of research as a set of activities. Usability tests, interviews, surveys, prototype reviews. Frame it that way and scaling becomes a throughput problem, which leads you straight to questions about how to run more studies and collect more feedback faster. Those are reasonable questions and they quietly skip the part that matters, which is that research does not create value because studies happened. It creates value because decisions got better.
Once that clicked, the goal stopped being more research and became more evidence-informed decisions, and the distribution question changed with it. We stopped asking whether designers could run a usability test, which is mostly an execution skill, and started asking whether they could recognize uncertainty, pick a method that fits it, gather evidence, and understand what that evidence does not cover. That second one is judgment, it is much harder to scale, and it is where nearly all the value sits.
Second-order effects
Scaling research is a systems design problem, because research outputs never land in isolation. They touch product planning, design decisions, engineering investment, roadmap prioritization, customer enablement, support. Every change ripples. If anyone can run a study, throughput goes up, and if governance disappears alongside it, confidence in findings goes down. Findings people are unsure about are easy to ignore, and once they are ignored research quietly loses its influence. A system tuned entirely for speed can end up costing you the organization's trust, which was the constraint we ended up designing around. Not maximum activity. Sustainable evidence.
Rigor is not the method
The misconception I run into most is that rigor comes from sophisticated methods. A badly framed interview does not become rigorous because it was moderated, and a leading usability test does not become rigorous at ten participants. A small, quick evaluative study can be enormously valuable if it answers a decision question someone actually has. So the bar we used had four questions in it: what decision is being made, what is uncertain about it, what evidence would reduce that uncertainty, and what action changes depending on the result. If a study cannot answer those, it is generating activity rather than insight, and I would rather not run it.
What surprised me was how closely this mirrored customer enablement work. In both cases the instinct when things get complicated is to throw documentation, training, or tooling at it. But tools rarely fix an understanding problem. Customers struggling with an implementation usually were not missing functionality, they were missing a model of the system they had built. Designers did not need access to a testing platform so much as a way to think about when research belongs in a decision, what it can answer, and what it cannot.
Research maturity is not how many studies you run. It is how well the organization turns evidence into better decisions.
The outcome I care about from that work was not a new process. It was moving the conversation from how do we run more studies to how do we make better decisions at scale, which is a much less satisfying thing to put on a slide and a much more durable thing to build on.
Filed by Erin Naylor — June 23, 2026