Your data team shouldn’t start with an org chart. It should begin with the problems you need solved.
Do you need someone to establish trusted business metrics? Explain why conversion declined? Determine whether a product change actually caused an improvement? Forecast what happens next? Build a system that predicts which customers are likely to churn?
Those are different problems. They require different capabilities. Yet companies often start somewhere else. They decide they need a Data Scientist, then look for Data Science work.
Start with the decisions the business needs to make and the uncertainty preventing it from making them well. Then determine the capability you need, whether your data is ready to support it, and whether the problem is important and recurring enough to specialize around. Only then decide what role to hire.
Start with the problem. Let that define the role. Worry about the title later.
A framework for deciding what data capability you need
When I think about building a data organization, I ask three questions:
1. What problem are you trying to solve?
The first question is not “Do we need a Data Scientist?” It is: “What uncertainty is getting in the way of a better decision?”
Most data problems fall into a few broad categories:
Understand: What is happening in the business? Where is it happening? For whom? Why might it be happening?
Measure: Did something actually cause the outcome we observed? Did a product, pricing, marketing, or operational change work?
Predict: What is likely to happen next? Which customers are likely to churn? What will demand look like next month?
Optimize or automate: Given what we know, what decision should we repeatedly make at scale?
Enable: Can people reliably access and interpret the data at all? Are the underlying data, definitions, and infrastructure trustworthy enough to support the other questions?
That last category is easy to overlook. Sometimes your most important data problem is not analytical sophistication. It is that nobody agrees on the denominator.
Start with the decision
Imagine customer retention has declined. If you do not know where, descriptive analysis may be enough. If you know where but not why, investigate the drivers. If you suspect a recent product or pricing change, causal analysis may be appropriate. If you have an intervention, test it. And if that intervention works but cannot be targeted efficiently, prediction may finally create leverage.
Each method can be valuable. But sophistication is not the objective. The most valuable method is the one appropriate for the uncertainty standing between the business and a better decision.
This is how I think about Data Science broadly: its job is to reduce uncertainty around important decisions. Sometimes that requires a simple query. Sometimes it requires causal inference, forecasting, or machine learning. The method should follow the problem, not the other way around.
2. Is your organization ready to use that capability?
The next question is whether your data and organization can actually support the work you want someone to do.
Companies often hire for the version of themselves they hope to become. They recruit a machine learning specialist because they believe machine learning will matter eventually. Then that person arrives and discovers that event tracking is unreliable, metric definitions are inconsistent, and nobody agrees on what counts as an active customer!
Instead of building ML models, they spend six months reconciling dashboards and fixing data tables. None of that work is beneath them. But it may not be the work they wanted to do, and it may not be the person you actually needed to hire.
Before adding a sophisticated capability, ask:
Is the necessary data captured?
Can we trust it?
Are the core entities and metrics defined?
Can people retrieve the data without heroic effort?
Do we have enough history or volume for the methodology we want to use?
Will the business actually use the output to make a decision?
If not, your constraint may be foundational rather than analytical.
Match the hire to the data you actually have, not the model you wish you had the inputs for.
3. Is the problem important and recurring enough to specialize around?
This is where I think organizations often specialize too early.
A capable generalist can solve an enormous range of analytical problems. I would specialize when a class of work becomes important, recurring, and complex enough that the generalist model itself starts creating a bottleneck: quality becomes uneven, the same infrastructure keeps getting rebuilt, or demand is persistent enough to justify dedicated ownership.
Roles should separate when specialization solves a real problem, not simply because the company has become large enough to draw another box on the org chart.
The answers also tell you what kind of person to hire. Early on, that is often a broad, hands-on generalist. As the work changes, the team should change with it.
Of course, we did not sit down at early DoorDash with this framework and design the organization perfectly. Much of this became clear in retrospect, as we saw which problems kept recurring and where the generalist model started to strain.
What this looks like as a company grows
There is no magic employee count for any of this. But company size can still be a useful shorthand because the nature of the problems tends to change as organizations grow. I would think about the progression in stages, not rules.
Very early: the business is still learning what matters
In many companies, the very early stage looks roughly like 10–50 people. At that point, a specialized data organization is usually premature, although a data-intensive product, marketplace, fintech, or highly regulated business may need dedicated expertise earlier. For most companies, founders and operators may still be able to understand the most important metrics directly.
The signal to hire is not simply that the company has accumulated data. It is that important decisions are repeatedly constrained by the company’s ability to trust, retrieve, or interpret it. You start hearing the same symptoms repeatedly: teams disagree on metrics, nobody knows why an important number moved, simple questions require an engineer, or the same analysis gets rebuilt from scratch.
That is when a dedicated data owner starts to become valuable.
At this stage, I would generally bias toward one senior, hands-on generalist. The problems are still broad, changing quickly, and often poorly defined. Range matters more than specialization.
At DoorDash, this was essentially my role after transitioning from a General Manager. As we began building the team, the sweet spot was often around an L6 Data Scientist. The level itself is not the point, and levels vary enormously across companies. What mattered was the shape: senior enough to independently own an ambiguous business problem and work with senior stakeholders, but still hands-on enough to write the query, inspect the data, and build the analysis themselves.
That hands-on requirement matters whether this person is the first member of the team or the person leading it. Building a data function from zero to one is a very different job from leading a mature organization. There may be little infrastructure, few established processes, and nobody to delegate the hard work to.
I have seen companies hire a Director+ leader too early because they are optimizing for the person who might eventually run a large organization. Early on, I would optimize instead for a player-coach: someone senior enough to set direction and build credibility with the business, but willing and able to roll up their sleeves and build alongside the team.
Growing: one generalist can no longer absorb everything
As the company grows beyond 100 people, more teams rely on data, and one person can no longer absorb all the demand. That does not necessarily mean it is time to specialize. Often, the right next step is simply to build a team of strong generalists.
At early DoorDash, that is largely what we did. Data Scientists might support different parts of the business, but the shape of the role remained broad. The same person could investigate a marketplace problem, define a metric, design an experiment, build a data table, and work directly with a business or product leader on next steps.
The first sign that one generalist is overloaded is usually a reason to add capacity, not necessarily a reason to specialize.
AI may allow this stage to last longer than it used to. A strong generalist can now write queries faster, prototype lightweight workflows, navigate unfamiliar code, and communicate findings more efficiently. If the problems are still broad and ambiguous, another strong problem solver may create more value than a specialist whose expertise applies to only a fraction of them.
At DoorDash, we started to specialize after hiring four generalists. We added a machine learning scientist to the Analytics team. We added another ML scientist and two BI engineers after we crossed ten people on the Analytics team. As the company crossed the 1,000-employee mark, the Analytics team was about 20 people.
Scaling: specialization starts to create leverage
Eventually, more demand becomes a different kind of problem. Certain types of work stop being occasional and become important, recurring, and complex.
At DoorDash, we started seeing the same kind of problems recur across more and more teams. Experimentation was no longer an occasional technique; it was becoming central to product decisions. Shared data models and metric definitions increasingly affected whether different teams reached the same answer. Forecasts became inputs into planning, and machine learning began moving from analysis and prototyping into systems that had to operate reliably in production.
The signal was not that our generalists had suddenly become incapable. It was that solving these problems well now required consistency, infrastructure, and depth that were hard to create one project at a time.
The question had changed. It was no longer “Can someone on the team solve this?” It was “Should we build a real capability around solving this repeatedly?”
That is when specialization starts to earn its keep. Shared models and definitions may justify dedicated Analytics Engineering. Experimentation may need common expertise and standards. Forecasting or Machine Learning may become critical enough to require specialized ownership. Our few specialists were overloaded, and they needed more support.
The trigger is the work, not the headcount.
At scale: specialized work becomes organizational capability
At scale, the specialist’s job changes too. They are no longer just solving the hard problem themselves. They are building the systems that let the company solve it reliably: experimentation infrastructure and standards, trusted data models and metric definitions, production systems that are monitored and maintained.
The progression is roughly:
No dedicated data team → first generalist → team of generalists → specialized capabilities → scaled systems
Company size can help you understand where to look. But the problems themselves should tell you when to move from one stage to the next.
Titles describe centers of gravity, not hard boundaries
The exact titles vary enormously across companies. That is part of the problem. One company’s Data Scientist is another company’s Data Analyst, Product Analyst, Decision Scientist, Applied Scientist, or Machine Learning Scientist.
I find it more useful to think of these titles as different centers of gravity. Some roles sit closer to understanding the business and improving human decisions. Others sit closer to building reusable data, infrastructure, or automated decision systems. Data Science can span a surprisingly large portion of that landscape.
The point is not to draw hard walls. Some crossover is healthy, especially while a team is exploring what is valuable. Once something becomes important, reusable, or business-critical, clearer ownership matters more.
Overlap in skills is healthy. Duplication in ownership is not. When role boundaries become confusing, I would inspect the roadmap rather than the org chart: What decisions does each team own? Where are multiple teams answering the same question? And where are important questions falling between them?
Design the organization around owned problems and decisions. Then decide where the role boundaries should sit.
Define the job before you recruit for the title
That ambiguity creates a very practical hiring problem. LinkedIn usually tells you someone’s title. It does not necessarily tell you what they actually did. Three people with the title of Data Scientist may have fundamentally different experience: one may build machine learning models, another may run experiments, and another may primarily do product analytics.
That is why the hiring manager has to define the job before asking a recruiter to source it. This is also where recruiter calibration matters. If the hiring manager says “find me Data Scientists” without explaining what that means in practice, the recruiter has little choice but to use titles as a proxy. Then, several interviews later, you discover the candidate pool is full of people with the right label but the wrong experience. The recruiter needs to know what evidence to look for beyond the title.
Before opening a req, write a one-page role brief that answers six questions:
The first five define the job. The last determines what to call it.
The title comes last
Data organizations should evolve because their problems evolve. Early on, the right hire may be a hands-on generalist. Later, the same company may need specialized experimentation, Analytics Engineering, forecasting, or Machine Learning. Those changes should happen because the work changed, not because the company crossed some arbitrary headcount threshold.
Company size can tell you roughly where to look. The problems tell you what to build.
The problem defines the capability. The capability defines the role. The title comes last.
Acknowledgements: The ideas and writing are mine, but they’ve been shaped by current and former teammates who challenged my thinking and helped evolve how we built and ran the Analytics team. Special thanks to my Chief of Staff, Anita Chan, for brainstorming, editing, and AI wizardry on the images, and to ChatGPT and Claude for serving as editorial critics.



Really like this framing. It’s such a good reminder that titles can be misleading, and the more useful question is what problem someone actually owns and how their work creates value.
The idea of starting with the capability you need, rather than forcing everything into a title, really resonated with me.
Reading this on a plane and I'm just nodding my way right to my destination. I've never read a piece on data titles and role descriptions better explained than this piece. I feel like all the newsletters you've written so far speaks directly to the stage in which our data team is in. I really appreciate you sharing all of this - it's so validating!