Skip to the content
Work

The First Hire on a Two-Person Data Team Must Solve the Bottleneck You Can Actually Name

The first additional hire on a growing data team should be a data engineer if raw data infrastructure is the bottleneck, a data scientist if structured data already exists and business questions remain unanswered, and a data analyst when the priority is operationalizing reporting.

5 min read

A slide rule, an instrument that answers only the question it is set up for
A slide rule, an instrument that answers only the question it is set up for. Photo: Jacek Halicki · Wikimedia Commons · CC BY-SA 4.0

This taxonomy, from Databricks, strips away the usual title confusion and forces a harder question: can you actually identify which kind of work is choking your two-person operation?

Most teams cannot. They hire by résumé keyword instead of constraint, and the result is predictable. Two people with overlapping generalist skills end up doubling the same failure mode rather than removing it.

Diagnose the Bottleneck Before the Job Description

The four roles commonly considered for an early data hire solve genuinely different problems. A data engineer develops and maintains data architecture and pipelines, according to IBM's May 2022 guidance, while Microsoft describes the same role as managing and securing the flow of structured and unstructured data from multiple sources. An analytics engineer, by contrast, brings together data sources so consolidated insights can be produced repeatedly. Data scientists study large data sets using advanced statistical analysis and machine learning algorithms. Data analysts profile, clean, and transform data, then extract actionable insights from what the engineers have prepared.

University of Wisconsin–Madison InterPro's November 2024 analysis clarifies the handoffs: analytics engineering bridges data engineering and data analysis, using platforms and processes created by the data engineer. The analyst works downstream of both. The scientist may work across the stack but needs structured, accessible data to function.

The Databricks framework is explicit: choose the role that matches your constraint. This sounds obvious until you observe how teams actually decide. A founder who spent a previous role drowning in dashboards assumes they need an analyst. A CTO who inherited brittle ETL scripts concludes a data engineer will fix everything. Neither has logged what actually consumed the past week's labor, or what sat waiting while the two existing team members scrambled elsewhere.

The diagnostic week is straightforward in principle. Log every incoming request. Classify it: ingest and quality work, modeling, question-answering, or product-built systems. Measure which category generates the queue. The bottleneck is the category where requests accumulate while your two people work elsewhere. The logic is reconstructible from constraint-based hiring frameworks.

The Four Job Shapes and Where They Sit

Understanding the distinctions prevents the common error of treating the four roles as interchangeable skill levels rather than different functions.

Data engineering sits furthest upstream. Without it, nothing else functions. The analytics engineer depends on this foundation, consolidating sources into repeatable insight pipelines. The analyst then extracts actionable insights from prepared data sets. The scientist, where applicable, performs descriptive and predictive analytics on what the stack delivers.

The University of Wisconsin–Madison InterPro description captures the practical dependency: analytics engineering uses the platforms and processes created by the data engineer. Attempting to hire an analytics engineer or analyst when your data engineer is already underwater simply adds another person waiting on the same blocked pipe.

Microsoft's July 2023 role definitions emphasize security and flow management for engineers, profiling and transformation for analysts. These are not aesthetic preferences. They describe different daily work with different tooling, different success metrics, and different failure modes.

Why the Wrong First Hire Misfires

Hiring the wrong specialist does not produce half-value. It produces negative value. The new hire requires onboarding, management attention, and platform access. If their skills address a non-constraint, they join the queue for the actual bottleneck's output. The existing two-person team now splits focus between their original overload and the new colleague's blocked work.

The structure of Databricks' guidance implies familiarity with the pattern. They lead with infrastructure-first hiring for a reason: raw data problems are the most common and least visible constraint. Teams notice unanswered business questions more readily than they notice missing lineage documentation or failed incremental loads.

The reverse failure also occurs. Teams with clean, modeled data and a burning business question hire an engineer to "build the foundation properly," deferring insight generation for quarters. The constraint was never infrastructure. It was analytical bandwidth and statistical capability. The hire satisfies a technical aesthetic while the actual bottleneck persists.

Describing the Role by Its Output

Given verified source constraints, the job description should articulate the bottleneck being solved and the artifact it must produce. Not a list of tools. A description of what stops happening when this role succeeds.

If the diagnostic week shows ingest and quality as the constraint, the data engineer role addresses failure in data flow from multiple sources. The advert emphasizes pipeline reliability, schema evolution, and observability. If modeling is the bottleneck, the analytics engineer role emphasizes transform logic, testing frameworks, and stakeholder-accessible datasets. If question-answering queues, the analyst role emphasizes turnaround time on defined business metrics and communication with requestors. If predictive or inferential work waits, the data scientist role emphasizes experimental design and deployment of statistical models.

What the advert must not do is blend these. "We need someone who can do a bit of everything" translates to two people doing the same scattered work, now with a third. The specificity of the Databricks framework, applied honestly, prevents this.

When the Platform Is the Problem

There are circumstances where no hire is the right move. If the diagnostic week reveals that the existing two-person team spends most hours repairing a vendor platform, fighting access controls, or manually compensating for missing orchestration, the constraint may be tooling rather than headcount.

No verified source explicitly states a platform-first, no-hire rule for two-person teams. The inference is reconstructible: hiring specialized labor to work around broken infrastructure consumes salary budget that could replace the infrastructure. This is particularly true in data engineering, where IBM and Microsoft both emphasize architecture and security management. A platform that cannot be secured or scaled absorbs engineering hours that a replacement platform would free.

The practical test is whether the bottleneck would persist with a technically perfect hire. If a senior data engineer would still spend weeks on access provisioning and brittle connector maintenance, the problem is upstream of staffing.

The Hiring Decision Is Only as Good as the Diagnosis

If your two-person team cannot state which of the four work categories generated last week's waiting queue, the hiring decision is premature. The Databricks framework, the IBM and Microsoft role definitions, and the University of Wisconsin–Madison InterPro analysis all assume a legible constraint. Without that legibility, the best-case outcome is a well-credentialed generalist who doubles your existing failure mode. The advert you write should name the bottleneck explicitly, select for the role that removes it, and repel applicants whose skills address the wrong problem.

Sources

  1. Microsoft Learn, “Roles in data” — learn.microsoft.com, 2023-07-03
  2. Databricks, “Data Science vs Data Engineering: Choosing Analysis or Infrastructure” — databricks.com, 2026-01-05
  3. University of Wisconsin–Madison InterPro, “Demystifying Data Roles: Data Engineer, Analytics Engineer” — interpro.wisc.edu, 2024-11-14
  4. IBM, “Data Engineer vs Data Scientist vs Analytics Engineer” — ibm.com, 2022-05-26

More from Work & Leadership

Section index
Work

How to apply, as the firm described it

Join a team of professionals dedicated to one another’s success. Contact us to learn about our job opportunities and to find out why the firm has been consistently identified as one of the…

29 Sep 2026
2 min

Work

Careers at a Pittsburgh promotions firm

Innovation is the key to success in business. We believe this so strongly at the firm that we consider team development to be one of our highest priorities. By building a thoroughly trained…

25 Sep 2026
2 min

Work

Replacing resolutions with habits

It seems that every year, the masses make new resolutions and vow to keep them, only to falter a few weeks into their commitments. At the firm, we came across one study indicating that…

31 Aug 2026
2 min

Independent trade desk. We take no commission on anything we describe and run no affiliate programme of our own. Every figure on this page names the standard, filing or organisation it comes from; where a number could not be verified the page says so. How we work and how we correct. Reviewed:

Cookies, and what this site stores. The Dispatch sets no advertising or analytics cookies and loads no third-party tracker. Closing this notice writes one key — icd-notice — into your browser’s local storage, so that the notice does not return. Nothing else is kept. The one thing a page here sends onward is what a reader types into the form on the contact page, and that is described before the form is used. What the policy says.