Data
AI 



 

AI in Data Engineering: The Agent Writes the Pipeline. The Team Still Gets More Expensive.  

Written by
Benjamin Gnädig VP Data & Cloud
 Benjamin Gnädig is VP Data & Cloud at HICO Group and specialises in turning data into strategic value through scalable architectures and cloud solutions. With extensive consulting experience, he drives innovation across Data Engineering, Analytics and Cloud Transformation.  


Publication date
October 2, 2026
Share Article

There is an assumption showing up in 2027 budget discussions right now that nobody says out loud, but almost everyone is already calculating with: if the agent writes the pipelines, the data team needs fewer people.

I think that assumption is wrong. Not on principle, but on arithmetic. The work does not disappear. It moves from implementation into specification and verification. And both are more expensive per hour than much of the work that disappears.

The Number You Should Know Before Cutting the Data Team

In July 2025, METR did something that is still rare in this debate: a randomized controlled trial. Sixteen experienced open-source developers worked on 246 real issues from repositories with more than one million lines of code. The tasks took around two hours on average. These were not unfamiliar codebases. The developers had contributed to the repositories for years.

The result: with AI tools, the developers were 19 percent slower. Before the study, they had expected AI to make them 24 percent faster. Afterwards, they still believed they had been around 20 percent faster.

The interesting number is not the 19 percent. It is the gap of roughly 40 percentage points between perceived and measured performance. People who worked with these tools every day did not recognise their own slowdown.

The study has limitations. It is small, it used models from early 2025, and it looks at a very specific setting: experienced developers working on large, mature open-source repositories they know well. I still cite it because it is one of the cleanest measurements I know, and because the pattern matches what I see in teams.

What Actually Shifts When AI Writes the Pipeline

In the data teams I work with, the workload roughly breaks down like this: 45 percent implementation, 20 percent specification and business alignment, 20 percent review and testing, and 15 percent operations and incident handling. Those are my field numbers, not a benchmark.

Now the agent arrives, and it is good. It cuts implementation effort in half. Forty-five percent becomes 22. In a team of six, that looks like roughly 1.4 freed-up full-time positions. That is exactly the number that then appears in the budget draft.

What is missing from the same draft is the offset. The agent needs a precise target state. Not “rebuild the revenue logic,” but which sources, which granularity, how cancellations are handled, how late postings across a period boundary are treated. If you do not specify that first, you get code that runs and is wrong. In my projects, specification effort rises from around 20 to about 28 percent.

Then comes verification. The DORA report calls this the verification tax and identifies it as one of the reasons for the initial dip in the AI productivity J-curve.

For application code, that verification burden is inconvenient. For data logic, it is something else entirely, and I think we talk about that far too little.

Broken application code crashes. Someone sees an error message. Someone gets called. Someone rolls it back.

A broken transformation often does not crash. It produces a number. That number flows into reporting, planning or commission calculations, and it looks exactly like every other number.

In DORA’s illustrative model, the change failure rate increases from 5 to 6 percent after AI adoption, creating a modeled negative downtime impact of $344,000. That is software. For data, one thing is missing from that calculation: time to detection.

In one case I saw, an incorrect aggregation ran for eleven weeks before anyone noticed. Fixing the logic took an afternoon. Unwinding the decisions made on top of those numbers took a quarter.

That is why review effort rises more sharply in data teams than in many development teams. My field estimate is from around 20 to 35 percent.

The Arithmetic Behind AI in Data Engineering

Take six data engineers with fully loaded internal costs of €110,000 per person. Annual team budget: €660,000.
Implementation: minus 23 percentage points.
Specification: plus 8 percentage points.
Review and verification: plus 15 percentage points.
Total: zero.

The team does not get smaller in year one. It does different work.

And this is the part that does not show up in the budget line. The 23 points that disappear are largely work a junior engineer can do after a solid onboarding period. The 23 points that appear are work that requires domain knowledge and decision authority.

Headcount cost may stay similar. The skill mix moves upward. If you have three juniors and three seniors today, in two years you may need something closer to one junior and four seniors.

Run the numbers. In the projects I see, fully loaded internal costs are roughly €85,000 for a junior and €135,000 for a senior. Three juniors and three seniors equal €660,000. One junior and four seniors equal €625,000 with five people. Keep the team at six people, however, and you are closer to €760,000.

The gap between those two scenarios, €135,000 a year, is the real decision. It is rarely treated as one. Instead, it emerges quietly from who leaves the organisation over the next two years and who gets replaced.

The savings currently written into many 2027 plans are not there yet. They may come later, but only under one condition.

The Condition Is Verifiability

DORA’s broader point is that the largest returns do not come from the tools themselves. They come from the organisational and engineering foundations around them. For data teams, that becomes very concrete.

Verification gets cheaper when you have data contracts that automatically flag violations. When test cases exist for the twenty most important metrics and run with every change. When lineage reaches down to column level so that a review does not start from zero.

The starting point is smaller than most organisations assume. You do not need full coverage. You need the metrics that decisions actually depend on, and a test for each of them that runs every time something changes.

In the organisations I see, that is somewhere between twelve and thirty metrics. Not two hundred. Building that foundation takes weeks, not a transformation programme.

Teams that have it can win back much of the additional review effort. Teams that do not have it verify manually, and then the agent simply accelerates a foundation that cannot support the speed.

That is why I am stubborn about the order. First verifiability. Then speed. Do it the other way around and you create exactly the situation METR measured: everyone feels faster, nobody is, and nobody notices.

The Counterargument I Take Seriously

The strongest argument against everything I have written so far is this: you are measuring the first year and treating it as the end state.

That is fair. The METR study used models from early 2025. The tools inside modern data platforms have improved significantly since then. The J-curve described by DORA starts with a dip and then rises. If you wait because of the dip, you may end up two years from now without mature tooling and without a team that knows how to use it.

There is also a labour-market argument. Gartner predicts that by 2027, 75 percent of hiring processes will include certifications or testing for workplace AI proficiency. If you deny your team that experience, you make them harder to place in the market and harder to retain.

Both arguments are valid. My objection is narrow, but I think it matters: neither argument justifies putting the savings into the 2027 budget today.

Start using the tools, yes. Book the savings in advance, no. If you do, you cut exactly the positions you need for specification and verification, and you create the instability that the DORA model warns about.

What This Means for Data Leadership

The most uncomfortable consequence is not the cost calculation. It is the entry point into the profession.

The work that makes data engineers senior is increasingly the work the agent can take over: building pipelines, hunting bugs, learning why a join duplicated rows, understanding what happens when supposedly simple logic meets messy real data.

If you have never gone through those years, judging an agent’s output becomes difficult. You can only believe it.

If we remove entry-level positions because the agent can do the junior work, then five years from now we may have nobody left who learned how to verify the agent’s work.

That is not a labour-market problem that somebody else will solve eventually. It is a decision every organisation makes for itself, usually without recognising it as a decision.

My position is simple: the role does not disappear. The entry path disappears if we do not protect it deliberately. And that is the more expensive of the two outcomes.

What I Would Like to Know

Have you actually gained capacity in your data team since introducing AI agents? And if you have: was it measured, or did it just feel that way?

I mean the second part seriously.

If you are currently looking at how AI agents could change the setup of your data team, book a 30-minute exploration call with us. We can look at where implementation effort is really falling, where specification and verification are increasing, and what needs to be in place before those gains can safely turn into actual capacity.

Book an Ex​​​​ploration Call


Sources

METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
DORA: ROI of AI-Assisted Software Development
Gartner: Top Predictions for Data and Analytics in 2026​​​​