Welcome to Part Two of our three-part series on dbt project health where we look at the three areas that most commonly hold data platforms back: pipeline reliability, compute costs and deployment velocity. In this blog we argue that cost inefficiency is a design debt problem concentrated in a small number of assets, most often driven by materialisation strategy, environment sprawl and underused platform features, and explains why fixing it for good requires named accountability and cost attribution rather than a one-off cleanup.
Blog Series Introduction
Our three-part series on dbt project health is written specifically for CDOs and Heads of Data and looks at the three areas that most commonly hold data platforms back: pipeline reliability, compute costs (this blog) and deployment velocity.
Most data platforms look healthy on the surface. Jobs run, dashboards load and dashboards show green. But underneath that surface, three separate problems tend to accumulate quietly: data that’s wrong but undetected, compute spend that’s climbing for no clear reason and delivery timelines that keep stretching further from what the business expects. None of these show up as a single dramatic failure. They build up gradually, and by the time they’re visible enough to act on, the cost (in credibility, budget or lost momentum) has already been paid.
When Cloud Costs Stop Making Sense
At some point in the life of every established data platform, the bill for cloud services stops making sense. Volume hasn’t doubled. Headcount hasn’t surged, but spend is climbing and nobody can explain exactly why.
The gap between what you’re paying and what you’d expect to pay is almost always a design debt problem, not an infrastructure problem. The platform has grown incrementally, and each increment has added cost without anyone stepping back to ask whether the underlying approach still holds. By the time finance asks the question, the answer is buried across dozens of models, environments and workload patterns that no single person has full visibility over.
For leaders, this creates a specific organisational challenge. The people best positioned to diagnose the cost growth are often the last to notice it, because cost efficiency is rarely a named responsibility on any data engineering team. It accumulates, silently, skulking around in the shadows until a budget conversation forces it out into the open.
Where the Excess Spend Actually Lives
The distribution of cost inefficiency in data platforms is highly concentrated. In practice, a small number of data assets account for the majority of excess compute spend. This concentration is both the cause of the problem and the reason it’s manageable to fix.
The most common driver is the materialisation strategy, how and when data is computed and stored. The default behaviour in most data transformation frameworks is to calculate data on demand – which sounds efficient, but means that a heavily-used data asset recalculates itself every time it’s queried. A single asset queried two hundred times a day by dashboards and downstream processes can consume more compute than the entire rest of the platform. Materialising it once, calculating it on a schedule and retaining the result, reduces that to a single computation with no downstream cost.
The second driver is environmental and workload sprawl. Platforms that don’t clearly separate development, testing and production workloads bill all three against the same compute resource, with no visibility into which activities are generating which costs. Development workloads running on production-grade infrastructure is a consistent finding in platform cost reviews.
Finally, platform-specific optimisations. Features that Snowflake and Databricks provide to reduce data scanned per query, is the third lever and the one most frequently left. These features can dramatically reduce compute costs, but they require deliberate configuration aligned with actual query patterns rather than speculative setup. This can be greatly amplified by pipelines that have been lazyly migrated from legacy technologies, which can actively vex optimisation features of modern systems.
The Governance Question Behind the Cost Question
Cost efficiency in data platforms doesn’t stay solved without governance. Organisations that address it as a one-time project typically find costs climbing again within six to twelve months, as new assets are added with the same design patterns that created the original inefficiency.
Organisations that manage data platform costs most effectively have made two structural decisions:
- Someone is accountable for cost efficiency as an ongoing responsibility. Not as a project owner when finance raises a concern, but as a named part of someone’s role, with visibility into cost attribution and the authority to prioritise optimisation work.
- Cost-per-output visibility exists. The ability to attribute platform spend to specific data assets, business functions or consuming teams changes the conversation from “the cloud spend is high” to “this specific report costs £XXX to produce”. This level of attribution is achievable with the tooling most platforms already have, it just requires deliberate setup.
What Addressing It Actually Delivers
The return on addressing cost inefficiency properly, rather than reactively, is typically direct and significant, but the full value goes beyond the immediate spend saving.
Fixing materialisation strategy and environment sprawl reduces the current bill, but it also makes future spend more predictable. When costs scale in line with value delivered rather than accumulated design debt, investment decisions become easier to justify. A CDO who can attribute every pound of platform spend to a specific output is in a fundamentally stronger position when making the case for expanded AI and analytics capabilities.
The governance improvements that keep costs under control, named accountability, cost attribution, design standards, raise the overall standard of platform engineering. Teams that build these habits produce more reliable, more maintainable platforms across the board. The savings fund the governance work, the governance work prevents the costs from climbing again.
The Strategic Opportunity
There’s a forward-looking argument for investing in cost efficiency that goes beyond reducing the current cloud expenditure.
Organisations leading on data and AI aren’t necessarily the ones with the largest data budgets, they’re the ones where investment is managed with enough rigour that the returns are visible and the case for more is credible and justifiable. A data platform that can demonstrate cost efficiency, spend is attributable, governed and proportionate to business value, is a platform that earns the right to grow.
That case is easier to make from a position of control than from a position of catch-up.
What This Means for Data Leaders
The diagnostic question worth asking teams today: “If I asked you which ten data assets are the most expensive to run, could you tell me by end of week?”
If the answer is no, cost attribution isn’t in place – and optimisation decisions are being made without the information needed to make them well.
Analytics8 has spent over two decades working with data organisations across industries. We bring the cross-industry pattern recognition to identify where excess spend is concentrated, and the senior delivery accountability to fix it without disrupting production. The same team that assesses your platform builds the solution and ensures your people own it at the end. No handoffs. No diagnosis without delivery.
Our dbt Project Health Check includes a structured cost efficiency assessment alongside pipeline reliability and deployment velocity – giving you a complete picture of where your platform stands and a clear, prioritised roadmap for improvement.
Ready to find out where the cost is hiding? Book a free, no-pressure strategy session with one of our senior consultants. We’ll talk through where you are, what’s blocking progress, and what a practical path forward could look like.