Author David Beaty 9 minute read

Welcome to our three-part series on dbt project health where we look at the three areas that most commonly hold data platforms back: pipeline reliability, compute costs and deployment velocity. In Part One we focus on The Data Reliability Gap and explore the strategic case for treating data quality as an engineering discipline, not a reactive fix.

Blog Series Introduction

Our three-part series on dbt project health is written specifically for CDOs and Heads of Data and looks at the three areas that most commonly hold data platforms back: pipeline reliability (this blog), compute costs and deployment velocity.

Most data platforms look healthy on the surface. Jobs run, dashboards load and dashboards show green. But underneath that surface, three separate problems tend to accumulate quietly: data that’s wrong but undetected, compute spend that’s climbing for no clear reason and delivery timelines that keep stretching further from what the business expects. None of these show up as a single dramatic failure. They build up gradually, and by the time they’re visible enough to act on, the cost (in credibility, budget or lost momentum) has already been paid.

Silent data failures cost more than crashes

The most dangerous failure mode in a modern data platform isn’t the one that stops everything. It’s the one that lets everything keep running while quietly shipping wrong data to the business.

Hard failures – jobs that crash, models that error, pipelines that halt – are operationally painful but strategically manageable. They’re visible, they’re loud and they’re fast to diagnose. The failure mode that should concern data leaders more is the silent one: a pipeline that completes successfully, passes every automated check and delivers subtly incorrect data to every report and dashboard built on top of it. By the time that failure surfaces, it usually surfaces through a stakeholder – and the damage to data credibility is already done.

Why Coverage, Not Architecture, Is the Real Differentiator

The temptation when facing reliability problems is to reach for architectural solutions – better infrastructure, more sophisticated tooling, a platform migration. These investments rarely address the actual cause.

Data reliability is primarily a coverage problem. The best-engineered platform in the world will ship bad data if nobody has thought carefully about what to assert, where to assert it and how to surface failures before they reach the business. The gap between teams that catch their own problems and teams whose stakeholders catch them for them is almost never a technology gap. It’s a test strategy gap.

Three areas consistently have insufficient coverage in mature data platforms:

  • Source contracts: The assumption that upstream data quality is someone else’s responsibility is one of the most dangerous and expensive assumptions in data engineering. Source systems change – formats drift, schemas evolve, feeds go silent – and without explicit assertions on what your platform expects to receive, those changes propagate invisibly downstream.
  • Intermediate layers: Staging and transformation models that sit between source data and business-facing outputs are frequently under-tested. A silent failure here compounds across every downstream asset simultaneously.
  • Business rule validation: Structural tests – does this field have values, are they unique – are necessary but not sufficient. The assertions that catch the failures stakeholders actually notice are the business rule ones: revenue within expected ranges, customer counts not dropping overnight, conversion rates inside plausible bounds. These require a genuine conversation with the business about what “correct” looks like – and most organisations haven’t had it.

The Strategic Shift: From Reactive to Proactive Quality

The maturity trajectory for data reliability follows a predictable pattern. Teams start reactive – tests get added after something breaks. Over time, coverage becomes uneven: dense around the models that have caused pain before, sparse everywhere else. The organisation is always one upstream change away from an incident it won’t find before the business does.

The shift to proactive quality requires three things that are organisational as much as technical:

  • A defined minimum standard by data layer: Mature organisations define what adequate coverage looks like at the source, transformation and mart layers – and treat deviations from that standard as technical debt to be scheduled and addressed, not accumulated indefinitely.
  • Source freshness monitoring as a first-class concern: Knowing whether your data arrived – before your transformations run – is a basic capability that is consistently underinvested. The setup cost is hours. The diagnostic value is catching ingestion failures before they become reporting incidents.
  • Alerting that reaches the right people before stakeholders do: The goal of operational monitoring isn’t to log failures. It’s to surface them to the people who can act on them, fast enough that they don’t become business incidents. If the first notification of a data quality failure comes from outside the data team, the monitoring posture needs immediate attention.

What Good Looks Like – and What It Delivers

Across the dbt projects Analytics8 has assessed and delivered, the pattern is consistent. Organisations that treat reliability as an engineering discipline – not a reactive process – don’t just avoid incidents. They build data credibility that changes how the business uses data.

More decisions are made with confidence. Less time is spent validating numbers before presenting them. There’s a genuine appetite to expand what the data platform is asked to do. That shift – from reactive to proactive, from fixing to preventing – is the foundation that makes everything else in a data strategy possible.

It’s also the shift that takes longer than most teams expect, because the barriers aren’t primarily technical. They’re cultural: defining what “good” means, agreeing standards across the organisation and building the habits that keep coverage high as the platform grows. That’s where experienced outside perspective consistently accelerates progress.

What This Means for Data Leaders

The question worth putting to your team today: “How would we know if critical data was wrong before a stakeholder told us?”

If the honest answer is uncertain, the reliability posture needs investment – not in infrastructure, but in coverage, monitoring and the engineering culture that treats quality as a first-class concern.

Analytics8 has worked with data organisations across industries for over two decades. We know what good reliability engineering looks like at every layer of the stack – and we know the organisational patterns that make it stick. The same senior team that assesses your platform builds the solution and ensures your people own it at the end. No handoffs. No shelf-ware. No disappearing after the diagnosis.

Our dbt Project Health Check is the right starting point – a structured, independent assessment that surfaces exactly where your coverage gaps are and sets out a clear, prioritised path to close them. Book a free, no-pressure strategy session with one of our senior consultants. We’ll talk through where you are, what’s blocking progress, and what a practical path forward could look like.