Data Quality Management

Data quality management establishes whether data is fit for a particular use and provides a repeatable process for detecting and correcting defects. Fitness depends on the consumer: the same dataset can be adequate for a weekly trend report yet unsafe for an automated payment decision or a machine learning model.

Poor data quality rarely fails loudly. Dashboards keep rendering, pipelines keep succeeding, and models keep predicting — just wrongly. A data quality program makes expectations explicit, tests them automatically where data moves, assigns owners who can fix problems, and tracks whether fixes actually worked.

TL;DR

Quick Example

A data quality plan for one data product:

The same rules as dbt tests on the warehouse model:

Core Concepts

Quality Dimensions

Every score needs a population and denominator. "98% complete" means little until you know 98% of what — and averages can hide a critical segment that's 50% broken.

Validity Is Not Accuracy

A postal code can match the expected format while belonging to the wrong customer. Validity checks syntax and allowed values; accuracy compares against a trusted reference. Don't label syntactically valid data as proven correct.

Data Contracts and Ownership

A data contract documents a dataset's meaning, schema, semantics (units, time zones, status definitions), quality rules, freshness, and change process. The producer commits to it; consumers rely on it. An accountable owner decides acceptable use and prioritizes fixes; a steward coordinates issues day to day. Contracts turn "the data team's problem" into a shared agreement.

Where Checks Run

Handling Failures

Whatever you choose, never silently drop records from totals.

Data Observability

Beyond explicit rules, data observability tools monitor freshness, volume, schema changes, distribution shifts, and lineage automatically, catching the defects nobody wrote a test for. Lineage shows which dashboards and models a broken table affects.

Best Practices

Start With One High-Impact Dataset

Interview its consumers, identify critical fields and failure consequences, write a handful of meaningful rules, and baseline them before setting thresholds.

Keep Every Rule Explainable

Each rule needs a business purpose, owner, scope, threshold, and response. A few actionable checks beat hundreds of ignored alerts.

Enforce Contracts in CI

Run schema and data tests on pull requests to transformation code, and fail builds on breaking changes. See dbt.

Version Semantic Changes

Changing a status definition or unit of measure can break consumers without changing a column type. Treat it as a contract change with notice.

Fix at the Source

Patching bad data downstream in every consumer multiplies work and hides the cause. Fix the producing system or workflow, then replay.

Reconcile After Remediation

After correcting and reprocessing, compare counts and totals with the source and notify consumers whose reports or decisions changed.

Common Mistakes

Silent Filtering

WHERE customer_id IS NOT NULL in a transformation quietly drops revenue from reports.

Alert Floods

Hundreds of noisy checks without owners train teams to ignore alerts entirely.

Testing Only Schema

Correct types and non-null columns don't catch a currency conversion bug or duplicated events.

No Owner for Remediation

Detecting a defect without anyone accountable for fixing it produces dashboards of known problems.

Duplicates From Replays

Reprocessing without idempotent loads (merge/upsert on keys) turns one defect into two.

FAQ

What is data quality management?

A discipline for defining what "good enough" data means for each use, measuring it with automated checks, assigning ownership, and correcting defects so decisions and systems rely on trustworthy data.

What are the main data quality dimensions?

Completeness, validity, accuracy, uniqueness, consistency, and timeliness. Some frameworks add integrity, conformity, and lineage.

What is a data contract?

An agreement between a data producer and its consumers describing a dataset's schema, meaning, quality rules, freshness, and how changes are communicated.

Which tools test data quality?

dbt tests and contracts, Great Expectations, Soda, and database constraints for rule-based checks; data observability platforms for anomaly detection, freshness, and lineage.

Should failing data block the pipeline?

It depends on consequences. Block for critical automated decisions, quarantine failing rows for most operational data, and flag or allow with alerts for low-impact analytics.

Related Topics

References