Data Quality Management
Data quality management establishes whether data is fit for a particular use and provides a repeatable process for detecting and correcting defects. Fitness depends on the consumer: the same dataset can be adequate for a weekly trend report yet unsafe for an automated payment decision or a machine learning model.
Poor data quality rarely fails loudly. Dashboards keep rendering, pipelines keep succeeding, and models keep predicting — just wrongly. A data quality program makes expectations explicit, tests them automatically where data moves, assigns owners who can fix problems, and tracks whether fixes actually worked.
TL;DR
- Define quality relative to a consumer and a business consequence, not in the abstract.
- Measure separate dimensions: completeness, validity, accuracy, uniqueness, consistency, timeliness — each with a clear denominator.
- Write data contracts between producers and consumers covering meaning, schema, and quality rules.
- Test at pipeline boundaries (ingestion, transformation, delivery) with tools like dbt tests or Great Expectations.
- Decide per rule whether failures reject, quarantine, flag, or allow — never silently drop records.
- Fix defects at the source, then replay, reconcile, and notify affected consumers.
Quick Example
A data quality plan for one data product:
The same rules as dbt tests on the warehouse model:
Core Concepts
Quality Dimensions
Every score needs a population and denominator. "98% complete" means little until you know 98% of what — and averages can hide a critical segment that's 50% broken.
Validity Is Not Accuracy
A postal code can match the expected format while belonging to the wrong customer. Validity checks syntax and allowed values; accuracy compares against a trusted reference. Don't label syntactically valid data as proven correct.
Data Contracts and Ownership
A data contract documents a dataset's meaning, schema, semantics (units, time zones, status definitions), quality rules, freshness, and change process. The producer commits to it; consumers rely on it. An accountable owner decides acceptable use and prioritizes fixes; a steward coordinates issues day to day. Contracts turn "the data team's problem" into a shared agreement.
Where Checks Run
Handling Failures
- Reject — stop the load (for critical, all-or-nothing data).
- Quarantine — route failing rows to a holding table with reasons, continue with the rest.
- Flag — load with a quality flag consumers can filter on.
- Allow and alert — for low-impact issues while the fix is in progress.
Whatever you choose, never silently drop records from totals.
Data Observability
Beyond explicit rules, data observability tools monitor freshness, volume, schema changes, distribution shifts, and lineage automatically, catching the defects nobody wrote a test for. Lineage shows which dashboards and models a broken table affects.
Best Practices
Start With One High-Impact Dataset
Interview its consumers, identify critical fields and failure consequences, write a handful of meaningful rules, and baseline them before setting thresholds.
Keep Every Rule Explainable
Each rule needs a business purpose, owner, scope, threshold, and response. A few actionable checks beat hundreds of ignored alerts.
Enforce Contracts in CI
Run schema and data tests on pull requests to transformation code, and fail builds on breaking changes. See dbt.
Version Semantic Changes
Changing a status definition or unit of measure can break consumers without changing a column type. Treat it as a contract change with notice.
Fix at the Source
Patching bad data downstream in every consumer multiplies work and hides the cause. Fix the producing system or workflow, then replay.
Reconcile After Remediation
After correcting and reprocessing, compare counts and totals with the source and notify consumers whose reports or decisions changed.
Common Mistakes
Silent Filtering
WHERE customer_id IS NOT NULL in a transformation quietly drops revenue from reports.
Alert Floods
Hundreds of noisy checks without owners train teams to ignore alerts entirely.
Testing Only Schema
Correct types and non-null columns don't catch a currency conversion bug or duplicated events.
No Owner for Remediation
Detecting a defect without anyone accountable for fixing it produces dashboards of known problems.
Duplicates From Replays
Reprocessing without idempotent loads (merge/upsert on keys) turns one defect into two.
FAQ
What is data quality management?
A discipline for defining what "good enough" data means for each use, measuring it with automated checks, assigning ownership, and correcting defects so decisions and systems rely on trustworthy data.
What are the main data quality dimensions?
Completeness, validity, accuracy, uniqueness, consistency, and timeliness. Some frameworks add integrity, conformity, and lineage.
What is a data contract?
An agreement between a data producer and its consumers describing a dataset's schema, meaning, quality rules, freshness, and how changes are communicated.
Which tools test data quality?
dbt tests and contracts, Great Expectations, Soda, and database constraints for rule-based checks; data observability platforms for anomaly detection, freshness, and lineage.
Should failing data block the pipeline?
It depends on consequences. Block for critical automated decisions, quarantine failing rows for most operational data, and flag or allow with alerts for low-impact analytics.
Related Topics
- Data Engineering — Pipelines where quality checks live
- dbt — Tests and contracts in the transformation layer
- Data Warehousing — Modeling data consumers rely on
- Change Data Capture — Moving source changes reliably
- Compliance & Privacy — Governance obligations on data