Part 6 · 1 chapters · ~8 min

Data Quality and Data Contracts

Dimensions of data quality (freshness, completeness, validity, uniqueness, consistency), tests in dbt and Great Expectations or Soda, anomaly detection on volumes, reconciliation against source systems, data contracts between producers and consumers, and ownership of data incidents.

8

Tests, reconciliation, contracts

code
# a data contract (excerpt, in the style of the Open Data Contract Standard)
dataset: core.transfers
version: 3.0.0
owner: team-payouts
consumers: [finance-reporting, risk-models, growth-dashboards, regulator-returns]
schema:
  - { name: id, type: uuid, required: true, unique: true }
  - { name: amount_kobo, type: bigint, required: true, description: "minor units, always positive" }
  - { name: status, type: string, enum: [pending, completed, failed, reversed] }
quality:
  - freshness: "max(created_at) within 15 minutes"
  - reconciliation: "sum(amount_kobo) where status = completed equals ledger postings per day"
quality dimensiontest
freshnesslatest row within the agreed delay
completenessrow counts versus source; volume anomaly alerts (a 40% drop is a bug, not a holiday, until proven otherwise)
validitytypes, enums, ranges (amounts positive)
uniquenessprimary keys unique after every load
consistencytotals reconcile with the ledger and with finance reports
A DATA CONTRACT CATCHES A BREAKING CHANGE
the producer's CI knows who consumes the table
backend engineerproducer CIcontract registrydata teamPR: rename transfers.amount → amount_kobo
swipe the figure sideways, or tap expand for full screen
1/4
the usual failure
A backend engineer renames a column in a production table. CDC carries the change downstream and the next morning's dashboards are broken or silently wrong.
upstream changes break downstreamfound by the CFO, not CI