Loading...
This recipe automatically scans every table in your data workspace and grades it on completeness, uniqueness, validity, accuracy, consistency, timeliness, integrity, and conformance — turning raw data into a clear health scorecard. It goes beyond scores: it detects values that break the expected format (such as an email column that starts receiving plain digits or text), flags potential PII (emails, phone numbers), and names the exact tables, columns, and rows affected. Each finding comes with a plain-language explanation, example values, the number and share of records impacted, and, where a table has an event date, when the problem started. Findings are ranked into prioritized recommendations, so teams know exactly what to fix first and why it matters. Results are visualized in an interactive Data Quality Officer dashboard that shows which tables are AI-ready, which need urgent attention, and how fresh the underlying data is.
Most teams only find data quality problems after a report, a model, or a customer exposes them. This recipe automatically scans every table added to your workspace, with no rules to write, and grades it across data quality dimensions — completeness, uniqueness, validity, accuracy, consistency, timeliness, integrity, and conformance — while flagging potential PII. It goes beyond counting nulls and duplicates: it detects values that break the expected format (for example, an email column that starts receiving plain digits or text), placeholder values such as N/A or TBD, orphaned and missing keys, and records that disagree across tables. For every finding it shows the table, the column, the affected rows, example values, how much data is affected, and, where the table has an event date, when the problem started. Everything lands in one prioritized remediation list and an interactive dashboard, so teams know exactly which tables need attention, why, and what to fix first.
Step 1 — Profile, detect, score & recommend
The recipe scans every table added to the workspace and works through the following, with no rules to configure:
• Profile — row and column counts, data types, null and blank rates, distinct counts, top values, min/max, and string lengths for every column, plus the business role of each column (identifier, date, contact, measure, status, and so on).
• Detect column-level issues — nulls and blanks, repeated or non-unique key values, columns holding a single value, very long text, and exact duplicate rows.
• Detect format violations — for every text column, the recipe works out the expected format from the column name, the column content, or the column's dominant pattern, then flags values that break it. It recognizes email, phone and date columns as well as code formats such as ORD-000123. Each bad value gets a reason (for example, missing @, digits only, placeholder text, too few digits, different pattern), and severity rises with the share of affected values.
• Detect PII — columns that look like email addresses or phone numbers are flagged separately from general quality issues.
• Check across tables — related tables are found through shared identifiers. The recipe compares matching columns for consistency and tests parent/child relationships for orphaned or missing records.
• Find when it started — where a table has an event date (such as an order or created date), the recipe shows whether a problem has been present from the start or began on a particular date.
• Collect evidence — for each finding it records the affected column, number and percentage of records, example values, row numbers, and the failure reason, and writes the affected rows to a separate evidence table.
• Score & prioritize — everything is rolled into a weighted Completeness, Uniqueness, Validity, Accuracy, Consistency, Timeliness, Integrity, Conformance, and Security & PII scorecard per table (each status names the columns behind the score), one AI-readiness composite score, and a single severity-ranked list of recommended fixes.
Step 2 — Visualize & prioritize
The recipe turns those results into eight linked dashboard views: completeness ranked worst-first, duplicate rows alongside validity issues, PII exposure by table, recommendations broken down by severity per table (so you see exactly which table needs urgent work), recommendations broken down by issue category, the AI-readiness composite score, data freshness (how stale the latest record is), and the advanced accuracy/consistency/integrity/conformance checks. Each panel renders independently, so a table missing one metric never blocks the rest of the dashboard.
| Dimension | Issues it finds | Example |
|---|---|---|
| Completeness | Null and blank values per column; missing values in identifier columns | 89.5% of a remarks column is empty |
| Uniqueness | Exact duplicate rows; repeated values in key columns | A transaction ID that appears twice |
| Validity | Columns that hold a single value; unusually long text values | A country column that is always the same value |
| Accuracy (format) | Values that break the expected format for emails, phone numbers, dates, or code patterns; placeholder text such as N/A, TBD or UNKNOWN | “abc123” or “99887766” in an email column; a date of 2025-13-40 |
| Consistency | The same record described differently in two related tables | A customer’s city differs between two tables |
| Integrity | Orphaned foreign keys; null identifiers where a key is expected | Orders that reference a customer who does not exist |
| Conformance | Phone and date formats; categorical columns whose values stray from the dominant set | A status column with unexpected variants |
| Timeliness | How old the most recent record is | A table whose latest record is 9 months old |
| Security & PII | Columns that look like email addresses or phone numbers | Contact columns that need masking or access controls |
| Insight Category | What the recipe discovered | Business Impact |
|---|---|---|
| Headline health check | Ranking tables by completeness and AI-readiness score immediately surfaces the two or three tables dragging down overall workspace quality, instead of an even spread of minor issues. | Leadership can see, in one glance, exactly which datasets are safe to build on and which need work before any AI or reporting use case. |
| A field that quietly stops being valid | An email column was clean for the first several hundred records, then 3% of new values arrive as plain digits, text or “NA”. The finding names the column, lists the affected rows and example values, and shows the date the problem started. | Teams catch a form change or broken integration before campaigns bounce or a model trains on bad contacts, and know where to look. |
| Codes that change shape | Order references follow ORD-000123 for most records, but a later batch arrives as bare numbers. A ZIP code column loses its leading zero for a small share of rows. | Joins, lookups and address matching stop working silently; the evidence table shows exactly which rows to correct. |
| Placeholders hiding as data | A supplier email column looks populated, but 13% of values are “TBD”, “N/A” or a bare domain with no @ symbol. | Completeness numbers alone would have looked healthy; the recipe shows the field is not actually usable for outreach or analytics. |
| Broken relationships between tables | Cross-table checks catch identifier columns with nulls where a key is expected, and foreign-key columns pointing to records that no longer exist in the parent table — with the orphaned values and row numbers listed. | Fixing these before they reach a join or a dashboard prevents undercounted or duplicated figures in downstream reports. |
| Compliance-ready remediation queue | Columns carrying potential PII (email, phone) are flagged separately from general quality issues, and every recommendation — PII or not — is ranked by severity into one ordered backlog. | Governance teams get a ready-made, defensible worklist instead of having to triage scattered findings themselves. |
Make sure the following ingredient is available: