is the practice of monitoring and managing data.
It ensures their quality, availability, and reliability
At its core is tracking data quality metrics that explain how they are measured, quantitatively or qualitatively.
Observability practice tracks the following metrics: data-to-error ratio.
To calculate this ratio, measure the number of known data errors in the dataset and compare it to the total dataset size.
When errors decrease while the data volume remains constant or increases, data quality improves. Number of empty values.
Empty values indicate that important information is missing or was entered into the wrong field. Data transformation errors.
Problems converting data from one format to another indicate quality issues. Volume of dark data.
Dark data is collected and stored by a company but never used.
Large volumes of dark data signal data quality problems, since no one bothers to examine it.
Many companies do not fully realize the potential value of the data they hold.
To use this data, it must be extracted and its correctness, consistency, and completeness assessed. Data storage costs.
If data storage costs rise while the volume of data in use stays the same, some of the data is poor quality.
Time to value from data.
A high number of errors during data conversion or a need for manual cleansing points to low data quality.
The faster a team turns data into business value, the higher the data quality.
Email bounce rates.
Sales and marketing can only succeed with a high-quality email address list.
Data on current and potential customers can degrade quickly, reducing dataset quality and campaign effectiveness. Percentage of duplicate records.
Duplicate records appear due to data-entry errors, system issues or other causes.
Their number reflects the quality of data management. Data update delays.
Irregular data updates lead to decisions based on outdated information.
This metric helps keep data current and relevant. Data pipeline incidents.
Data pipelines are systems that collect, process, and transfer data from one place to another.
Monitoring the number of incidents in pipelines, such as failures or data loss, helps teams identify points where data integrity may be compromised. Table availability.
This is an aggregate indicator that measures the overall health of a database table.
It includes the number of missing values, data range, and record integrity in a table, metrics that provide a comprehensive assessment of data quality for individual datasets.