Data quality was for a long time a very popular topic for large companies, just like probably every area concerning data. With the arrival of tools and technologies that made working with data accessible even to much smaller players without huge data engineering or data analytics teams and the like, much smaller companies on the market also started dealing with data quality.
For this reason the market for tools that would address data quality and that smaller companies could also afford is not very mature.
For a company with 500 employees it is no great problem to start working with tools from, for example, Atacama, which provides tools for metadata management, including Data Quality Management. Unfortunately a lot of these tools are aimed squarely at business to enterprise, and the typical integrator or reseller is a firm such as the big four (Accenture, Deloitte and so on).
Not least for this reason, many solutions are built by small internal teams, where the integration of these solutions unfortunately often places great demands on data engineers.
Categories of Data Quality
Data quality is very often classed under the category of metadata management. This is because metadata is very often used for data quality checks, e.g. table sizes, the date of the last change, the schemas themselves and so on. Data quality generally solves the problem that if you use an ETL (ELT) pipeline, or even simply work with data, you need to be able to rely on that data, or at least know how far you can rely on it. Because when working with data an endless number of problems can occur, whether on the technical or the human side.

A user who works with a dataset, whether in a BI tool, in emailing or in any other tool, does not want the first thing they run into to be that half of the products in the warehouse have no purchase price. Or that the revenue in the dataset comes from four sources, but each source uses a different time format. And on top of that, the data engineer who joined the datasets has decided to ignore this fact, or the datasets have changed over time.
News and Interesting Points in Data Quality – How AI Can Be Involved
As in every technology sector, the terms of artificial intelligence, or more often rather machine learning, are making their way into data quality. This is mainly for 2 reasons:
- The first is to simplify the integration of data checks. Writing down which checks should be run on which table can be very tedious. It is enough to own even just a few hundred datasets. Here a simple algorithm can identify, for example, columns containing phone numbers, dates and the like, and recommend the most common tests for these columns.
- The second reason is that we identify trends, whether when working with metadata or with the data itself. With every iteration 150–200 rows are written to the dataset, and suddenly 2,000 were written. This means a non-standard situation has occurred and it is possible that, for example, we wrote the same rows 10 times instead of once. We should know about this situation and possibly check it manually.
These are common problems that a data quality system can alert us to immediately. The event can then be handled by a person or by another algorithm that helps us confirm or rule out the alert.
Data Quality and Archetix
At Archetix we have recently been devoting great attention to data quality, because we have reached perhaps the worst phase. Clients wrote to us "this chart doesn't look right to me", "something here doesn't add up" or "this is wrong here"…
These queries take up a large amount of our time and are entirely justified. When I use reporting, I have to be able to rely on it.
At the moment we deal with data quality separately from the rest of metadata management. We think data quality is essential for all our clients regardless of their maturity and size.
I think we have come a fair way in this area, but we are still at the beginning. In the coming years the market in this sector will see significant development, and I am curious what changes await us in the near future.
