AI-Augmented Reconciliation: Automating Valid/Invalid Record Classification in Post-Load Validation
Main Article Content
Abstract
Post-load validation, the step in which records freshly loaded into a target system are checked against their source and against business rules before they are trusted for downstream use, remains one of the least automated stages of most data migration and integration programs. Traditional reconciliation relies on deterministic checks, row counts, checksums, key existence, and hand-written business rules, that catch obvious breakage but struggle with subtler failure modes such as silent truncation, character encoding drift, partial duplicate loads, and transformation logic that behaves correctly on most records but not all of them. This article evaluates whether machine learning can be layered onto deterministic reconciliation to classify each loaded record as valid or invalid more accurately, more quickly, and with better-calibrated confidence than rule-based checks alone
We propose a reference architecture for AI-augmented reconciliation that combines a deterministic rule engine, a supervised classification ensemble trained on historical validation outcomes, an unsupervised anomaly detection fallback for record categories with little labeled history, and a calibrated confidence scoring layer that routes ambiguous records to human review while allowing high-confidence classifications to proceed automatically. We evaluate this architecture against rule-based reconciliation, logistic regression, random forest, and isolation forest baselines on a simulated benchmark spanning five common target tables in an enterprise data platform, with invalid records injected across six recurring root cause categories
Across the benchmark, the hybrid ensemble achieved an F1 score of 0.91 for invalid record classification, compared with 0.56 for rule-based reconciliation alone, while reducing mean time to classify a load batch from 26 hours to 4.8 hours by the tenth simulated load cycle. Automated classification coverage rose from 41 percent to 94 percent of loaded records over the same period, as the feature store accumulated a growing history of confirmed valid and invalid outcomes. Referential breaks and transformation mismatches remained the hardest categories to classify reliably across every method tested, underscoring that record-level context, not just column-level statistics, is often required to catch the most consequential invalid records
We close with a discussion of design trade-offs, current limitations, threats to validity, and directions for future research, including active learning for feature store growth and standardized benchmarks for post-load validation classifiers
Article Details
Section
How to Cite
References
1. Abedjan, Z., Chu, X., Deng, D., Fernandez, R. C., Ilyas, I. F., Ouzzani, M., Papotti, P., Stonebraker, M., & Tang, N. (2016). Detecting data errors: Where are we and what needs to be done? Proceedings of the VLDB Endowment, 9(12), 993 to 1004.
2. Arocena, P. C., Glavic, B., Mecca, G., Miller, R. J., Papotti, P., & Santoro, D. (2015). Messing up with BART: Error generation for evaluating data cleaning algorithms. Proceedings of the VLDB Endowment, 9(2), 36 to 47.
3. Batini, C., & Scannapieco, M. (2016). Data and information quality: Dimensions, principles and techniques. Springer.
4. Baylor, D., Breck, E., Cheng, H. T., Fiedel, N., Foo, C. Y., Haque, Z., Haykal, S., Ispir, M., Jain, V., Koc, L., Koo, C. Y., Lew, L., Mewald, C., Modi, A. N., Polyzotis, N., Ramesh, S., Roy, S., Whang, S. E., Wicke, M., ... Zinkevich, M. (2017). TFX: A TensorFlow-based production-scale machine learning platform. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1387 to 1395.
5. Breck, E., Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2019). Data validation for machine learning. Proceedings of Machine Learning and Systems, 1, 334 to 347.
6. Breunig, M. M., Kriegel, H. P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, 93 to 104.
7. Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15.
8. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321 to 357.
9. Chu, X., Ilyas, I. F., Krishnan, S., & Wang, J. (2016). Data cleaning: Overview and emerging challenges. Proceedings of the 2016 ACM SIGMOD International Conference on Management of Data, 2201 to 2206.
10. Dasu, T., & Johnson, T. (2003). Exploratory data mining and data cleaning. Wiley.
11. English, L. P. (1999). Improving data warehouse and business information quality: Methods for reducing costs and increasing profits. Wiley.
12. Fan, W., & Geerts, F. (2012). Foundations of data quality management. Morgan and Claypool Publishers.
13. Ganti, V., & Sarma, A. D. (2013). Data cleaning: A practical perspective. Morgan and Claypool Publishers.
14. Haller, K. (2009). Towards the industrialization of data migration: Concepts and patterns for standard software implementation projects. Lecture Notes in Business Information Processing, CAiSE Forum 2009.
15. He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263 to 1284.
16. Krishnan, S., Wang, J., Wu, E., Franklin, M. J., & Goldberg, K. (2016). ActiveClean: Interactive data cleaning for statistical modeling. Proceedings of the VLDB Endowment, 9(12), 948 to 959.
17. Liu, F. T., Ting, K. M., & Zhou, Z. H. (2008). Isolation forest. Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, 413 to 422.
18. Naumann, F. (2014). Data profiling revisited. ACM SIGMOD Record, 42(4), 40 to 49.
19. Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2018). Data lifecycle challenges in production machine learning: A survey. ACM SIGMOD Record, 47(2), 17 to 28.
20. Provost, F., & Fawcett, T. (2001). Robust classification for imprecise environments. Machine Learning, 42(3), 203 to 231.
21. Rahm, E., & Do, H. H. (2000). Data cleaning: Problems and current approaches. IEEE Data Engineering Bulletin, 23(4), 3 to 13.
22. Redman, T. C. (1998). The impact of poor data quality on the typical enterprise. Communications of the ACM, 41(2), 79 to 82.
23. Rekatsinas, T., Chu, X., Ilyas, I. F., & Re, C. (2017). HoloClean: Holistic data repairs with probabilistic inference. Proceedings of the VLDB Endowment, 10(11), 1190 to 1201.
24. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781 to 1794.
25. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503 to 2511.
26. Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5 to 33.