Generative AI for Data Mapping: Evaluating AI-Assisted Approaches to Legacy-to-S/4HANA Field Reconciliation
DOI:
https://doi.org/10.15662/IJEETR.2025.0705021Keywords:
generative artificial intelligence, large language models, schema matching, entity reconciliation, retrieval-augmented generation, SAP S/4HANA, ERP data migration, data governance, embeddings, human-in-the-loop review.Abstract
Converting a legacy enterprise resource planning landscape to SAP S/4HANA requires reconciling tens of thousands of source fields against a simplified, restructured target data model that consolidates entities such as customer and vendor master records into a unified business partner model and replaces multiple finance tables with a single universal journal. Field-level reconciliation, deciding which target field a given source field should map to, what transformation it requires, and whether it should be retained at all, has traditionally been a slow, manual, and expertise-intensive exercise performed by functional consultants. This article evaluates whether generative artificial intelligence, specifically large language models combined with semantic embeddings and retrieval-augmented generation, can materially accelerate this process without compromising mapping quality or introducing unacceptable risk
We propose a reference architecture for AI-assisted field reconciliation that combines embedding-based candidate generation, large language model reasoning over field context, calibrated confidence scoring, and a human-in-the-loop review workflow governed by an auditable mapping catalog. We compare this hybrid retrieval-augmented approach against four baselines, rule-based name matching, classic schema matching, embedding similarity alone, and unaided large language model prompting, on a simulated benchmark spanning five common conversion modules, including finance, materials management, sales and distribution, business partner, and custom Z-fields
Across the benchmark, the hybrid retrieval-augmented approach achieved an F1 score of 0.91 for mapping correctness, compared with 0.53 for rule-based matching, while reducing human review effort from an initial 42 hours per 1,000 fields to 5 hours per 1,000 fields by the eighth simulated migration wave. Critically, the hybrid approach also reduced the rate of confidently incorrect, or hallucinated, mappings to under 2.5 percent across every module tested, compared with rates above 14 percent for unaided prompting on the most structurally complex custom field category. These results suggest that generative AI can substantially reduce reconciliation effort, provided it is grounded in retrieved organizational context and paired with calibrated confidence scoring and governed human review, rather than deployed as an unaided, free-form reasoning step
We close with a discussion of design trade-offs, data privacy and model deployment considerations specific to prompting proprietary schema information, current limitations, threats to validity, and directions for future research, including active learning for mapping catalog growth and standardized field reconciliation benchmarks.
References
1. Bernstein, P. A., Madhavan, J., & Rahm, E. (2011). Generic schema matching, ten years later. Proceedings of the VLDB Endowment, 4(11), 695 to 701.
2. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877 to 1901.
3. Cappuzzo, R., Papotti, P., & Thirumuruganathan, S. (2020). Creating embeddings of heterogeneous relational datasets for data integration tasks. Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, 1335 to 1349.
4. Davenport, T. H. (1998). Putting the enterprise into the enterprise system. Harvard Business Review, 76(4), 121 to 131.
5. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, 4171 to 4186.
6. Do, H. H., & Rahm, E. (2002). COMA: A system for flexible combination of schema matching approaches. Proceedings of the 28th International Conference on Very Large Data Bases, 610 to 621.
7. Doan, A., Madhavan, J., Domingos, P., & Halevy, A. (2002). Learning to map between ontologies on the semantic web. Proceedings of the 11th International World Wide Web Conference, 662 to 673.
8. Fernandez, R. C., Abedjan, Z., Koko, F., Yuan, G., Madden, S., & Stonebraker, M. (2018). Aurum: A data discovery system. Proceedings of the 2018 IEEE 34th International Conference on Data Engineering, 1001 to 1012.
9. Haller, K. (2009). Towards the industrialization of data migration: Concepts and patterns for standard software implementation projects. Lecture Notes in Business Information Processing, CAiSE Forum 2009.
10. Hulsebos, M., Hu, K., Bakker, M., Zgraggen, E., Satyanarayan, A., Kraska, T., Demiralp, C., & Hidalgo, C. (2019). Sherlock: A deep learning approach to semantic data type detection. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1500 to 1508.
11. Koutras, C., Siachamis, G., Ionescu, A., Psarakis, K., Brons, J., Fragkoulis, M., Lofi, C., Bonifati, A., & Katsifodimos, A. (2021). Valentine: Evaluating matching techniques for dataset discovery. Proceedings of the 2021 IEEE 37th International Conference on Data Engineering, 468 to 479.
12. Li, Y., Li, J., Suhara, Y., Doan, A., & Tan, W. C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1), 50 to 60.
13. Loshin, D. (2010). Master data management. Morgan Kaufmann.
14. Madhavan, J., Bernstein, P. A., & Rahm, E. (2001). Generic schema matching with Cupid. Proceedings of the 27th International Conference on Very Large Data Bases, 49 to 58.
15. Miller, R. J. (2018). Open data integration. Proceedings of the VLDB Endowment, 11(12), 2130 to 2139.
16. Narayan, A., Chami, I., Orr, L., & Re, C. (2022). Can foundation models wrangle your data? Proceedings of the VLDB Endowment, 16(4), 738 to 746.
17. Nargesian, F., Zhu, E., Miller, R. J., Pu, K. Q., & Arocena, P. C. (2019). Data lake management: Challenges and opportunities. Proceedings of the VLDB Endowment, 12(12), 1986 to 1989.
18. Peeters, R., & Bizer, C. (2023). Using ChatGPT for entity matching. arXiv preprint arXiv:2305.03423.
19. Rahm, E., & Bernstein, P. A. (2001). A survey of approaches to automatic schema matching. VLDB Journal, 10(4), 334 to 350.
20. Rahm, E., & Do, H. H. (2000). Data cleaning: Problems and current approaches. IEEE Data Engineering Bulletin, 23(4), 3 to 13.
21. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3982 to 3992.
22. Suhara, Y., Li, Y., Li, Y., Zhang, D., Demiralp, C., Chen, C., & Tan, W. C. (2022). Annotating columns with pre-trained language models. Proceedings of the 2022 International Conference on Management of Data, 1493 to 1503.
23. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998 to 6008.
24. Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5 to 33.





