When AI Refactors Code: An Expert-informed Framework for Assessing Long-term Maintainability and Architectural Drift

Authors

  • Sourabh Jhawar

Keywords:

AI-assisted refactoring, Architectural drift, Human-in-the-loop governance, Large language models, Software maintainability, Technical debt

Abstract

Artificial Intelligence (AI), including large language models, is increasingly used to identify, recommend, and implement software refactoring. Existing evaluations primarily emphasize immediate structural outcomes, including reduced complexity, improved cohesion, increased testability, and removal of code smells. However, limited guidance exists for determining whether such improvements remain beneficial across multiple development cycles or contribute to architectural drift, technical debt, and knowledge loss. This study develops an expert-informed framework for assessing long-term maintainability in AI-refactored software systems. A focused narrative synthesis of literature on software refactoring, maintainability, architectural erosion, software evolution, and AI-assisted software engineering informed a preliminary conceptual model. The model was subsequently refined through semi-structured interviews with three senior practitioners representing quality engineering, enterprise architecture, cloud infrastructure, and production operations. Hybrid deductive-inductive thematic analysis produced four integrated findings: long-term maintainability is multidimensional and temporal; repeated context-limited optimization may accumulate into architectural drift; assessment requires structural, architectural, contextual, governance, and operational evidence; and expert knowledge should be embedded through persistent architectural context, machine-enforceable controls, risk-based review, and operational feedback. These findings informed the AI Refactoring Maintainability and Drift Assessment Framework (AI-RMDAF), which links an approved system baseline, multidimensional assessment, risk classification, deployment monitoring, and continuous organizational learning. The study contributes an exploratory conceptual framework that provides a foundation for future empirical and longitudinal validation rather than a statistically or predictively validated model. It offers researchers a structured agenda for longitudinal evaluation and provides practitioners with a governance-oriented approach for preserving architectural intent while retaining the productivity advantages of AI-assisted refactoring.

References

M. Fowler, Refactoring: Improving the Design of Existing Code, 2nd ed. Boston, MA, USA: Addison-Wesley Professional, 2018.

T. Mens and T. Tourwé, “A survey of software refactoring,” IEEE Transactions on Software Engineering, vol. 30, no. 2, pp. 126–139, Feb. 2004

X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large language models for software engineering: A systematic literature review,” ACM Transactions on Software Engineering and Methodology, vol. 33, no. 8, Art. no. 220, 2024

B. Liu, Y. Jiang, Y. Zhang, N. Niu, G. Li, and H. Liu, “Exploring the potential of general purpose LLMs in automated software refactoring: An empirical study,” Automated Software Engineering, vol. 32, Art. no. 26, 2025.

S. R. Chidamber and C. F. Kemerer, “A metrics suite for object-oriented design,” IEEE Transactions on Software Engineering, vol. 20, no. 6, pp. 476–493, Jun. 1994.

T. J. McCabe, “A complexity measure,” IEEE Transactions on Software Engineering, vol. SE-2, no. 4, pp. 308–320, Dec. 1976.

E. Murphy-Hill, C. Parnin, and A. P. Black, “How we refactor, and how we know it,” IEEE Transactions on Software Engineering, vol. 38, no. 1, pp. 5–18, Jan.–Feb. 2012.

D. Silva, N. Tsantalis, and M. T. Valente, “Why we refactor? Confessions of GitHub contributors,” In Proceedings of the 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering, Seattle, WA, USA, 2016, pp. 858–870.

L. de Silva and D. Balasubramaniam, “Controlling software architecture erosion: A survey,” Journal of Systems and Software, vol. 85, no. 1, pp. 132–151, Jan. 2012.

J. Van Gurp and J. Bosch, “Design erosion: Problems and causes,” Journal of Systems and Software, vol. 61, no. 2, pp. 105–119, 2002.

N. Tsantalis, M. Mansouri, L. M. Eshkevari, D. Mazinanian, and D. Dig, “Accurate and efficient refactoring detection in commit history,” In Proceedings of the 40th International Conference on Software Engineering, Gothenburg, Sweden, 2018, pp. 483–494.

A. A. B. Baqais and M. Alshayeb, “Automatic software refactoring: A systematic literature review,” Software Quality Journal, vol. 28, no. 2, pp. 459–502, 2020.

ISO/IEC 25010:2023, Systems and Software Engineering—Systems and Software Quality Requirements and Evaluation (SQuaRE)—Product Quality Model. Geneva, Switzerland: International Organization for Standardization and International Electrotechnical Commission, 2023.

N. Nagappan and T. Ball, “Use of relative code churn measures to predict system defect density,” In Proceedings of the 27th International Conference on Software Engineering, St. Louis, MO, USA, 2005, pp. 284–292.

P. Kruchten, R. L. Nord, and I. Ozkaya, “Technical debt: From metaphor to theory and practice,” IEEE Software, vol. 29, no. 6, pp. 18–21, Nov.–Dec. 2012.

ISO/IEC 42001:2023, Information Technology—Artificial Intelligence—Management System. Geneva, Switzerland: International Organization for Standardization and International Electrotechnical Commission, 2023.

K. Malterud, V. D. Siersma, and A. D. Guassora, “Sample size in qualitative interview studies: Guided by information power,” Qualitative Health Research, vol. 26, no. 13, pp. 1753–1760, Nov. 2016.

H. Kallio, A.-M. Pietilä, M. Johnson, and M. Kangasniemi, “Systematic methodological review: Developing a framework for a qualitative semi-structured interview guide,” Journal of Advanced Nursing, vol. 72, no. 12, pp. 2954–2965, Dec. 2016.

V. Braun and V. Clarke, “Using thematic analysis in psychology,” Qualitative Research in Psychology, vol. 3, no. 2, pp. 77–101, 2006.

Published

2026-07-31

How to Cite

Jhawar, S. (2026). When AI Refactors Code: An Expert-informed Framework for Assessing Long-term Maintainability and Architectural Drift. Journal of Computer Based Parallel Programming, 11(2), 50–65. Retrieved from https://www.matjournals.net/engineering/index.php/JoCPP/article/view/3926

Issue

Section

Articles