SP-CARE: A Risk-Aware Framework for Continuous System Prompt Evaluation and Adaptive Revalidation in Agentic AI Systems
DOI:
https://doi.org/10.15662/IJEETR.2026.0805004Keywords:
system prompt evaluation, large language models, agentic AI, LLM-as-a-judge, prompt robustness, behavioral testing, AI governance, model regression, tool-use evaluation, risk-aware evaluationAbstract
System prompts have become a consequential behavioral control layer in large language model applications and agentic AI systems. They shape instruction adherence, task completion, tool selection, verification, safety boundaries, escalation, style, and resource use. Yet system-prompt changes are commonly assessed with a small number of manually selected examples, static benchmarks, or uncalibrated model-based judgments. These practices provide limited visibility into rare but severe failures, interactions among instructions, prompt-model coupling, non-deterministic behavior, tool-mediated risk, and regressions introduced by model or harness upgrades. This paper proposes SP-CARE, a System Prompt Continuous Assessment and Risk Evaluation framework that treats the system prompt as a versioned production control artifact. SP-CARE combines requirement decomposition, risk-based test portfolio construction, controlled repeated experiments, multi-source evaluation, release gating, production monitoring, and adaptive revalidation. It defines a multidimensional taxonomy covering instruction fidelity, task effectiveness, claim reliability, tool behavior, verification discipline, safety, recovery, robustness, user experience, and operational efficiency. The framework also introduces a risk-adjusted decision index that penalizes execution variability, tail-failure severity, cost, and latency while enforcing non-compensatory safety gates. A reproducible empirical protocol is specified for comparing SP-CARE with static benchmark and single-judge baselines across prompt variants, model configurations, tool environments, and repeated runs. The intended contribution is not a claim of completed empirical superiority, but a falsifiable framework and research design for determining whether a system-prompt change is reliable enough to release and portable enough to survive system evolution.References
1. M. Kim, P. Chapman, S. S. Somarouthu, J. Agrawal, and M. K. Ramanathan, "Continuous Prompt Evaluation: How We Use LLM Judges and Live Signals to Improve Kiro Agent Quality," Kiro, Aug. 21, 2026. Available: https://kiro.dev/blog/continuous-prompt-evaluation/
2. P. Liang et al., "Holistic Evaluation of Language Models," arXiv:2211.09110, 2023.
3. M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh, "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList," in Proc. 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 4902-4912, doi: 10.18653/v1/2020.acl-main.442.
4. M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I Learned to Start Worrying about Prompt Formatting," in Proc. 12th International Conference on Learning Representations, 2024.
5. K. Zhu et al., "PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts," arXiv:2306.04528, 2024.
6. T. P. Zollo, T. Morrill, Z. Deng, J. C. Snell, T. Pitassi, and R. Zemel, "Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models," in Proc. 12th International Conference on Learning Representations, 2024.
7. L. Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," in Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 2023.
8. Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, "G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment," in Proc. 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2511-2522, doi: 10.18653/v1/2023.emnlp-main.153.
9. H. Hashemi, J. Eisner, C. Rosset, B. Van Durme, and C. Kedzie, "LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts," arXiv:2501.00274, 2024.
10. L. Shi, C. Ma, W. Liang, W. Ma, and S. Vosoughi, "Judging the Judges: A Systematic Investigation of Position Bias in Pairwise Comparative Assessments by LLMs," arXiv:2406.07791, 2024.
11. X. Liu et al., "AgentBench: Evaluating LLMs as Agents," in Proc. 12th International Conference on Learning Representations, 2024.
12. S. Yao, N. Shinn, P. Razavi, and K. Narasimhan, "Tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," arXiv:2406.12045, 2024.
13. E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramer, "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents," in Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 2024.
14. K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, "Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection," arXiv:2302.12173, 2023.
15. B. Wang et al., "DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models," in Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 2023.
16. C. Autio, R. Schwartz, J. Dunietz, S. Jain, M. Stanley, E. Tabassi, P. Hall, and K. Roberts, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, National Institute of Standards and Technology, 2024, doi: 10.6028/NIST.AI.600-1.
17. R. T. Rockafellar and S. Uryasev, "Optimization of Conditional Value-at-Risk," Journal of Risk, vol. 2, no. 3, pp. 21-42, 2000, doi: 10.21314/JOR.2000.038.
18. Kumar, S. N. P. (2026). Advanced architectural frameworks for scalable, production-grade agentic RAG pipelines. International Journal of Research and Applied Innovations, 9(1), 13491–13498. https://doi.org/10.15662/IJRAI.2026.0901001
19. Kumar, S. N. P. (2025). A secure accountability framework for multi-modal agent systems: Detecting, mitigating, and auditing data-poisoning attacks via Model Context Protocol servers. Journal of Computer Science and Technology Studies, 7(12), 1–5. https://doi.org/10.32996/jcsts.2025.7.12.1
20. Kumar, S. N. P. (2025). Building scalable and reliable agentic AI systems: A technical blueprint for autonomous intelligence. Global Journal of Engineering and Technology Research, 1(3). https://doi.org/10.65150/ep-gjetr/v1e3/2025-07
21. Kumar, S. N. P. (2025). Regulating autonomous AI agents: Prospects, hazards, and policy structures. Journal of Computer Science and Technology Studies, 7(10), 393–399. https://doi.org/10.32996/jcsts.2025.7.10.41
22. Kumar, S. N. P. (2025). Hallucination detection and mitigation in large language models: A comprehensive review. Journal of Information Systems Engineering and Management, 10(60S), 435–443. https://doi.org/10.52783/jisem.v10i60s.13133
23. Kumar, S. N. P. (2025). Multi-agent AI systems in finance: Models, applications, and challenges. International Journal of Advanced Research in Computer Science & Technology, 8(1), 11555–11573. https://doi.org/10.15662/IJARCST.2025.0801006
24. Kumar, S. N. P. (2025). Scalable cloud architectures for AI-driven decision systems. Journal of Computer Science and Technology Studies, 7(8), 416–421. https://doi.org/10.32996/jcsts.2025.7.8.46
25. Kumar, S. N. P. (2025). Recent innovations in cloud-optimized retrieval-augmented generation architectures for AI-driven decision systems. European Modern Studies Journal, 9(4), 870–883. https://doi.org/10.59573/emsj.9(4).2025.81
26. Kumar, S. N. P., Gangurde, R., & Mohite, U. L. (2025). RMHAN: Random multi-hierarchical attention network with RAG-LLM-based sentiment analysis using text reviews. International Journal of Computational Intelligence and Applications, 24(4). https://doi.org/10.1142/S1469026825500075
27. Kumar, S. N. P. (2025). Ethical frameworks for AI-driven decision systems: A comprehensive analysis. Global Journal of Computer Science and Technology, 25(1), 53–60. https://doi.org/10.34257/gjcstdvol25is1pg53
28. Kumar, S. N. P. (2025). Real-time justice intelligence: AI-enabled data governance transforming city criminal justice. Global Journal of Engineering and Technology Research, 1(3). https://doi.org/10.65150/ep-gjetr/v1e3/2025-08
29. Kumar, S. N. P. (2025). Fraud detection in banking using generative AI. Sarcouncil Journal of Engineering and Computer Sciences, 4(11), 133–145. Publisher record
30. Kumar, S. N. P. (2025). AI and cloud data engineering transforming healthcare decisions. Sarcouncil Journal of Engineering and Computer Sciences, 4(8), 76–82. Publisher record
31. Kumar, S. N. P. (2025). Quantum-enhanced AI decision systems: Architectural approaches for cloud-based machine learning applications. Sarcouncil Journal of Multidisciplinary, 5(8). Publisher record
32. Kumar, S. N. P. (2025). Navigating the AI horizon: Transformations, ethical imperatives, and pathways to responsible innovation. Sarcouncil Journal of Applied Sciences, 5(10), 34–43. Publisher record
33. Choudhury, A., Balasubramaniam, S., Kumar, A. P., & Kumar, S. N. P. (2025). PSSO: Political squirrel search optimizer-driven deep learning for severity-level detection and classification of lung cancer. International Journal of Information Technology & Decision Making, 24(8), 2373–2406. https://doi.org/10.1142/S0219622023500189
34. Preetham, A., Vyas, S., Kumar, M., & Kumar, S. N. P. (2024). Optimized convolutional neural network for land-cover classification via improved lion algorithm. Transactions in GIS, 28(4), 769–789. https://doi.org/10.1111/tgis.13150
35. Mahalakshmi, T., Beevi, S. Z., Navaneethakrishnan, M., Ramya, P., & Kumar, S. N. P. (2024). Optimized attention-driven bidirectional convolutional neural network for Facebook sentiment classification. International Journal of Business Data Communications and Networking, 19(1), 1–20. https://doi.org/10.4018/IJBDCN.349572
36. Choudhury, A., Vuppu, S., Singh, S. P., Kumar, M., & Kumar, S. N. P. (2023). ECG-based heartbeat classification using exponential-political optimizer-trained deep learning for arrhythmia detection. Biomedical Signal Processing and Control, 84, 104816. https://doi.org/10.1016/j.bspc.2023.104816
37. Masthan, M., Pazhanikumar, K., Chavan, M., Mandala, J., & Kumar, S. N. P. (2023). SCSLnO-SqueezeNet: Sine cosine-sea lion optimization-enabled SqueezeNet for intrusion detection in IoT. Network: Computation in Neural Systems, 34(4), 343–373. https://doi.org/10.1080/0954898X.2023.2261531
38. Nalavade, J. E., Kolli, C. S., & Kumar, S. N. P. (2023). Deep embedded clustering with matrix factorization-based user-rating prediction for collaborative recommendation. Multiagent and Grid Systems, 19(2), 169–185. https://doi.org/10.3233/MGS-230039
39. Boopathi, M., Chavan, M., Jebanazer, J. J., & Kumar, S. N. P. (2023). An approach for DoS attack detection in cloud computing using sine cosine anti-coronavirus optimized deep Maxout network. International Journal of Pervasive Computing and Communications, 19(5), 666–688. https://doi.org/10.1108/IJPCC-05-2022-0197
40. Kumar, S. N. P. (2022). Improving fraud detection in credit-card transactions using autoencoders and deep neural networks [Doctoral dissertation, The George Washington University]. GW ScholarSpace. https://scholarspace.library.gwu.edu/concern/gw_etds/cv43nx607
41. Kumar, S. N. P. (2022). Text classification: A comprehensive survey of methods, applications, and future directions. International Journal of Technology, Management and Humanities, 8(3), 39–49. https://doi.org/10.21590/ijtmh.8.03.04
42. Kumar, S. N. P. (2022). Machine learning regression techniques for modeling complex industrial systems: A comprehensive summary. International Journal of Humanities and Information Technology, 4(1–3), 67–79. https://doi.org/10.21590/ijhit.06.01-3.07





