Prompting, Retrieval, and Fine-Tuning: Foundations of Enterprise Language Model Adaptation

Authors

  • Karan Nisar Independent Researcher, USA Author

DOI:

https://doi.org/10.15662/IJEETR.2024.0606030

Keywords:

Large language models, prompt engineering, retrieval-augmented generation, parameter-efficient fine-tuning, low-rank adaptation, enterprise architecture, model adaptation, in-context learning

Abstract

General-purpose language models must be specialised to an organisation’s documents, policies, vocabulary and output formats before they produce dependable value. Three mechanisms dominate practice: prompt engineering, which conditions a frozen model at inference time; retrieval-augmented generation (RAG), which attaches an external, non-parametric memory to a frozen model; and parameter-efficient fine-tuning, which modifies a small number of weights on curated task data. These are usually described separately, in the vocabulary of the research communities that produced them, which leaves enterprise architects without a common frame in which to compare them. This paper establishes that frame. We review the mechanism, enterprise strengths and enterprise limitations of each technique, together with the hybrid architectures that combine retrieval with weight modification, and we argue that the four resulting archetypes differ along a single structural axis: where the task knowledge required for a correct answer is stored, and which component is mutable at run time. Under prompting, knowledge is parametric and the input is mutable; under RAG, knowledge is external and an index is mutable; under fine-tuning, knowledge and behaviour are parametric and the weights are mutable; under hybrids, both are mutable on different cycles. We show that this one difference, rather than a list of independent trade-offs, accounts for the divergent update latency, citation capability, entitlement enforcement, erasure behaviour and maintenance profile of the four archetypes. The paper is deliberately descriptive: it does not rank the archetypes or prescribe a choice, but supplies the structural vocabulary on which such comparisons and decision procedures can be built

References

[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems 30, 2017, pp. 5998–6008.

[2] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems 33, 2020, pp. 1877–1901.

[3] R. Bommasani, D. A. Hudson, E. Adeli, et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, Aug. 2021.

[4] P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys, vol. 55, no. 9, pp. 1–35, Sep. 2023, doi: 10.1145/3560815.

[5] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems 33, 2020, pp. 9459–9474.

[6] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learning Representations (ICLR), 2022.

[7] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, "Language models are unsupervised multitask learners," OpenAI Tech. Rep., Feb. 2019.

[8] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, "Scaling laws for neural language models," arXiv preprint arXiv:2001.08361, Jan. 2020.

[9] J. Hoffmann, S. Borgeaud, A. Mensch, et al., "Training compute-optimal large language models," in Advances in Neural Information Processing Systems 35, 2022, pp. 30016-30030.

[10] H. Touvron, L. Martin, K. Stone, et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, Jul. 2023.

[11] A. Q. Jiang, A. Sablayrolles, A. Mensch, et al., “Mistral 7B,” arXiv preprint arXiv:2310.06825, Oct. 2023.

[12] J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in Proc. Int. Conf. Learning Representations (ICLR), 2022.

[13] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in Advances in Neural Information Processing Systems 35, 2022, pp. 27730–27744.

[14] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” in Advances in Neural Information Processing Systems 36, 2023, pp. 53728–53741.

[15] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, art. 248, pp. 1–38, Dec. 2023, doi: 10.1145/3571730.

[16] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” arXiv preprint arXiv:2311.05232, Nov. 2023.

[17] S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer, “Rethinking the role of demonstrations: What makes in-context learning work?,” in Proc. 2022 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2022, pp. 11048–11064, doi: 10.18653/v1/2022.emnlp-main.759.

[18] Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh, “Calibrate before use: Improving few-shot performance of language models,” in Proc. 38th Int. Conf. Machine Learning (ICML), PMLR vol. 139, 2021, pp. 12697–12706.

[19] M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr, “Quantifying language models’ sensitivity to spurious features in prompt design, or: How I learned to start worrying about prompt formatting,” in Proc. Int. Conf. Learning Representations (ICLR), 2024.

[20] S. Chen, S. Wong, L. Chen, and Y. Tian, “Extending context window of large language models via positional interpolation,” arXiv preprint arXiv:2306.15595, Jun. 2023.

[21] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 157–173, 2024, doi: 10.1162/tacl_a_00638.

[22] S. Schulhoff, M. Ilie, N. Balepur, K. Kahadze, A. Liu, C. Si, Y. Li, A. Gupta, H. Han, et al., “The Prompt Report: A systematic survey of prompting techniques,” arXiv preprint arXiv:2406.06608, Jun. 2024.

[23] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems 35, 2022, pp. 24824–24837.

[24] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in Proc. Int. Conf. Learning Representations (ICLR), 2023.

[25] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proc. 29th ACM Symp. Operating Systems Principles (SOSP), 2023, pp. 611–626, doi: 10.1145/3600006.3613165.

[26] K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang, “Retrieval augmented language model pre-training,” in Proc. 37th Int. Conf. Machine Learning (ICML), PMLR vol. 119, 2020, pp. 3929–3938.

[27] V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 6769–6781, doi: 10.18653/v1/2020.emnlp-main.550.

[28] G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proc. 16th Conf. European Chapter of the Association for Computational Linguistics (EACL), 2021, pp. 874–880, doi: 10.18653/v1/2021.eacl-main.74.

[29] Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang, “Retrieval-augmented generation for large language models: A survey,” arXiv preprint arXiv:2312.10997, Dec. 2023.

[30] A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to retrieve, generate, and critique through self-reflection,” in Proc. Int. Conf. Learning Representations (ICLR), 2024.

[31] W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W.-t. Yih, “REPLUG: Retrieval-augmented black-box language models,” in Proc. 2024 Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), vol. 1, 2024, pp. 8371–8384, doi: 10.18653/v1/2024.naacl-long.463.

[32] N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive text embedding benchmark,” in Proc. 17th Conf. European Chapter of the Association for Computational Linguistics (EACL), 2023, pp. 2014–2037, doi: 10.18653/v1/2023.eacl-main.148.

[33] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3982–3992, doi: 10.18653/v1/D19-1410.

[34] Y. A. Malkov and D. A. Yashunin, “Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, pp. 824–836, Apr. 2020, doi: 10.1109/TPAMI.2018.2889473.

[35] R. Nogueira and K. Cho, “Passage re-ranking with BERT,” arXiv preprint arXiv:1901.04085, Jan. 2019.

[36] S. Barnett, S. Kurniawan, S. Thudumu, Z. Brannelly, and M. Abdelrazek, “Seven failure points when engineering a retrieval augmented generation system,” in Proc. IEEE/ACM 3rd Int. Conf. AI Engineering: Software Engineering for AI (CAIN), 2024, pp. 194–199, doi: 10.1145/3644815.3644945.

[37] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection,” in Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023, pp. 79–90, doi: 10.1145/3605764.3623985.

[38] N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in Proc. 36th Int. Conf. Machine Learning (ICML), PMLR vol. 97, 2019, pp. 2790–2799.

[39] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Advances in Neural Information Processing Systems 36, 2023, pp. 10088–10115.

[40] C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy, “LIMA: Less is more for alignment,” in Advances in Neural Information Processing Systems 36, 2023, pp. 55006–55021.

[41] Y. Luo, Z. Yang, F. Meng, Y. Li, J. Zhou, and Y. Zhang, “An empirical study of catastrophic forgetting in large language models during continual fine-tuning,” arXiv preprint arXiv:2308.08747, Aug. 2023.

[42] D. Biderman, J. Portes, J. J. Gonzalez Ortiz, M. Paul, P. Greengard, C. Jennings, D. King, S. Havens, V. Chiley, J. Frankle, C. Blakeney, and J. P. Cunningham, “LoRA learns less and forgets less,” arXiv preprint arXiv:2405.09673, May 2024.

[43] N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in Proc. 30th USENIX Security Symposium, 2021, pp. 2633–2650.

[44] N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramèr, and C. Zhang, “Quantifying memorization across neural language models,” in Proc. Int. Conf. Learning Representations (ICLR), 2023.

[45] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, Gaithersburg, MD, USA, Jan. 2023, doi: 10.6028/NIST.AI.100-1.

[46] T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez, “RAFT: Adapting language model to domain specific RAG,” arXiv preprint arXiv:2403.10131, Mar. 2024.

Downloads

Published

2024-11-16

How to Cite

Prompting, Retrieval, and Fine-Tuning: Foundations of Enterprise Language Model Adaptation. (2024). International Journal of Engineering & Extended Technologies Research (IJEETR), 6(6), 9310-9319. https://doi.org/10.15662/IJEETR.2024.0606030