Copilots for Self-Service BI: A Practitioner Comparison of Generative Analytics Assistants across Enterprise Data Platforms
DOI:
https://doi.org/10.15662/IJEETR.2025.0706050Keywords:
generative BI, natural-language query, semantic layer, text-to-SQL, row-level security, self-service analytics, BI competency centerAbstract
Every major business intelligence (BI) and data platform vendor now ships a generative assistant that promises to open analytics to non-specialists through natural language. Microsoft offers Copilot in Fabric and Power BI, SAP offers Joule and the Just Ask experience in SAP Analytics Cloud, and Snowflake offers Cortex Analyst. Many enterprises run several of these platforms side by side, often as a result of acquisitions, and they have little neutral guidance on how the assistants behave on their own governed data. This paper proposes a vendor-neutral evaluation framework and applies it to the three assistants in a manufacturing and finance setting. The framework scores answer accuracy on governed semantic models, sensitivity to semantic-layer maturity, handling of ambiguous business terms, conformance with row-level security, explainability, and cost and licensing behavior. We use 90 business questions drawn from order-to-cash, procurement and general-ledger analysis and spread across four difficulty tiers. Each question runs against a basic and a curated semantic model, together with a cumulative ablation of the curation steps and 216 security probes issued under restricted personas. In representative results, curation raised accuracy from 43–57% to 74–80% for all three anonymized assistants and narrowed the spread between them from 13.3 to 5.6 percentage points. No assistant returned rows outside a persona's entitlement, but all three sometimes presented partial totals as enterprise totals. The central finding is that semantic-model maturity explains more of the answer quality than the choice of assistant. We close with adoption guidance for BI competency centers
References
[1] P. Alpar and M. Schulz, "Self-service business intelligence," Business & Information Systems Engineering, vol. 58, no. 2, pp. 151–155, 2016.
[2] R. Kimball and M. Ross, The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling, 3rd ed. Indianapolis, IN, USA: Wiley, 2013.
[3] H. J. Watson and B. H. Wixom, "The current state of business intelligence," Computer, vol. 40, no. 9, pp. 96–99, 2007.
[4] Microsoft, "Overview of Copilot for Power BI," Microsoft Learn documentation, 2024.
[5] SAP SE, "SAP Analytics Cloud: Just Ask," SAP Help Portal product documentation, 2024.
[6] Snowflake Inc., "Cortex Analyst," Snowflake Documentation, 2024.
[7] Gartner, Inc., "Magic Quadrant for Analytics and Business Intelligence Platforms," Gartner Research, Jun. 2024.
[8] T. Yu et al., "Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task," in Proc. EMNLP, 2018, pp. 3911–3921.
[9] J. Li et al., "Can LLM already serve as a database interface? A BIg bench for large-scale database grounded text-to-SQLs," in Proc. NeurIPS (Datasets and Benchmarks Track), 2023.
[10] N. Rajkumar, R. Li, and D. Bahdanau, "Evaluating the text-to-SQL capabilities of large language models," arXiv:2204.00498, 2022.
[11] M. Pourreza and D. Rafiei, "DIN-SQL: Decomposed in-context learning of text-to-SQL with self-correction," in Proc. NeurIPS, 2023.
[12] D. Gao et al., "Text-to-SQL empowered by large language models: A benchmark evaluation," Proc. VLDB Endowment, vol. 17, no. 5, pp. 1132–1145, 2024.
[13] A. Floratou et al., "NL2SQL is a solved problem... Not!," in Proc. CIDR, 2024.
[14] J. Sequeda, D. Allemang, and B. Jacob, "A benchmark to understand the role of knowledge graphs on large language model's accuracy for question answering on enterprise SQL databases," arXiv:2311.07509, 2023.
[15] OpenAI, "GPT-4 technical report," arXiv:2303.08774, 2023.
[16] Z. Ji et al., "Survey of hallucination in natural language generation," ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023.
[17] H. Chen, R. H. L. Chiang, and V. C. Storey, "Business intelligence and analytics: From big data to big impact," MIS Quarterly, vol. 36, no. 4, pp. 1165–1188, 2012.
[18] Microsoft, "Row-level security (RLS) with Power BI," Microsoft Learn documentation, 2024.
[19] SAP SE, "Joule," SAP Help Portal product documentation, 2024.
[20] Snowflake Inc., "Understanding row access policies," Snowflake Documentation, 2024.
[21] G. Katsogiannis-Meimarakis and G. Koutrika, "A survey on deep learning approaches for text-to-SQL," The VLDB Journal, vol. 32, no. 4, pp. 905–936, 2023.
[22] L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Proc. NeurIPS (Datasets and Benchmarks Track), 2023.
[23] W. H. Inmon, Building the Data Warehouse, 4th ed. Indianapolis, IN, USA: Wiley, 2005.
[24] E. F. Codd, "A relational model of data for large shared data banks," Communications of the ACM, vol. 13, no. 6, pp. 377–387, 1970.
[25] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS, 2020, pp. 9459–9474.
[26] J. Cohen, "A coefficient of agreement for nominal scales," Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960.
[27] E. B. Wilson, "Probable inference, the law of succession, and statistical inference," Journal of the American Statistical Association, vol. 22, no. 158, pp. 209–212, 1927.
[28] Q. McNemar, "Note on the sampling error of the difference between correlated proportions or percentages," Psychometrika, vol. 12, no. 2, pp. 153–157, 1947.





