RAG Pipeline Accuracy Benchmarking in Healthcare Supply Chain: Medical Item Classification Tasks
DOI:
https://doi.org/10.22105/kmisj.vi.129Keywords:
Retrieval-augmented generation, Healthcare Supply Chain, Medical item classification, Dense retrieval, BM25, Large language models, Retrieval-augmented generation evaluationAbstract
The range of medical products within healthcare Supply Chains is extensive and must be precisely classified into hierarchical categories for purchasing, inventory control, expenditure analysis, and regulatory purposes. Traditional classification methods may not suit the short, mixed nature of product descriptions, or they may produce undue or hallucinated labels (e.g., Large Language Models (LLMs)). Retrieval-Augmented Generation (RAG) introduces a way to ground LLM outputs using an external source of domain knowledge. This work proposes a benchmark for RAG-based pipelines for the hierarchical classification of medical supplies. The performance evaluation of a three-level taxonomy with 214 leaf-level classification nodes was conducted on a benchmark dataset comprising 950 descriptions of medical supply items. We evaluated four pipelining settings by combining dense and BM25 retrieval with two instruction-tuned generation models. Retrieval quality was assessed with MAP@5, Recall@5, NDCG@5, Context Relevance Score (CRS), and Context Coverage (CCO), and the E2E quality was measured by hierarchical classification accuracy, macro-F1, and Answer Fidelity (AF). Dense retrieval outperformed all other metrics in terms of retrieval scores. The dense-retrieval models also yield higher hierarchical classification accuracy, with the best configuration achieving 87.3% for the subcategory. The MAP@5 is strongly correlated with subcategory accuracy, as indicated by partial correlation analysis. A human assessment of 240 misclassified queries revealed context dilution, label hallucination, retrieval collapse, and formatting inconsistency as the main failure patterns. The results provide a rigorous framework for assessing and refining RAG pipelines for domain-specific health care supply classification.
References
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., ... & Fung, P. (2023). Survey of hallucination in natural language generation. ACM computing surveys, 55(12), 1-38. https://doi.org/10.1145/3571730
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in neural information processing systems (Vol. 33, pp. 9459–9474). NeurIPS. https://proceedings.neurips.cc/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., ... & Zhou, D. (2023). Large language models can be easily distracted by irrelevant context. International conference on machine learning (pp. 31210-31227). PMLR. https://proceedings.mlr.press/v202/shi23a.html
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., ... & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. https://arxiv.org/abs/2312.10997
Pavate, A., Mourya, V., Digra, M., Kumavat, K., & Jalalov, M. (2025). Unsupervised vs. supervised RAG: A comparative study of retrieval-augmented language models and their output variability. International conference on intelligent optimization and big data management (pp. 418-431). Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-10400-7_37
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., ... & Grave, E. (2023). Atlas: Few-shot learning with retrieval augmented language models. Journal of machine learning research, 24(251), 1-43. https://www.jmlr.org/papers/v24/23-0037.html
Manning, C. D., Raghavan, P., & Schütze, H. (2008). XML retrieval. Introduction to information retrieval, 1403, 1404. https://www.internetdownloadmanager.com/register/new_faq/chrome_extension2.html
Es, S., James, J., Anke, L. E., & Schockaert, S. (2024). Ragas: Automated evaluation of retrieval augmented generation. Proceedings of the 18th conference of the european chapter of the association for computational linguistics: System demonstrations (pp. 150-158). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2024.eacl-demo.16
Espinosa, O., Basto, S., Ordóñez, A., Arias, M. L., Sepúlveda, D., & Bhutta, M. F. (2026). A review of artificial intelligence in healthcare Supply Chains: Untapped potential?. BMC health services research, 26(742). https://doi.org/10.1186/s12913-026-14559-2
US., G. H. (2022). Implementation guideline: Applying the GS1 system of standards for U.S. FDA unique device identification (UDI). https://documents.gs1us.org/.../Implementation-Guideline-Using-the-GS1-System-for-US-FDA-UDI-R
Roberson, A. (2021). Applying machine learning for automatic product categorization. Journal of official statistics, 37(2), 395–410. https://doi.org/10.2478/jos-2021-0017
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2025 Karshi Multidisciplinary International Scientific Journal

This work is licensed under a Creative Commons Attribution 4.0 International License.