RAG Pipeline Accuracy Benchmarking in Healthcare Supply Chain: Medical Item Classification Tasks

Authors

DOI:

https://doi.org/10.22105/kmisj.vi.129

Keywords:

Retrieval-augmented generation, Healthcare Supply Chain, Medical item classification, Dense retrieval, BM25, Large language models, Retrieval-augmented generation evaluation

Abstract

The range of medical products within healthcare Supply Chains is extensive and must be precisely classified into hierarchical categories for purchasing, inventory control, expenditure analysis, and regulatory purposes. Traditional classification methods may not suit the short, mixed nature of product descriptions, or they may produce undue or hallucinated labels (e.g., Large Language Models (LLMs)). Retrieval-Augmented Generation (RAG) introduces a way to ground LLM outputs using an external source of domain knowledge. This work proposes a benchmark for RAG-based pipelines for the hierarchical classification of medical supplies. The performance evaluation of a three-level taxonomy with 214 leaf-level classification nodes was conducted on a benchmark dataset comprising 950 descriptions of medical supply items. We evaluated four pipelining settings by combining dense and BM25 retrieval with two instruction-tuned generation models. Retrieval quality was assessed with MAP@5, Recall@5, NDCG@5, Context Relevance Score (CRS), and Context Coverage (CCO), and the E2E quality was measured by hierarchical classification accuracy, macro-F1, and Answer Fidelity (AF). Dense retrieval outperformed all other metrics in terms of retrieval scores. The dense-retrieval models also yield higher hierarchical classification accuracy, with the best configuration achieving 87.3% for the subcategory. The MAP@5 is strongly correlated with subcategory accuracy, as indicated by partial correlation analysis. A human assessment of 240 misclassified queries revealed context dilution, label hallucination, retrieval collapse, and formatting inconsistency as the main failure patterns. The results provide a rigorous framework for assessing and refining RAG pipelines for domain-specific health care supply classification. 

References

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., ... & Fung, P. (2023). Survey of hallucination in natural language generation. ACM computing surveys, 55(12), 1-38. https://doi.org/10.1145/3571730

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in neural information processing systems (Vol. 33, pp. 9459–9474). NeurIPS. https://proceedings.neurips.cc/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf

Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., ... & Zhou, D. (2023). Large language models can be easily distracted by irrelevant context. International conference on machine learning (pp. 31210-31227). PMLR. https://proceedings.mlr.press/v202/shi23a.html

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., ... & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. https://arxiv.org/abs/2312.10997

Pavate, A., Mourya, V., Digra, M., Kumavat, K., & Jalalov, M. (2025). Unsupervised vs. supervised RAG: A comparative study of retrieval-augmented language models and their output variability. International conference on intelligent optimization and big data management (pp. 418-431). Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-10400-7_37

Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., ... & Grave, E. (2023). Atlas: Few-shot learning with retrieval augmented language models. Journal of machine learning research, 24(251), 1-43. https://www.jmlr.org/papers/v24/23-0037.html

Manning, C. D., Raghavan, P., & Schütze, H. (2008). XML retrieval. Introduction to information retrieval, 1403, 1404. https://www.internetdownloadmanager.com/register/new_faq/chrome_extension2.html

Es, S., James, J., Anke, L. E., & Schockaert, S. (2024). Ragas: Automated evaluation of retrieval augmented generation. Proceedings of the 18th conference of the european chapter of the association for computational linguistics: System demonstrations (pp. 150-158). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2024.eacl-demo.16

Espinosa, O., Basto, S., Ordóñez, A., Arias, M. L., Sepúlveda, D., & Bhutta, M. F. (2026). A review of artificial intelligence in healthcare Supply Chains: Untapped potential?. BMC health services research, 26(742). https://doi.org/10.1186/s12913-026-14559-2

US., G. H. (2022). Implementation guideline: Applying the GS1 system of standards for U.S. FDA unique device identification (UDI). https://documents.gs1us.org/.../Implementation-Guideline-Using-the-GS1-System-for-US-FDA-UDI-R

Roberson, A. (2021). Applying machine learning for automatic product categorization. Journal of official statistics, 37(2), 395–410. https://doi.org/10.2478/jos-2021-0017

Published

2026-06-10

How to Cite

Nasir, S. M. ., Pavate, A. ., & Yusuf, A. . (2026). RAG Pipeline Accuracy Benchmarking in Healthcare Supply Chain: Medical Item Classification Tasks. Karshi Multidisciplinary International Scientific Journal, 3(2), 153-164. https://doi.org/10.22105/kmisj.vi.129

Similar Articles

1-10 of 22

You may also start an advanced similarity search for this article.