BIRALAB
Conference paper2023

Drug Classification Based on Machine Learning Models with a Combination of Data Binning and SMOTE Technique

Authors

Tran Anh Vu, Tran Minh Hieu, Hoang Thi Mai Linh, Hoang Quang Huy, Pham Thi Viet Huong

BIRALAB members are shown in bold and link to their profile page.

Abstract

Prescribing medications to patients by physicians involves navigating through several intricate and multifaceted stages to determine the optimal and safest pharmaceutical intervention. This necessitates the physical presence of the patient for medical evaluation, resulting in significant temporal and financial costs. To expedite this process and assist both patients and physicians in identifying the most suitable medications, our team has leveraged the capabilities of Artificial Intelligence (AI) to address this challenge. In this investigation, we propose the amalgamation of Data Binning and the Synthetic Minority Over-sampling Technique (SMOTE) in the initial phase of data extraction and processing, preceding its integration into the training framework. The dataset is incorporated into various machine learning models, including Logistic Regression, k-nearest neighbors Classifier, Naive Bayes, Stochastic Gradient Descent Classifier, and Gradient Boosting models. Desired results have been achieved using the drug classification models, particularly through the application of Logistic Regression, which yields an impressive accuracy of 97.2% and an fl-score of 96.6%. Additionally, by employing Gradient Boosting, an accuracy of 96.9% and an fl-score of 96.8% were attained. The machine learning algorithms implemented are different from related research and yield good results by comparison. Our findings substantiate the development of a robust and meaningful drug classification model. Considering these results, our team aspires to expand the application of this model to encompass broader drug categories and more diverse patient groups.