Artificial Intelligence in Breast Tumor Classification: A Comparative Study of Logistic Regression, SVM, and KNN Algorithms Using the WBCD Dataset
DOI:
https://doi.org/10.22153/kej.2026.02.003Keywords:
Breast tumor classification; support vector machine (SVM); logistic regression; K-nearest neighbors (KNN); artificial intelligence (AI); machine learning (ML); wisconsin breast cancer dataset (WBCD)Abstract
Breast cancer as such, is one of the two fallaciously debilitating diseases known to woman (the other being cardiac disease) so in that context and with some qualifications it really shouldn't come as that much of a surprise. The concern, however, is: Yes, you could do it right — early detection is our nearest weapon to the Silent Killer (cancer). However, standard diagnostics are far from 100% accurate, can be mislead by virtually any kind of fake and often miss occasions that cry for assistance but which have not yet proceeded to the malignancy that will eventually confirm. More recently you would see the triggers of real momentum with machine learning(ML) one of the last great fortresses where the majority of frontier artificial intelligence (AI) resides. ML can also be leveraged to guide the accuracy and precision of a breast cancer diagnosis when considering a dataset carefully that includes patient data.
The aim in this study was to compare three popular supervised ML algorithms on Logistic Regression (LR), Support Vector Machines (SVM) and K-Nearest Neighbors (KNN). In this domain, the main objective was to effectively categorize malignant breast depending on the benign breast tissue of widely used Wisconsin Breast Cancer Dataset (WBCD) features. An accurate extraction and integration of records belonging to 699 patients was performed for data formatting suitable for analysis. All the required cleaning of data was performed first. This is key to building reliable and credible outcomes, we standardized our functions in a thoroughgoing manner before feeding the model.
For feature selection, Gain Ratio is used. Most data points can be covered in a fair manner and performance has to be squeezed up fully. It is an old-school style but it fits our case perfectly: we keep about 30% of records as test samples and the rest (70%) for training (so, we split by 70/30). With each algorithm you performed a full 10-fold cross validation. No stone unturned! We took a fairly wide range of metrics. You covered sensitivity and specificity but what got the most play was with benchmarked net accuracy rates, much less so about false positives/negatives. When we went to so much trouble with other parts of it, it seems an odd time to write off Precision or AUC-ROC as adjuncts to fill in this sector of the picture.
They made use of ANOVA to provide statistical comparison among those sets of algorithms that were tested against the data — with a McNemar's test functioning as an objective measuring stick. Now the exciting part began, SVM model absolutely jumped to end up at amazing 97.1% accuracy and 0.994 on AUC score! Logistic Regression followed, with 96.4% accuracy(AUC=0.992), KNN takes the third place with a combined score of 95.6%(AUC=0.978).
Downloads
References
[1] World Health Organization, "WHO report on cancer: setting priorities, investing wisely and providing care for all," Geneva: WHO, 2020. https://books.google.iq/books?hl=en&lr=&id=anYOEQAAQBAJ&oi=fnd&pg=PA6&dq=%5B1%5D%09+World+Health+Organization,+%22WHO+report+on+cancer:+setting+priorities,+investing+wisely+and+providing+care+for+all,%22+Geneva:+WHO,+2020.&ots=N3K_sCB6IG&sig=7MI8y9ig07gMlZ_aMSZneWW5VQA&redir_esc=y#v=onepage&q&f=false
[2] N. Azamjah, Y. Soltan-Zadeh, and F. Zayeri, "Global Trend of Breast Cancer Mortality Rate: A 25-Year Study," Asian Pac J Cancer Prev, vol. 20, no. 7, pp. 2015–2020, 2019. https://doi.org/10.31557/APJCP.2019.20.7.2015
[3] J. G. Elmore et al., "Variability in interpretive performance at screening mammography and radiologists' characteristics," J Natl Cancer Inst, vol. 101, no. 5, pp. 327–337, 2009. https://doi.org/10.1148/radiol.2533082308
[4] J. C. van Zelst et al., "Multireader study on the diagnostic accuracy of ultrafast breast magnetic resonance imaging for breast cancer screening," Invest. Radiol., vol. 53, no. 10, pp. 579–586, Oct. 2018. https://doi.org/10.1097/RLI.0000000000000494
[5] Baimukashev, Rashid, Shirali Kadyrov, and Cemil Turan. "Systematic Survey of Deep Fuzzy Computer Vision in Biomedical Research." Fuzzy Information and Engineering 16.3 (2024): 220-243.https://doi.org/10.26599/FIE.2024.9270043
[6] T. Saba et al., "Breast cancer detection using machine learning and convolutional neural networks: a comparative analysis," PeerJ, vol. 7, p. e6201, 2019, doi: 10.7717/peerj.21576
PeerJ.
[7] Nemade, Varsha, Sunil Pathak, and Ashutosh Kumar Dubey. "Deep learning-based ensemble model for classification of breast cancer." Microsystem Technologies 30.5 (2024): 513-527. https://doi.org/10.1007/s00542-023-05469-y
[8] Palarimath, Suresh, et al. "Empowering Breast Cancer Detection with AI: A Modified Support Vector Machine Approach for Improved Classification Accuracy." 2024 International Conference on Expert Clouds and Applications (ICOECA). IEEE, 2024. https://doi.org/10.1109/ICOECA62351.2024.00159
[9] A. Esteva et al., "Dermatologist-level classification of skin cancer with deep neural networks," Nature, vol. 542, pp. 115–118, 2017. https://doi.org/10.1038/nature21056
[10] M. A. Khan et al., "Multimodal brain tumor classification using deep learning and robust feature selection: A machine learning application for radiologists," Diagnostics, vol. 10, no. 8, p. 565, 2020. https://doi.org/10.3390/diagnostics10080565
[11] M. A. Mohammed et al., "Evaluating the performance of machine learning techniques in the classification of Wisconsin Breast Cancer," Int J Eng Technol, vol. 7, no. 4, pp. 160–166, 2018. https://doi.org/10.14419/ijet.v7i4.36.23737
[12] R. Kaifi, "A review of recent advances in brain tumor diagnosis based on AI-based classification," Diagnostics, vol. 13, no. 18, p. 3007, 2023. https://doi.org/10.3390/diagnostics13183007
[13] S. Sharma and R. Mehra, "Conventional machine learning and deep learning approach for multi-classification of breast cancer histopathology images—a comparative insight," J. Digit. Imaging, vol. 33, no. 3, pp. 632–654, Jun. 2020. https://doi.org/10.1007/s10278-019-00307-y
[14] A. S. Assiri et al., "Breast tumor classification using ensemble SVM," J Imaging, vol. 6, no. 6, p. 39, 2020. https://doi.org/10.3390/jimaging6060039
[15] A. T. Alhasani et al., "A comparative analysis of methods for detecting and diagnosing breast cancer based on data mining," Methods, vol. 7, no. 9, pp. 1–10, 2023. https://doi.org/10.54216/JAIM.040201
[16] C. Cortes and V. Vapnik, "Support-vector networks," Mach Learn, vol. 20, no. 3, pp. 273–297. https://doi.org/10.1007/BF00994018
[17] Guido, Rosita, et al. "An overview on the advancements of support vector machine models in healthcare applications: a review." Information 15.4 (2024): 235. https://doi.org/10.3390/info15040235
[18] Dinesh, Paidipati, A. S. Vickram, and P. Kalyanasundaram. "Medical image prediction for diagnosis of breast cancer disease comparing the machine learning algorithms: SVM, KNN, logistic regression, random forest and decision tree to measure accuracy." AIP Conference Proceedings. Vol. 2853. No. 1. AIP Publishing LLC, 2024. https://doi.org/10.1063/5.0203746
[19] X. Dai, L. Xiang, T. Li, and Z. Bai, "Cancer hallmarks, biomarkers and breast cancer molecular subtypes," J. Cancer, vol. 7, no. 10, pp. 1281–1294, Jun. 2016. https://doi.org/10.7150/jca.13141
[20] N. Dey et al., Machine Learning in Bio-Signal Analysis and Diagnostic Imaging. Springer, 2019. https://doi.org/10.1016/B978-0-12-816086-2.00004-7
[21] J. S. Ahn et al., "Artificial intelligence in breast cancer diagnosis and personalized medicine," J. Breast Cancer, vol. 26, no. 5, p. 405, Oct. 2023. https://doi.org/10.4048/jbc.2023.26.e45
[22] C. Gain et al., "Variability in mitosis counting: Implications for ML," Bioorg Med Chem, vol. 37, p. 116112, 2021. https://doi.org/10.1016/j.bmc.2021.116112
[23] S. A. Muneam et al., "Clinical evaluation of liver function tests and carcinoembryonic antigen levels in colorectal cancer associated with hepatic metastases," Govaresh, vol. 30, no. 2, pp. 82–91, Summer 2025. http://www.govaresh.org/index.php/dd/article/view/2758
[24] S. M. Lundberg and S. I. Lee, "A unified approach to interpreting ML models," Adv Neural Inf Process Syst, vol. 30, pp. 4765–4774, 2017. https://doi.org/10.48550/arXiv.1705.07874
[25] M. T. Ribeiro, S. Singh, and C. Guestrin, "Why should I trust you?: Explaining the predictions of any classifier," KDD, pp. 1135–1144, 2016. https://doi.org/10.1145/2939672.2939778
[26] M. Altalhan, A. Algarni, and M. T. Alouane, "Imbalanced data problem in machine learning: A review," IEEE Access, vol. 13, pp. 13686–13699, Jan. 2025. https://doi.org/10.1109/ACCESS.2025.3531662
[27] S. Patel, Z. Hassan, S. Iniyan, and U. Desai, "Multi cancer prediction using deep learning and cnn algorithm," in 2024 Second International Conference on Inventive Computing and Informatics (ICICI), Jun. 2024, pp. 214–221. https://doi.org/10.1109/ICICI62254.2024.00044
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Al-Khwarizmi Engineering Journal

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright: Open Access authors retain the copyrights of their papers, and all open access articles are distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided that the original work is properly cited. The use of general descriptive names, trade names, trademarks, and so forth in this publication, even if not specifically identified, does not imply that these names are not protected by the relevant laws and regulations. While the advice and information in this journal are believed to be true and accurate on the date of its going to press, neither the authors, the editors, nor the publisher can accept any legal responsibility for any errors or omissions that may be made. The publisher makes no warranty, express or implied, with respect to the material contained herein.






