PENERAPAN ALGORITMA K-NEAREST NEIGHBOR DENGAN REDUKSI DIMENSI BERBASIS ANALISIS KOMPONEN UTAMA DALAM MENDETEKSI KANKER PAYUDARA
Abstract
Breast cancer is basically curable if detected early. However, 70% of breast cancer patients were only detected when it was already severe, making it incurable and leading to death. In an effort to overcome this problem, this study aims to classify data from routine health checks to detect breast cancer. The data taken from Coimbra University Hospital, Portugal, is classified into healthy or breast cancer classes using the K-Nearest Neighbor with dimension reduction based on Principal Component Analysis. The variables used are age, BMI, glucose, insulin, HOMA, leptin, adiponectin, resistin, and MCP.1. Results show that the Principal Component Analysis reduced nine attributes down to four with a cumulative variance proportion of 76.15%. The dominating variables in the four principal components in descending order were insulin, leptin, and glucose. The data transformed into the principal component axis was then classified using the K-Nearest Neighbor algorithm. The performance of K-Nearest Neighbor in this paper reaches the highest result when k=5, where the F1 score is 72%. This means that the performance of the model in considered good in finding most positive cases (whether a patient has breast cancer) without making too many mistakes.
Keywords: Breast cancer, classification, K-Nearest Neighbor, F1 score, Principal Component Analysis
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Interval : Jurnal Ilmiah Matematika

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.





