Analisis Perbandingan Algoritma K-Nearest Neighbor (K-NN) dan Random Forest dalam Klasifikasi Kanker Payudara
Abstrak
Abstrak
Kanker payudara merupakan salah satu penyakit yang memiliki tingkat kejadian dan kematian yang tinggi pada perempuan di seluruh dunia. Deteksi dini yang akurat sangat penting untuk meningkatkan peluang kesembuhan pasien dan mendukung pengambilan keputusan medis. Penelitian ini bertujuan untuk menganalisis dan membandingkan kinerja algoritma K-Nearest Neighbor (K-NN) dan Random Forest dalam klasifikasi kanker payudara menggunakan Breast Cancer Wisconsin (Diagnostic) Dataset. Penelitian menggunakan metode Cross Industry Standard Process for Data Mining (CRISP-DM) yang meliputi tahapan Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, dan Deployment. Dataset yang digunakan terdiri dari 569 data dengan 30 atribut numerik dan satu atribut target diagnosis. Data dibagi menjadi data latih sebesar 80% dan data uji sebesar 20%, serta dilakukan normalisasi menggunakan Min-Max Scaling. Evaluasi model dilakukan menggunakan metrik accuracy, precision, recall, F1-score, dan Area Under Curve (AUC). Hasil penelitian menunjukkan bahwa nilai k terbaik pada algoritma K-NN diperoleh pada k=5 dengan accuracy sebesar 96,49%, precision sebesar 100%, recall sebesar 90,48%, F1-score sebesar 95,00%, dan AUC sebesar 99,62%. Sementara itu, algoritma Random Forest memperoleh accuracy sebesar 97,37%, precision sebesar 100%, recall sebesar 92,86%, F1-score sebesar 96,30%, dan AUC sebesar 99,29%. Berdasarkan hasil evaluasi, Random Forest menunjukkan performa yang lebih baik dibandingkan K-NN karena memiliki nilai accuracy, recall, dan F1-score yang lebih tinggi. Oleh karena itu, Random Forest dapat menjadi algoritma yang lebih optimal untuk klasifikasi kanker payudara pada Breast Cancer Wisconsin (Diagnostic) Dataset.
Kata kunci: Kanker Payudara, Klasifikasi, Machine Learning, K-Nearest Neighbor, Random Forest, CRISP-DM.
Abstract
Breast cancer is a disease with a high incidence and mortality rate in women worldwide. Accurate early detection is crucial to improve the chances of patient recovery and support medical decision-making. This study aims to analyze and compare the performance of the K-Nearest Neighbor (K-NN) and Random Forest algorithms in breast cancer classification using the Breast Cancer Wisconsin (Diagnostic) Dataset. The study uses the Cross Industry Standard Process for Data Mining (CRISP-DM) method, which includes the stages of Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment. The dataset used consists of 569 data with 30 numeric attributes and one diagnostic target attribute. The data is divided into 80% training data and 20% testing data, and normalization is performed using Min-Max Scaling. Model evaluation is carried out using accuracy, precision, recall, F1-score, and Area Under Curve (AUC) metrics. The results showed that the best k value in the K-NN algorithm was obtained at k = 5 with an accuracy of 96.49%, precision of 100%, recall of 90.48%, F1-score of 95.00%, and AUC of 99.62%. Meanwhile, the Random Forest algorithm obtained an accuracy of 97.37%, precision of 100%, recall of 92.86%, F1-score of 96.30%, and AUC of 99.29%. Based on the evaluation results, Random Forest showed better performance than K-NN because it had higher accuracy, recall, and F1-score values. Therefore, Random Forest can be a more optimal algorithm for breast cancer classification in the Breast Cancer Wisconsin (Diagnostic) Dataset.
Keyword: Breast Cancer, Classification, Machine Learning, K-Nearest Neighbor, Random Forest, CRISP-DM.