Implementation Of Data Mining For Inpatient Patient Data Clustering Using The K-Means Clustering Algorithm
Main Article Content
Abstract
The large availability of electronic medical record inpatient data at Wajo Public Health Center has not been optimally utilized and is only used for administrative purposes and routine reporting. The utilization of these data has not been directed toward analyzing patient characteristic patterns to support strategic health service planning. This study aims to implement data mining techniques using the K-Means Clustering algorithm to group inpatient data at Wajo Public Health Center to map patients' demographic and clinical characteristics in a structured manner.The research method used is a quantitative approach with a cluster analysis design. Data were extracted from the EMR system of Wajo Public Health Center in 2026, with a total sample of 1,527 patient records after undergoing preprocessing stages including data cleaning, attribute selection (age and diagnosis), and data transformation using Python programming in Google Colab. The determination of the optimal number of clusters was evaluated using the Elbow Method.The results showed that the most optimal number of clusters is 4 clusters ($K=4$). Cluster 0 (217 patients) is dominated by adults with diverse complaints (abdominal colic, vertigo, partus, ARI). Cluster 1 is the largest group (553 patients) dominated by adults with digestive system disorders (dyspepsia and GEA). Cluster 2 (449 patients) is dominated by infants and children with acute infectious disease patterns (GEA, ARI, febris). Cluster 3 (308 patients) is dominated by adult patients with a combination of acute and chronic/metabolic diseases (dyspepsia, GEA, ARI, diabetes, UTI). In conclusion, the K-Means algorithm is effective in grouping complex medical record data into structured information as a basis for targeted health service planning recommendations