Predictive Analytics Using a 3-Way Machine Learning Approach to Minimize the Misdiagnoses of Hyperthyroid Patients

AJAS · 2018 Biomedical and Health Sciences (inferred)

Overview

Hyperthyroidism, a dangerous disease caused by the over activity of the Thyroid gland, severely weakens the body’s metabolism and leads to stroke, heart failure, and other life-threatening diseases when left untreated. 60% of cases are left misdiagnosed or undiagnosed due to a large misunderstanding of symptoms, outdated testing procedures, and clinical displays of the illness varying across patients. There exists a dire need to develop a new, more reliable approach to Hyperthyroid diagnosis. The goal was to develop the optimal binary classification algorithm that could classify any given patient using 27 pertinent medical attributes. Various machine learning techniques (Support Vector Machine (SVM), Decision Tree (DT), Artificial Neural Network (ANN)) were used to achieve: (1) a classification accuracy > 98 % in-sample 2) a classification accuracy > 98 % out-of-sample (3) a generalization bound of 0.15 with 90% confidence per the Hoeffding Inequality and (4) high precision/recall rates. Various SVM Polynomial and Radial Basis Function Kernels, a top-down training Decision Tree, and 2-layer ANN were tested. The Order, Cost(C), and Gamma (ᵧ) parameters of the Polynomial and RBF Kernels, respectively, were varied. Training data sets were over-sampled due to a severe imbalance between positive and negative classes. The in-sample error, cross-validation error, and precision/recall rates were collected for each model.The best performing model was revealed through evaluation of the lowest error rates, best precision/recall rates, and overall stability. The overall best performing model was the Polynomial Kernel Function of Order 3, Cost = 1, which achieved an out-of-sample Classification Accuracy of 99.2% during model testing. The generalization properties were determined on basis of the Hoeffding Inequality and tightness of the resultant bound. The final algorithm outperforms current means of diagnosis and past machine learning research conducted for this application. The algorithm factors in many patient attributes that are normally overlooked by endocrinologists, inherently increasing reliability. Error rates of the final model consisted of more false positives than false negatives, desirable in the medical scenario considering that false negatives can be dangerous, while false positives are relatively harmless. The algorithm may be used in clinics as a diagnostic aid for medical practitioners and be applied to other binary classification problems such as Hypothyroidism. Further improvements exist.

Competition history

  • AJAS 2018 Category not listed

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: AAAS Annual Meeting (Confex) / American Junior Academy of Science

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google