Data generation method for disease prediction on imbalanced data
Gai, Ruochen (2024)
Kandidaatintyö
Gai, Ruochen
2024
School of Engineering Science, Tietotekniikka
Kaikki oikeudet pidätetään.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe2024051530808
https://urn.fi/URN:NBN:fi-fe2024051530808
Tiivistelmä
Disease prediction plays an increasingly important role in practical application, but also brings some challenges. In real situations, many diseases, especially rare diseases, often have the problem that there are fewer cases than controls, which brings a great negative impact on disease prediction. Due to the shortcomings of traditional solutions, generative adversarial network (GAN), a new data augmentation method, has been invented. With the deepening of research, GAN has many variants with better performance. This research is dedicated to building GAN and its variants to alleviate the data imbalance problem in disease prediction, thereby improving the performance of the prediction model. After reviewing previous literature, the research uses GAN, CGAN, and WGAN models to conduct experiments. Find the experimental data set on relevant data platforms, use PyTorch to build these models to generate minority category data, and then use XGboost to predict the generated data. Finally, performance indicators are used to evaluate these models, and the model with the strongest ability to alleviate data distribution imbalance is determined through comparison.
