Categories
- DATA SCIENCE / AI
- AFIR / ERM / RISK
- ASTIN / NON-LIFE
- BANKING / FINANCE
- DIVERSITY & INCLUSION
- EDUCATION
- HEALTH
- IACA / CONSULTING
- LIFE
- PENSIONS
- PROFESSIONALISM
- THOUGHT LEADERSHIP
- MISC
ICA LIVE: Workshop "Diversity of Thought #14
Italian National Actuarial Congress 2023 - Plenary Session with Frank Schiller
Italian National Actuarial Congress 2023 - Parallel Session on "Science in the Knowledge"
Italian National Actuarial Congress 2023 - Parallel Session with Lutz Wilhelmy, Daniela Martini and International Panelists
Italian National Actuarial Congress 2023 - Parallel Session with Kartina Thompson, Paola Scarabotto and International Panelists
1 views
0 comments
0 likes
0 favorites
AAE
Data quality and availability are central challenges in statistical and machine learning, particularly when modeling rare events. This work investigates the use of synthetic data generation as a means to enhance the performance of standard learning approaches in imbalanced settings, covering both classification and regression tasks. Generating realistic synthetic tabular data is inherently difficult, as it requires faithfully capturing complex inter-variable relationships. While generation can be performed directly in the original feature space, we demonstrate that representation learning offers significant advantages by uncovering non-linear correlations and improving the overall quality of the generated samples. We first provide a comprehensive analysis of data imbalance, characterizing its impact both empirically and theoretically. We then propose several data generation strategies designed to mitigate this imbalance and improve model accuracy and generalization. By extending imbalanced data management beyond binary classification to regression and more complex real-world settings where rare outcomes are often the most consequential, this work opens new perspectives for statistical modeling and machine learning in high-stakes applications.
0 Comments
There are no comments yet. Add a comment.