News

EAA Web Session 'Generative AI for Synthetic Actuarial Data' on 10 November 2026

 

 

 

­

­

EAA Web Session 

­

­

 

Generative AI for Synthetic Actuarial Data

­

10 November 2026 | 9:00-13:30 CET

­

 

Access to realistic, publicly available datasets is a significant barrier to advancing the development of assets for insurance analytics and actuarial research. A substantial amount of data in the insurance industry is confidential and proprietary, making it challenging to validate new methodologies. Even when data owners are willing to share anonymised datasets, the effort required to mask the data and navigate bureaucratic processes can often prevent disclosure. This environment limits opportunities to measure progress in the field.

 

To address the scarcity of real insurance data, there is an opportunity to develop simulated datasets that capture features similar to those observed in practice, thereby reducing friction in data disclosure for development and research. Although synthetic data generation is an active area within the broader machine learning community, its specific application to insurance data remains underexplored.

This web session is designed to demonstrate how to build synthetic insurance datasets using various generative models.

Several models will be employed, starting with Gaussian Mixture Models and extending to the conditional format for the typical non-life insurance dataset used to model frequency and severity. Results will be compared with those of the Conditional Diffusion Model. The seminar will explore the Variational Autoencoder and Generative Adversarial Networks using other insurance datasets. Lastly, we will also see the use of Large Language Models to generate synthetic data by exploiting prompt engineering. 

The data generation process will be conducted by running several trials: reproducing the full number of records in the data fed into the generative models, applying data augmentation, and then omitting sensitive variables to preserve privacy.

Several techniques will be used for the validation of the generated datasets:

  • Consistency Tests: to verify that synthetic insurance records strictly adhere to industry business logic.
  • Kolmogorov-Smirnov Test: to ensure that the generated data maintains the underlying statistical properties of the original datasets.
  • Data Visualization: using univariate analysis, correlation matrices, and other spatial techniques.
  • Actuarial Modelling: evaluating predictive utility also through feature importance comparisons.

Early-bird discount is available for bookings made by 29 September 2026.

 

 

 

­

Coming soon...

 

CERA, Module B: Taxonomy, Modelling and Mitigation of Risks | 7-10 September 2026

Fit4AI Compact | 6/7 October 2026

Special Actuarial Topics in Cyber (Re)Insurance | 4 November 2026

Explore our website for more information and discover all our upcoming events. For more insights, updates, and a bit of actuarial fun, feel free to follow us on LinkedIn!