From Text to Actuarial Modelling: The Role of NLP and LLMs in Cyber Risk Assessment

  • 4 views

  • 0 comments

  • 0 favorites

  • AAE AAE
  • 249 media
  • uploaded July 31, 2026

Cyber risk has become one of the main challenges facing the insurance industry today. For the seventh consecutive year, cyber risk has been ranked as the Number 1 concern for the insurance sector, ahead of climate risk by France Assureurs in the Prospective Mapping Report 2025. As an emerging risk, it is characterized by scarce historical data, high claim heterogeneity, and the occurrence of extreme events with major financial impacts.   To address these challenges, this research explores the potential of textual data as a novel source of actuarial information. The Privacy Rights Clearinghouse (PRC) database, which has recorded thousands of data breach incidents since 2005, provides a key foundation for analysis. Prior work by Kher, Lopez and Rapior (2023) demonstrated that textual incident descriptions can be leveraged through Natural Language Processing (NLP) and neural networks to assess claim severity even in the absence of quantitative information. Using the updated PRC 2025 extraction, this study extends the analysis by mobilizing Artificial Intelligence and Large Language Models (LLMs) to structure and exploit unstructured text, thereby improving the actuarial modelling of cyber claims. The main methodological steps include:

  • A comparative analysis of PRC databases (2019 vs 2025), data harmonization, and the application of Extreme Value Theory to characterize cyber loss severity; 
  • The classification of incidents using machine learning algorithms; 
  • Severity modelling through sequential neural networks (LSTM), which outperform classical models (logistic regression, SVM, random forests, XGBoost), particularly for high-severity claims (starting at the 95th–99th percentiles of the severity distribution); 
  • The generation of synthetic incident descriptions using LLMs to enrich training datasets and simulate extreme scenarios; 
  • The evaluation of model robustness and the contribution of synthetic data to improving predictive effectiveness. 

Results confirm the significant contribution of NLP and Deep Learning to cyber risk quantification. While traditional models remain suitable for moderate losses, LSTM architectures prove more effective at identifying and characterizing severe claims. The integration of feature extraction through regular expressions, and the use of generative LLMs further enhance model accuracy and robustness. Beyond its technical dimension, this workshop illustrates that exploiting textual data represents a strategic opportunity for the insurance sector: it enables richer claim databases, faster and more accurate loss assessment upon incident notification, and improved prudential anticipation (Solvency II, ORSA). Ultimately, this approach highlights the complementarity between actuarial expertise and data science, fostering a deeper understanding of cyber risks and strengthening the resilience of the insurance industry in the face of rapidly evolving digital threats.

Tags:

Additional files

More Media in "DATA SCIENCE / AI"

0 Comments

There are no comments yet. Add a comment.