Additive Tree Latent Variable Models for Insurance Loss Prediction

Authors

  • Jiajun Li School of Mathematics and Statistics, Ningbo University, Ningbo, Zhejiang, 315211, China

DOI:

https://doi.org/10.54097/94tgta52

Keywords:

Additive Tree Models, Latent Variable Models, Insurance Loss Prediction, Gradient Boosting, Tweedie Distribution

Abstract

Accurate prediction of insurance losses is essential for premium pricing, reserve estimation, and risk management in the property and casualty insurance industry. Traditional actuarial models, such as generalized linear models, often rely on predefined distributional assumptions and struggle to capture complex nonlinear interactions among risk factors. Meanwhile, tree based ensemble methods have demonstrated strong predictive power but typically lack the ability to model latent heterogeneity in policyholder risk profiles. This paper proposes an Additive Tree Latent Variable (ATLV) model that integrates gradient boosted additive tree structures with a latent variable framework to improve insurance loss prediction. A Latent Risk Index (LRI) is developed to capture unobserved policyholder risk heterogeneity through variational inference. A Tree Ensemble Loss Predictor (TELP) is formulated to combine the latent representations with observed covariates in a gradient boosting framework. A Distributional Calibration Module (DCM) is introduced to ensure that predicted loss distributions are well calibrated under the Tweedie compound Poisson family. Empirical analysis is conducted on the French Motor Third Party Liability (MTPL) insurance dataset comprising 678,013 policies. The results demonstrate that the ATLV model significantly outperforms both traditional generalized linear models and standard gradient boosting approaches across multiple evaluation metrics. The ATLV model achieves a mean Deviance of 1.247, a Gini coefficient of 0.312, and a calibration error of 0.034, representing improvements of 8.6%, 14.3%, and 41.4% respectively over the best performing baseline. An ablation study further reveals that the LRI module contributes most significantly to overall performance, confirming the importance of modeling latent risk heterogeneity in insurance loss prediction. These findings provide both theoretical insights and practical tools for actuaries seeking to enhance loss prediction accuracy through the integration of machine learning and latent variable modeling.

Downloads

Download data is not yet available.

References

[1] Frees, E. W., Derrig, R. A., & Meyers, G. (2014). Predictive modeling applications in actuarial science, Volume 1: Predictive modeling techniques. Cambridge University Press.

[2] De Jong, P., & Heller, G. Z. (2008). Generalized linear models for insurance data. Cambridge University Press.

[3] Wuthrich, M. V., & Merz, M. (2023). Statistical foundations of actuarial learning and its applications. Springer.

[4] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

[5] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451

[6] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785

[7] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T. Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (Vol. 30, pp. 3146–3154). Curran Associates.

[8] Yang, Y., Qian, W., & Zou, H. (2018). Insurance premium prediction via gradient tree-boosted Tweedie compound Poisson models. Journal of Business & Economic Statistics, 36(3), 456–470. https://doi.org/10.1080/07350015.2017.1345664

[9] Shi, P., & Valdez, E. A. (2011). A copula approach to test asymmetric information with applications to predictive modeling. Insurance: Mathematics and Economics, 49(2), 226–239. https://doi.org/10.1016/j.insmatheco.2011.03.007

[10] Tipping, M. E., & Bishop, C. M. (1999). Probabilistic principal component analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61(3), 611–622. https://doi.org/10.1111/1467-9868.00196

[11] Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. In Proceedings of the 2nd International Conference on Learning Representations.

[12] Jorgensen, B. (1997). The theory of dispersion models. Chapman and Hall.

[13] Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer.

[14] Blei, D. M., Kucukelbir, A., & McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American Statistical Association, 112(518), 859–877. https://doi.org/10.1080/01621459.2017.1285773

[15] Noll, A., Salzmann, R., & Wuthrich, M. V. (2020). Case study: French motor third-party liability claims. SSRN Electronic Journal. https://doi.org/xxxx

Downloads

Published

09-07-2026

Issue

Section

Articles

How to Cite

Li, J. (2026). Additive Tree Latent Variable Models for Insurance Loss Prediction. Academic Journal of Applied Sciences, 2(2), 48-54. https://doi.org/10.54097/94tgta52