Managing Uncertainty in AI-Enabled Occupational Injury Claims Review: An Observed-to-Expected Interpretation Framework (Preprint)
Le résumé fourni par la source
BACKGROUND Health and insurance administrators increasingly use statistical models and artificial intelligence methods to identify injury or work categories for closer review. An observed-to-expected (O/E) ratio compares total recorded sickness absence days in a category with the total expected by a prediction model. A ratio greater than 1 indicates more recorded than expected sickness absence. However, it does not explain why the difference occurred or what should be reviewed next. Interpretation may depend on the prediction model, the work information defining the expected comparison, and whether the difference is concentrated among claims with very long durations. OBJECTIVE This study aimed to identify occupational injury categories with more recorded sickness absence than expected, compare their O/E results across prediction models and definitions of the expected comparison, and determine whether positive individual claim differences were concentrated among claims with very long durations. METHODS We conducted a retrospective observational study of 86,305 closed occupational injury insurance claims in Hong Kong from 2005 to 2024. Repeated cross fitting with Tweedie generalized linear models estimated expected sickness absence for each claim. Five injury and work category groupings were examined. For each grouping, its defining variable was omitted, so expected duration was based on other recorded characteristics. Categories required an O/E ratio greater than 1.10, a mean recorded minus expected difference of at least 5 days per claim, and at least 500 claims. The selected categories were assessed by replacing the Tweedie model with CatBoost after overall scale alignment, omitting available work information from the expected duration model, and estimating the share of summed positive individual claim differences contributed by claims above the empirical 95th percentile of recorded duration. The first two assessments used a predefined criterion combining relative and absolute changes. RESULTS Overall recorded and expected sickness absence totals were closely aligned, but 14 categories met all selection criteria. Although CatBoost had lower overall prediction error, only 2 of 14 category comparisons met the predefined change criterion after replacing the Tweedie model with CatBoost. Omitting available work information met the criterion in 6 comparisons. In 10 of 14 selected categories, claims in the upper 5% of recorded durations contributed a greater share of summed positive individual claim differences than the corresponding share in the full cohort. CONCLUSIONS An elevated O/E ratio can identify categories for closer review but cannot explain the observed difference or determine an organizational response. Examining the same selected categories through these three questions separated category selection from subsequent interpretation. These questions provide a transparent way to examine additional evidence before comparisons generated by models are interpreted in organizational review.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Managing Uncertainty in AI-Enabled Occupational Injury Claims Review: An Observed-to-Expected Interpretation Framework (Preprint)
- Date Crossref
- 31/07/2026
- Éditeur
- JMIR Publications Inc.
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.