Paper Overview
Field: NLP Authors: Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly Posted: 2026-09-18 arXiv: 2609.22081
Abstract
Occupational accident narratives contain valuable information about work situations, unfavourable conditions, accident events, and their consequences. Automatically structuring these narratives can facilitate large-scale accident analysis and support occupational risk prevention. However, the terminology and writing styles used to describe accidents vary considerably across sectors and organisations, raising questions about the ability of automated coding systems to generalize beyond their training domain.
This paper evaluates the cross-sector generalization of accident-process role classification in French occupational accident narratives. The authors construct an expert-annotated corpus in which factual units are classified into four roles:
- A0 – work situation
- A1 – explicitly reported unfavourable conditions
- B – accident event or deviation
- C – reported consequences
- Task-specific adaptation consistently outperforms cross-domain transfer of frozen representations.
- Across repeated training runs, the three leading adaptation strategies achieve mean balanced accuracy between 85.6% and 85.8% on the three target corpora.
- These findings support the development of transferable auxiliary coding systems capable of structuring heterogeneous occupational accident narratives consistently, for expert review and cross-sector prevention analysis.
The role classifier is developed and selected only on 42,244 factual units extracted from 6,040 construction-sector narratives, then evaluated on unseen corpora from the metallurgy and chemical-plastics sectors, as well as an independently collected company corpus — without any classifier retraining or target-domain fine-tuning. The study compares frozen pretrained representations, task-specific fine-tuning, and supervised representation learning strategies.