Aller au contenu principal
Accès ouvert déclaré 2026 article

Hybrid Lexical–Semantic AI Architecture for Automated Cancer Registry Coding for the Vet-ICD-O-Canine-1 System from Free-Text Veterinary Pathology Reports

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : br, pt. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Background/Objectives: Free-text veterinary pathology diagnoses contain essential information for cancer registration but are difficult to convert into standardized ontology-based codes because of linguistic variability, contextual modifiers, and large ontology search spaces. This study evaluated a hybrid lexical–semantic architecture for the automated assignment of Vet-ICD-O-Canine-1 morphology codes. Methods: A retrospective single-registry benchmark included 211 diagnoses from the São Paulo Animal Cancer Registry. Of these, 190 contained sufficient morphological information for expert-reviewed reference coding, whereas 21 generic or insufficiently specified descriptions were retained as an exploratory challenge subset. Fuzzy lexical matching retrieved Top-10, Top-20, or Top-30 candidates from the complete 971-entry morphology ontology, followed by semantic selection using Claude Haiku 4.5 and structured JSON output. Performance and computational efficiency were compared to direct full-ontology inference. Results: Among the evaluated fuzzy metrics, token_set_ratio achieved the highest Top-30 reference-code retrieval rate of 89.5%. End-to-end exact-match agreement increased from 73.7% with Top-10 to 79.5% with Top-20 and 85.8% with Top-30 (95% CI, 80.1–90.0%). Top-30 generated non-null codes for 93.2% of the 190 evaluable diagnoses and achieved a conditional exact-match agreement of 92.1%. By contrast, the direct full-ontology baseline achieved 71.2% conditional exact-match agreement (42/59) among non-null predictions and 22.1% end-to-end exact-match agreement (42/190) when incorrect predictions, null outputs, and technical failures were considered non-concordant outcomes. Compared to direct full-ontology inference, Top-30 reduced input-token consumption by 92.4%, total token consumption by 92.2%, and inference cost by 91.3%, while avoiding the 118 API rate-limit failures observed with the direct baseline. Among the 21 insufficiently specified diagnoses, Top-30 returned null codes in 38.1% and non-null codes in 61.9%. Conclusions: Ontology-guided candidate reduction improved coding agreement, computational efficiency, and operational robustness within this retrospective single-registry benchmark. However, the reported performance estimates require confirmation in larger independent datasets, and an upstream data-sufficiency or abstention mechanism is needed before prospective operational deployment.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Hybrid Lexical–Semantic AI Architecture for Automated Cancer Registry Coding for the Vet-ICD-O-Canine-1 System from Free-Text Veterinary Pathology Reports
Date Crossref
23/08/2026
Éditeur
MDPI AG
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Biomedical Text Mining and OntologiesAI in cancer detectionData-Driven Disease Surveillance

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.