Aller au contenu principal
Accès ouvert déclaré 2026 article

From study design to executable code: automating target trial emulation with large language models

0Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : kr, us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Objective: Implementing target trial emulation (TTE) studies as standardized, reproducible analytic workflows is technically demanding. We developed Text-guided Health-study Estimation and Specification Engine Using Strategus (THESEUS), which uses large language models (LLMs) to translate free-text study descriptions into structured analytic specifications and Strategus R scripts within the Observational Health Data Sciences and Informatics (OHDSI) ecosystem. Materials and Methods: THESEUS executes 2 steps: an LLM maps study descriptions to a JavaScript Object Notation (JSON) schema, and validated specifications are converted into Strategus R scripts through rule-based logic. For standardization evaluation, we compared specifications generated by 8 LLMs using 15 OHDSI-based TTE studies and 15 non-OHDSI studies under primary-analysis and full-analyses settings. Results: Under the primary-analysis setting, overall standardization accuracy ranged from 0.93 to 0.97 across models in OHDSI studies and from 0.82 to 0.95 in non-OHDSI studies. Gemini-3.1-Pro achieved the highest overall accuracy in OHDSI studies, while Gemini-3.1-Pro and Gpt-5.5 jointly achieved the highest overall accuracy in non-OHDSI studies. Under the full-analyses setting, field-level sensitivity ranged from 0.83 to 0.97 in OHDSI studies, with 0.07-0.80 false positives (FPs) per study, and from 0.77 to 0.89 in non-OHDSI studies, with 0.53-1.20 FPs per study. Gpt-5.5 performed best at the field level. THESEUS was implemented as a web application and coding-agent tools. Discussion: Pairing a standardized data model with a structured analysis framework enables reliable LLM-assisted interpretation of study descriptions and deterministic workflow construction in observational research. Conclusion: THESEUS supports translation of natural language study descriptions into executable, shareable code in standardized observational research settings.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
From study design to executable code: automating target trial emulation with large language models
Date Crossref
07/07/2026
Éditeur
Oxford University Press (OUP)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Yonsei University Department of Biomedical Systems Informatics pays non établi dans la notice
    Université ou école supérieure
  • University Health System pays non établi dans la notice
    Établissement de santé
  • Yonsei University Health System pays non établi dans la notice
    Établissement de santé

Department of Biomedical Systems Informatics — Yonsei University, University Health System et Yonsei University Health System.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Scientific Computing and Data ManagementResearch Data Management PracticesMeta-analysis and systematic reviews

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.