Aller au contenu principal
Accès ouvert déclaré 2026 article

Trajectory-Level Runtime Integrity Monitoring for Deployed Conversational LLM Services

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Large language models (LLMs) are increasingly deployed as stateful, networked services in decision-support and automation pipelines. Because dialogue history functions as an implicit runtime state, service assurance extends beyond single-turn correctness to the question of whether observable conversational trajectories remain stable, auditable, and measurable over sustained interaction under opaque-model access. This paper studies transcript-derived conversational deviation under policy-compliant multi-turn interaction. We do not treat such deviation as direct evidence of behavioral instability, safety degradation, or system compromise. Instead, we examine Conversational State Drift (CSD) as a constrained transcript-level observability quantity that can be measured from externally visible dialogue histories. We define a transcript-derived state representation and quantify CSD as cosine displacement from a fixed reference state. We further derive session-level summaries capturing terminal drift, temporal accumulation, and trajectory volatility, together with benign-calibrated descriptive indicators. Controlled experiments on GPT-4o and GPT-4o-mini show that structured sessions can exhibit higher transcript-derived drift than matched benign controls under the primary operating point. However, additional validity analyses substantially narrow the interpretation of this separation. Construct-validity controls show that transcript length is a dominant shortcut feature, although drift retains limited residual information under controlled comparisons. Drift-only holdout evaluation weakens markedly, indicating limited transfer beyond the original operating point. Backend comparisons further show that separability depends strongly on representation choice rather than reflecting a backend-invariant semantic property. Probe-based behavioral analysis does not reveal a reliable predictive relationship between drift and downstream consistency outcomes. Overall, transcript-derived long-horizon drift is best interpreted as a limited, operating-point-dependent observability feature. It may support conservative auditing and diagnostic monitoring under controlled opaque-model settings, but it should not be treated as a strong standalone detector, a behaviorally validated instability measure, or a representation-invariant signal of runtime integrity loss.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Trajectory-Level Runtime Integrity Monitoring for Deployed Conversational LLM Services
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Software System Performance and ReliabilityDistributed systems and fault toleranceService-Oriented Architecture and Web Services

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.