Trajectory-Level Runtime Integrity Monitoring for Deployed Conversational LLM Services
Résumé fourni par la source
Large language models (LLMs) are increasingly deployed as stateful, networked services in decision-support and automation pipelines. Because dialogue history functions as an implicit runtime state, service assurance extends beyond single-turn correctness to the question of whether observable conversational trajectories remain stable, auditable, and measurable over sustained interaction under opaque-model access. This paper studies transcript-derived conversational deviation under policy-compliant multi-turn interaction. We do not treat such deviation as direct evidence of behavioral instability, safety degradation, or system compromise. Instead, we examine Conversational State Drift (CSD) as a constrained transcript-level observability quantity that can be measured from externally visible dialogue histories. We define a transcript-derived state representation and quantify CSD as cosine displacement from a fixed reference state. We further derive session-level summaries capturing terminal drift, temporal accumulation, and trajectory volatility, together with benign-calibrated descriptive indicators. Controlled experiments on GPT-4o and GPT-4o-mini show that structured sessions can exhibit higher transcript-derived drift than matched benign controls under the primary operating point. However, additional validity analyses substantially narrow the interpretation of this separation. Construct-validity controls show that transcript length is a dominant shortcut feature, although drift retains limited residual information under controlled comparisons. Drift-only holdout evaluation weakens markedly, indicating limited transfer beyond the original operating point. Backend comparisons further show that separability depends strongly on representation choice rather than reflecting a backend-invariant semantic property. Probe-based behavioral analysis does not reveal a reliable predictive relationship between drift and downstream consistency outcomes. Overall, transcript-derived long-horizon drift is best interpreted as a limited, operating-point-dependent observability feature. It may support conservative auditing and diagnostic monitoring under controlled opaque-model settings, but it should not be treated as a strong standalone detector, a behaviorally validated instability measure, or a representation-invariant signal of runtime integrity loss.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Trajectory-Level Runtime Integrity Monitoring for Deployed Conversational LLM Services
- Date Crossref
- 01/01/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.