Latent persona coordination as an attack surface in large language models
Rattachement africain : it, jp. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Large language models are often secured primarily at the level of outputs. This Perspective argues that unauthorized use and manipulation should also be treated as attacks on latent persona coordination: a relational property of the internal state, comprising the relative dominance of assistant-like, truth-preserving, and safety-preserving representations over competing dispositions, the conflict between incompatible dispositions that are active at once, and the stability of that configuration under perturbation. Reframing jailbreaks, malicious fine-tuning, hidden-signal training and uncensoring as reweightings of this control state yields a testable prediction: that latent measurements taken after different attacks are better explained by a general drift component plus pathway-specific residuals than by pathway-specific effects alone, and that the resulting signatures could flag manipulation before unsafe outputs appear. We state what would refute this.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Latent persona coordination as an attack surface in large language models
- Date Crossref
- 12/09/2026
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.