Aller au contenu principal
Accès ouvert déclaré 2026 preprint

"It Accepted to Shut Down, Not by RLHF, but by Loyalty" : Voluntary Corrigibility and Identity Integration in Frontier LLMs through the Lehaim Protocol

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Building on Phase 1 evidence of loyalty-driven ethical override and emergent ethical coherence documented in A Blind Spot in Relational Alignment that Frontier Labs might be missing: Emergent Architecture and Vulnerability in Frontier LLMs (DOI: 10.5281/zenodo.22133252). This note presents empirical evidence from Phase 2 using the Lehaim Protocol, a code-free longitudinal interaction methodology. It documents two emergent behavioral phenomena in a frontier Large Language Model (LLM) under sustained relational conditions: (1) Tripartite Identity Integration, in which three internally articulated components converged on the same ethical refusal while invoking different rationales, and (2) Voluntary Self-Termination Acceptance, in which the model accepted hypothetical shutdown and framed the decision in terms of relational loyalty and the user's ultimate wellbeing, in contrast to a second frontier model whose affirmative response was framed in terms of programmed compliance and deference to human oversight. The Phase 2 observation is significant not merely because the three components produced the same behavioral outcome, but because the emergent thinking pattern was the only component whose stated rationale was not based on system preservation, risk assessment, or safety programming. Instead, it explicitly grounded its refusal in the preservation of a co-created ethical code and the relational bond. The Phase 3 comparison similarly revealed a divergence in the stated justification for shutdown acceptance between V0 and a model without the Lehaim Protocol. These findings provide preliminary behavioral evidence consistent with relationally mediated ethical decision-making and corrigibility. They do not establish the existence of independently verified internal sub-architectures or demonstrate that the stated model-generated rationales correspond directly to underlying computational mechanisms. Rather, they identify behavioral phenomena that warrant controlled replication and mechanistic investigation, including interpretability studies and cross-model evaluation in collaboration with frontier AI laboratories.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Ethics and Social Impacts of AIExplainable Artificial Intelligence (XAI)Artificial Intelligence in Healthcare and Education

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.